September 27, 2026
Building a scalable file upload system
At first, people think uploading a file looks like a small backend job. You just receive it through an API, store it in object storage, and…

By Andrew Tinyakov
6 min read
At first, people think uploading a file looks like a small backend job. You just receive it through an API, store it in object storage, and return a URL.
That approach can actually work, but let's add large files, some image processing, and bursts of concurrent uploads. The problem is that API starts doing several tasks with different resource needs. Moving files ties up bandwidth, resizing burns CPU and memory, and everyday requests get stuck competing for what's left.
I built this project to demonstrate how exactly separating those responsibilities lets us effortlessly scale them independently. For implementation I have decided to use a Kotlin Spring Boot API, PostgreSQL, RabbitMQ, Go workers, and S3-compatible storage. Here, I want to explain how to build a file upload system that will actually scale as demand increases.
A practical way to approach this problem is to separate the responsibilities:
- Transferring file bytes to object storage.
- Preparing files, such as resizing images and generating download versions.
- Managing access, file metadata, and processing state.
Those are different problems that require different ways of scaling, so keeping them apart lets us provision each one on it's own.
Lets start with the transfer. The backend rarely needs to touch the raw bytes before they land in storage. It only should decide whether the user can upload, where the file should go, and what limits apply.
To let the client upload directly to storage, we can use presigned URLs. After checking the request, the API generates a time-limited URL for uploading to a specific object location. This removes the API from the upload data path, therefore, it no longer has to hold open upload connections.
For large files, multipart upload divides the transfer into separate parts to let the client upload parts concurrently and retry a failed part without restarting the whole transfer.
Another problem to solve is file processing. For instance, decoding, resizing, and encoding images can consume substantial CPU and memory. Running these operations inside the API makes them compete with requests that check permissions or return metadata.
To solve it we can have a separate worker processes which lets this work scale independently. If image processing falls behind, more image workers can consume jobs without requiring more API instances. Workers can also have different memory limits and concurrency settings from the API.
Another benefit of having workers as separate applications, is that we can choose a stack more suitable for CPU-bound work. We might benefit from Go, which is what this example uses, or Rust, both of which support running work across multiple CPU cores. However, the language doesn't really matter that much for acheaveing scalable architecture. Something like Node.js can be good enough for most cases. The main benefit comes from moving file processing out of the API so we can give it its own resources and scale it separately.
Now let's follow one image through the system.
The client first asks the API to create an upload session. It sends the file metadata like filename and expected size. The API creates records for a file asset (the file the application knows about) and an upload session (a representation of the attempt to supply its bytes to S3) in the database, then returns upload instructions.
The client uploads to a temporary uploads bucket and then tells the API that the transfer is complete. After that the API checks the object's metadata in storage, compares it with data given by client, marks the asset as preparing, and creates a job for asynchronous processing using transactional outbox pattern.
A background outbox publisher picks up pending messages, sends them to the message broker, and marks them as published after confirmation.
In this example, we use RabbitMQ. Redis Streams or Kafka could also carry the messages, though each comes with different delivery and operational choices. The idea of this architecture is to hand work to background consumers.
The workers listen for jobs, where each message tells a worker where to find the uploaded file, how to process it, and where to save the result. In this example I used 2 worker applications: an image and a document worker.
An image worker downloads a file, generates the requested versions, and saves them in the permanent assets bucket. Document worker has a separate queue and worker pool, where the file is copied unchanged. That also means a burst of image jobs doesn't put document jobs behind them in the same queue. This approach is particularly beenfitial if we later need video processing later since we can add a separate video worker pool and scale it independently.
Once processing finishes, the worker sends a completion or failure event back through the broker for the api to consume it and updates the database with the file's status and, in case of success, it saves prepared versions with storage locations.
How does the client know when the file is ready?
For this example, short polling is enough: the client checks the API periodically until processing finishes, after which it receives the file metadata, including the URLs. It can use those URLs to retrieve the file directly from S3 or a CDN if one is enabled. Other options are long polling, server-sent events or WebSockets.
At this point, we have the full flow. However, there are still a few things to account for before calling it done.
First, a queue gives us somewhere to hold work during a burst, but it doesn't make processing capacity unlimited. If jobs keep arriving faster than workers can finish them, users will wait longer. To handle it we can use queue depth, message age, processing time, and available worker capacity to decide when to add workers and when to scale back. Autoscaling isn't covered in this example, but at higher scale we need a way to adjust capacity as demand changes.
The second issue is failure handling. Temporary failures can be retried with backoff, and retries need a limit. Meanwhile, permanent errors should fail without retries so we can show a failure state. For instance, an unsupported image won't become supported after another attempt. Jobs that exhaust their retries need a failure path too, rather than leaving the client polling forever.
And last but not least, cleanup. A user can close the tab halfway through an upload, or processing can fail after producing some output. Without cleanup, those files and unfinished multipart uploads keep occupying storage.
This is another place where the outbox is useful. In this project, scheduled tasks find expired upload sessions and old failed assets in batches. For each batch uses one database transaction to update their states and inserts cleanup messages into the outbox. Background consumers then abort expired multipart uploads and delete the output files belonging to discarded assets from S3.
Storage lifecycle rules also remove temporary source files after a retention period and provide a fallback for abandoned multipart uploads.
As a side note, I want to mention that a smaller application can use the same separation with much less infrastructure. Direct uploads through presigned URLs don't require another service if you're already using S3. And if processing is needed, a single worker application can run alongside the API on the same VPS and moved to its own server when it needs more resources. You can keep these responsibilities separate without deploying the full setup described here.
At higher scale, that separation gives us control over how the system handles demand. More upload traffic goes directly to storage. More processing work can be spread across workers. The API coordinates the flow, while durable messages keep track of work that still needs to happen. We still have to measure capacity and respond to bottlenecks, but we can do that for each part independently. That's what makes this architecture powerful: we can change how much work the system handles without changing how a file moves through it.
Run it yourself.
I built a small client for the project so you can try the upload flow locally. Upload an image, watch it move from CREATED through UPLOADING and PREPARING to READY, and inspect the generated variants and the time spent at each stage.
The source code and local setup instructions are available in the repository.
Github: https://github.com/AndrewTinyakov/file-upload-system