September 29, 2026
Never Trust the Upload: R2 Event Notifications as a Validation Pipeline
EventAssets is an event media platform. Guests at a wedding scan a QR code, open a gallery, and upload photos and video from their phones…

By Tristan Trommer
8 min read
EventAssets is an event media platform. Guests at a wedding scan a QR code, open a gallery, and upload photos and video from their phones. A single event can take a thousand files, some of them a gigabyte each.
Those bytes do not pass through my API. They cannot — a Cloudflare Worker is not a file server, and proxying a gigabyte through one to put it in R2 would be absurd. So the client gets a presigned PUT URL and writes to R2 directly.
Which means that for a few seconds, a phone belonging to someone I have never met has write access to my bucket.
This post is about what happens next.
What the client actually controls
When a client asks to upload, it sends a declared size and a declared content type:
export const MediaFileUploadSchema = z.object({
file_size: z.number().positive().max(MEDIA_FILE_SIZE),
content_type: z.enum(MEDIA_MIME_TYPES)
});export const MediaFileUploadSchema = z.object({
file_size: z.number().positive().max(MEDIA_FILE_SIZE),
content_type: z.enum(MEDIA_MIME_TYPES)
});Both of those are claims. Zod will confirm that content_type is one of eight strings I allow and that file_size is a positive number under 1 GB, and that is the entire extent of what validation can do at this point, because the file does not exist yet. There are no bytes to inspect. The client is describing something it intends to upload.
A phone can say image/jpeg and PUT a 400 MB video. A script can say file_size: 1024 and PUT whatever it likes. The presigned URL does not enforce either claim, and neither does R2.
What the client does not control is the object key. I mint that server-side, and the way I mint it is what makes the rest of this possible.
The key knows which database it belongs to
EventAssets shards its data across 50 D1 databases, with the shard number encoded in the first two hex characters of every UUID. Object keys get the same treatment:
/**
* Object key for a file that belongs to an event: a UUID on the event's
* shard (so queue consumers can find the right database from the key
* alone) plus the extension for its content type.
*/
export function createEventScopedKey(
params: { event_id: string },
contentType: ContentType
): string {
const shardId = getShardFromUUID(params.event_id);
return `${generateShardedUUID(shardId)}.${CONTENT_TYPE_EXTENSIONS[contentType]}`;
}/**
* Object key for a file that belongs to an event: a UUID on the event's
* shard (so queue consumers can find the right database from the key
* alone) plus the extension for its content type.
*/
export function createEventScopedKey(
params: { event_id: string },
contentType: ContentType
): string {
const shardId = getShardFromUUID(params.event_id);
return `${generateShardedUUID(shardId)}.${CONTENT_TYPE_EXTENSIONS[contentType]}`;
}That comment is the load-bearing part. An R2 event notification tells you a bucket, an action, a key, a size and an eTag. It does not tell you which of my fifty databases has the row for it, and there is no place to attach that hint.
Because the key is a sharded UUID, the consumer recovers the database with parseInt(key.substring(0, 2), 16) and nothing else. No lookup table, no registry, no round trip. The object key routes itself.
Before handing out any URL, the API writes a pending row:
await prisma.media.createMany({
data: files.map(file => ({
key: file.key,
eventId: params.event_id,
albumId: params.album_id,
fileSize: file.fileSize,
contentType: file.contentType,
approved,
status: FileStatus.pending
}))
});await prisma.media.createMany({
data: files.map(file => ({
key: file.key,
eventId: params.event_id,
albumId: params.album_id,
fileSize: file.fileSize,
contentType: file.contentType,
approved,
status: FileStatus.pending
}))
});So at the moment the client starts writing, the system's belief is: I expect an object at this key, for this event, roughly this big, roughly this type, and I do not believe any of it yet.
R2 tells me when the bytes land
R2 event notifications push a message onto a Cloudflare Queue whenever an object is written. The wire format is hand-typed as a discriminated union, because it is Cloudflare's shape and not mine to redesign:
/**
* The R2 event notification payload, as Cloudflare delivers it to a queue
* consumer. Field names and casing mirror that wire format and are not ours
* to change.
*/
export type R2EventNotification =
| R2PutObjectEvent
| R2CopyObjectEvent
| R2CompleteMultipartUploadEvent
| R2DeleteObjectEvent
| R2LifecycleDeletionEvent;/**
* The R2 event notification payload, as Cloudflare delivers it to a queue
* consumer. Field names and casing mirror that wire format and are not ours
* to change.
*/
export type R2EventNotification =
| R2PutObjectEvent
| R2CopyObjectEvent
| R2CompleteMultipartUploadEvent
| R2DeleteObjectEvent
| R2LifecycleDeletionEvent;Only three of those five describe a written object, and only those three carry a size. TypeScript can extract exactly that subset:
/** Notification variants that describe a written object (with `size`). */
type R2UploadEvent = Extract<R2EventNotification, { object: { size: number } }>;/** Notification variants that describe a written object (with `size`). */
type R2UploadEvent = Extract<R2EventNotification, { object: { size: number } }>;A small thing, but it means the handler that reads event.object.size cannot be called with a DeleteObject event. The narrowing is structural rather than a runtime check I have to remember to write.
The ladder
Here is the whole validation sequence. Every rung either passes the object down or destroys it.
And the code, which is close to a transcription of that diagram:
private async handleUpload(event: R2UploadEvent): Promise<void> {
const { key, size: fileSize } = event.object;
const prisma = createPrismaClient(this.env, getShardFromUUID(key));
const record = await this.findRecord(prisma, key);
if (!record) {
this.log(`No pending record found, deleting: ${key}`);
await this.deleteFile(key);
return;
}
if (fileSize > this.maxFileSize) {
this.log(`File too large (${fileSize}), deleting: ${key}`);
await this.deleteFile(key);
await this.deleteRecord(prisma, key);
return;
}
const contentType = await this.detectContentType(key);
if (!contentType) {
this.log(`Invalid content type, deleting: ${key}`);
await this.deleteFile(key);
await this.deleteRecord(prisma, key);
return;
}
await this.completeRecord(prisma, key, {
fileSize,
contentType,
status: FileStatus.completed
});
await this.afterCompleted(prisma, record, key);
}private async handleUpload(event: R2UploadEvent): Promise<void> {
const { key, size: fileSize } = event.object;
const prisma = createPrismaClient(this.env, getShardFromUUID(key));
const record = await this.findRecord(prisma, key);
if (!record) {
this.log(`No pending record found, deleting: ${key}`);
await this.deleteFile(key);
return;
}
if (fileSize > this.maxFileSize) {
this.log(`File too large (${fileSize}), deleting: ${key}`);
await this.deleteFile(key);
await this.deleteRecord(prisma, key);
return;
}
const contentType = await this.detectContentType(key);
if (!contentType) {
this.log(`Invalid content type, deleting: ${key}`);
await this.deleteFile(key);
await this.deleteRecord(prisma, key);
return;
}
await this.completeRecord(prisma, key, {
fileSize,
contentType,
status: FileStatus.completed
});
await this.afterCompleted(prisma, record, key);
}Three details in there are worth pulling out.
No pending row means the object dies, and the row does not. There is nothing to delete — that branch is the one that catches an object written to a key I never issued, or to a key whose row was already cleaned up. Either way the object has no owner and no business existing.
Size is taken from the notification, not from the client. event.object.size is R2 reporting what it actually stored. The number the client declared at presign time is never compared against it and never written to the completed row; it is overwritten by the real one.
The type is sniffed from the stored bytes, and that is the next section.
Reading 4100 bytes to find out what a file really is
/** Bytes read from the object head to sniff its real content type. */
const FILE_TYPE_SAMPLE_BYTES = 4100;
private async detectContentType(key: string): Promise<ContentType | null> {
const object = await this.bucket.get(key, {
range: { offset: 0, length: FILE_TYPE_SAMPLE_BYTES }
});
if (!object) return null;
const fileType = await fileTypeFromBuffer(
new Uint8Array(await object.arrayBuffer())
);
const mimeType = fileType?.mime;
if (!mimeType || !this.allowedMimeTypes.includes(mimeType)) return null;
return MIME_TO_CONTENT_TYPE[mimeType as keyof typeof MIME_TO_CONTENT_TYPE];
}/** Bytes read from the object head to sniff its real content type. */
const FILE_TYPE_SAMPLE_BYTES = 4100;
private async detectContentType(key: string): Promise<ContentType | null> {
const object = await this.bucket.get(key, {
range: { offset: 0, length: FILE_TYPE_SAMPLE_BYTES }
});
if (!object) return null;
const fileType = await fileTypeFromBuffer(
new Uint8Array(await object.arrayBuffer())
);
const mimeType = fileType?.mime;
if (!mimeType || !this.allowedMimeTypes.includes(mimeType)) return null;
return MIME_TO_CONTENT_TYPE[mimeType as keyof typeof MIME_TO_CONTENT_TYPE];
}A range GET, not a full read. This is the difference between a validation step that works and one that falls over on the first large video: a 1 GB object sniffed by downloading 1 GB would blow the Worker's memory and its CPU budget, and would do it on exactly the uploads that matter most.
4100 bytes is file-type's recommended minimum sample — enough for every signature it knows, including the ones that are not at offset zero. ISO base media files (MP4, MOV) put their ftyp box near the start but not necessarily at the start, and some formats need a few kilobytes of slack. 4100 covers all of them and costs nothing.
The allowlist is per-subclass, and it is the same list the presign schema validates against:
export const MEDIA_MIME_TYPES = [
'image/jpeg',
'image/png',
'image/gif',
'image/webp',
'video/mp4',
'video/quicktime',
'video/x-msvideo',
'video/webm'
] as [string, ...string[]];export const MEDIA_MIME_TYPES = [
'image/jpeg',
'image/png',
'image/gif',
'image/webp',
'video/mp4',
'video/quicktime',
'video/x-msvideo',
'video/webm'
] as [string, ...string[]];The check is allowlist, not blocklist, and it runs against magic bytes rather than the extension or the declared header. An HTML file renamed .jpg and uploaded with Content-Type: image/jpeg produces no recognised signature, returns null, and is deleted along with its row. So does a PDF, a ZIP, and a polyglot file that is a valid GIF and something else — file-type reports the first signature it matches, and if that is image/gif the file is stored as a GIF and will be served with a GIF content type, which is the outcome I want.
Note what the sniff is not: it is not a virus scan, and it does not parse the file. A real JPEG containing something nasty in an EXIF field is a real JPEG and passes. What this closes is content-type confusion — an attacker getting a file with dangerous semantics stored under a type my application will trust and serve.
One base class, three buckets
Covers, media and QR images all go through the same ladder, differing only in their bucket, their limits, and which table holds their rows. That is an abstract class with three abstract accessors:
export abstract class R2UploadQueueHandler {
protected abstract readonly bucket: R2Bucket;
protected abstract readonly maxFileSize: number;
protected abstract readonly allowedMimeTypes: readonly string[];
protected abstract findRecord(
prisma: PrismaClient,
key: string
): Promise<{ eventId: string } | null>;
protected abstract deleteRecord(prisma: PrismaClient, key: string): Promise<unknown>;
protected abstract completeRecord(
prisma: PrismaClient,
key: string,
data: UploadRecordCompletion
): Promise<unknown>;
/** Hook for follow-up work once a record has been marked completed. */
protected async afterCompleted(
_prisma: PrismaClient,
_record: { eventId: string },
_key: string
): Promise<void> {}
}export abstract class R2UploadQueueHandler {
protected abstract readonly bucket: R2Bucket;
protected abstract readonly maxFileSize: number;
protected abstract readonly allowedMimeTypes: readonly string[];
protected abstract findRecord(
prisma: PrismaClient,
key: string
): Promise<{ eventId: string } | null>;
protected abstract deleteRecord(prisma: PrismaClient, key: string): Promise<unknown>;
protected abstract completeRecord(
prisma: PrismaClient,
key: string,
data: UploadRecordCompletion
): Promise<unknown>;
/** Hook for follow-up work once a record has been marked completed. */
protected async afterCompleted(
_prisma: PrismaClient,
_record: { eventId: string },
_key: string
): Promise<void> {}
}A concrete handler is then almost all declaration. The QR image one is fifty lines, and the interesting half is the hook:
export class QrImageQueueHandler extends R2UploadQueueHandler {
protected get bucket() { return this.env.EVENTASSETS_QR_IMAGES; }
protected readonly maxFileSize = QR_IMAGE_FILE_SIZE;
protected readonly allowedMimeTypes = QR_IMAGE_MIME_TYPES;
// ... findRecord / deleteRecord / completeRecord ...
/** An event keeps a single QR image: drop any previous ones. */
protected async afterCompleted(
prisma: PrismaClient,
record: { eventId: string },
key: string
): Promise<void> {
const otherRecords = await prisma.qrImage.findMany({
where: { eventId: record.eventId, key: { not: key } },
select: { key: true }
});
if (otherRecords.length === 0) return;
const otherKeys = otherRecords.map(other => other.key);
await prisma.qrImage.deleteMany({ where: { key: { in: otherKeys } } });
await Promise.all(otherKeys.map(otherKey => this.deleteFile(otherKey)));
}
}export class QrImageQueueHandler extends R2UploadQueueHandler {
protected get bucket() { return this.env.EVENTASSETS_QR_IMAGES; }
protected readonly maxFileSize = QR_IMAGE_FILE_SIZE;
protected readonly allowedMimeTypes = QR_IMAGE_MIME_TYPES;
// ... findRecord / deleteRecord / completeRecord ...
/** An event keeps a single QR image: drop any previous ones. */
protected async afterCompleted(
prisma: PrismaClient,
record: { eventId: string },
key: string
): Promise<void> {
const otherRecords = await prisma.qrImage.findMany({
where: { eventId: record.eventId, key: { not: key } },
select: { key: true }
});
if (otherRecords.length === 0) return;
const otherKeys = otherRecords.map(other => other.key);
await prisma.qrImage.deleteMany({ where: { key: { in: otherKeys } } });
await Promise.all(otherKeys.map(otherKey => this.deleteFile(otherKey)));
}
}"An event keeps exactly one QR image" is enforced after the new one is known to be valid, not before it is uploaded. Replace-then-delete rather than delete-then-replace: if the new upload turns out to be junk, it dies on the ladder and the old image is still there. There is no window in which an event has no QR code.
Failure handling, and the one thing I let slide
The outer handle is deliberately blunt:
async handle(message: Message<R2EventNotification>): Promise<void> {
try {
const event = message.body;
switch (event.action) {
case 'PutObject':
case 'CopyObject':
case 'CompleteMultipartUpload':
await this.handleUpload(event);
break;
default:
break;
}
message.ack();
} catch (error) {
Sentry.captureException(error);
message.retry();
}
}async handle(message: Message<R2EventNotification>): Promise<void> {
try {
const event = message.body;
switch (event.action) {
case 'PutObject':
case 'CopyObject':
case 'CompleteMultipartUpload':
await this.handleUpload(event);
break;
default:
break;
}
message.ack();
} catch (error) {
Sentry.captureException(error);
message.retry();
}
}Explicit ack() and retry() per message rather than letting the batch throw, so one poisoned notification cannot take nine healthy ones down with it. Queues are configured with max_retries: 3 and a dedicated dead-letter queue, so a message that genuinely cannot be processed ends up somewhere I can look at it instead of disappearing.
Deletion, though, is allowed to fail quietly:
protected async deleteFile(key: string): Promise<void> {
try {
await this.bucket.delete(key);
} catch (error) {
// An orphaned object is harmless for correctness but worth knowing about.
Sentry.captureException(error);
this.log(`Failed to delete file ${key}`, error);
}
}protected async deleteFile(key: string): Promise<void> {
try {
await this.bucket.delete(key);
} catch (error) {
// An orphaned object is harmless for correctness but worth knowing about.
Sentry.captureException(error);
this.log(`Failed to delete file ${key}`, error);
}
}That comment is a judgement call and I want to be explicit about it. If the delete fails, an object with no database row sits in R2 costing me storage. It is unreachable — nothing can serve it, because serving goes through rows, not through bucket listings. It is waste, not exposure. Throwing here would retry the whole message, which would re-run findRecord, now find nothing, and try to delete again — so the retry converges on the same state anyway. Logging and moving on is the honest version of that.
What this design gets wrong
Validation is asynchronous, so there is a window. Between the PUT completing and the queue consumer running, an oversized or fake-typed object exists in the bucket with a pending row pointing at it. Nothing serves pending rows, so it is not reachable — but "not reachable" is a property of my query filters, and a bug in one of those filters turns a latency window into an exposure window. Synchronous validation would close it, and synchronous validation is not available when the bytes never touch my Worker.
Retries are indiscriminate. A malformed message body and a transient D1 error take the same path: capture, retry, three times, then the DLQ. The malformed one can never succeed and burns three attempts proving it. Splitting "this will never work" from "try again" is a small change I have not made.
There is no quarantine. A file that fails the sniff is deleted immediately. That is the right default, but it means an upload that failed because file-type did not recognise a legitimate format is gone, and the user just sees a photo that vanished. Moving rejects to a quarantine prefix with a lifecycle rule would cost almost nothing and would have saved me two debugging sessions.
None of that changes the core, which I would build the same way again: let the client write directly, mint a key that routes itself, believe nothing until R2 tells you what actually landed, and check the bytes rather than the claim.
Next in this series: the one job in EventAssets that a Worker genuinely cannot do — building a 2 GB ZIP — and what it took to run it in a Cloudflare Container without ever holding the archive in memory.
EventAssets is a serverless event media platform built on Cloudflare Workers, D1, R2 and Queues. This post is part of a series on the architecture behind it.