For a long time, adding an image to a post here meant uploading it somewhere else, copying a URL and hoping it stayed alive. So I built a proper media pipeline into the site. It turned out to be a nice little tour of Django, S3, Celery and ffmpeg, so here's how it works.

Step 1: Just write markdown

Posts are markdown, so media should be too. Every upload gets a short name, and the post points at it with a made-up media: scheme:

![My cat judging me](media:grumpy-cat)

![Build timelapse](media:timelapse){.autoplay .loop .muted width=480}

![Conference talk](youtube:dQw4w9WgXcQ)

The renderer, markdown-it-py, spots media: and swaps in the right <img>, <video> or <audio> tag. The {...} options come from the attrs plugin, but with a twist: nothing in the braces reaches the HTML directly. Every attribute is set aside, and the renderer picks out the few it knows about:

renderer = MarkdownIt("default", {"html": False}).use(
    attrs_plugin, after=("image",), allowed=()
)

allowed=() means "keep everything aside for me", and html: False means no raw HTML sneaks through either. Markdown is lovely, but it's not a place to type onerror=.

Step 2: Trust the bytes, not the name

A file called cute.jpg might really be a PNG, a video, or something that shouldn't be there at all. So uploads are identified by their contents: magic bytes for PDFs and ZIPs, Pillow for images, and ffprobe for everything else.

FILE_SIGNATURES = {
    b"%PDF-": ("application/pdf", ".pdf"),
    b"PK\x03\x04": ("application/zip", ".zip"),
}

The file name the browser sent never decides anything. Stored files get a random name inside the article's folder:

def article_media_path(instance, filename):
    extension = Path(filename).suffix.lower()
    return f"blog/{instance.article_id}/{uuid.uuid4().hex}{extension}"

Step 3: Make it browser-friendly (in the background)

Once the upload is committed, a Celery task takes over. on_commit matters here: without it, the worker can go looking for a row the database hasn't saved yet.

transaction.on_commit(lambda: process_media.delay(pk))

The task does whatever the file needs:

  • Photos get resized WebP copies at 640, 1280 and 1920 pixels wide, served with srcset. The original is never served. A nice side effect: re-encoding drops EXIF data, so no GPS coordinates of my kitchen end up online.
  • MP4s that browsers can already play just get remuxed with +faststart, so they start playing before they've fully downloaded. No quality lost.
  • Everything else (MOV from a phone, MKV, odd audio) is transcoded to H.264/AAC, and videos get a generated poster frame.
ffmpeg(
    "-i", source,
    "-c:v", "libx264", "-crf", "23", "-pix_fmt", "yuv420p",
    "-c:a", "aac", "-b:a", "160k",
    "-movflags", "+faststart",
    target,
)

The result goes into a status field (pending → processing → ready or failed, with an error message), so the admin always shows what happened.

There's one race I had to handle: what if I replace a file while the old one is still being processed? The task only saves its result if the file is still the one it started with:

updated = ArticleMedia.objects.filter(pk=media.pk, file=original_name).update(**fields)
if not updated:
    delete_stored_files(saved)  # we lost the race; clean up our leftovers

Step 4: Big files skip the server entirely

Small files go through a normal Django form. But a 3 GB video would tie up a web worker for ages and run into every proxy body limit on the way. So the admin's drag-and-drop uploader asks Django for a presigned PUT URL, and the browser sends the file straight to the bucket:

url = client.generate_presigned_url(
    "put_object", Params=params, ExpiresIn=60 * 60, HttpMethod="PUT"
)

Once the upload finishes, the browser tells Django, which registers the file and queues processing. Django checks that the key it's told about is one it could have generated for that article, so the browser can't register any random object in the bucket.

Step 5: Serving it (and caching forever)

Rendered posts are cached for a week, so media URLs must never expire. That rules out signed URLs, so media lives in its own public bucket with random names and year-long immutable cache headers.

The render cache key includes a fingerprint of every media item in the post, so editing alt text or re-processing a video naturally gives a fresh render. There's no "clear cache" button to forget to press.

In front of the bucket sits Cloudflare, which caches everything at the edge and blocks other sites from hotlinking. The bucket itself only answers requests that carry a secret only Cloudflare knows, so going around the CDN gets you nothing.

Step 6: Clean up after yourself

Replace or delete a media item and its old files (the original, the WebP copies, the poster) are deleted from the bucket. That happens after the transaction commits, so a rolled-back save doesn't delete files that are still in use.

Was it worth it?

Absolutely. Adding a video to a post is now drag, drop, and type media:. The rest happens on its own. And every time a 40 MB phone photo turns into a 90 KB WebP, I feel a little glow of satisfaction.

Now I can share this picture of my dogs My pups