← Back to systems

Project / SYS-01

Video CDN

Tiered storage architecture for video delivery at scale

The problem

Video is expensive to serve. Every request either gets pulled from somewhere close to the user or it has to travel back to origin, and if you're not careful, most of your traffic ends up doing the expensive thing. Storing everything on fast media isn't affordable at scale, but storing everything on cheap media is too slow for anything popular. The system needs to know the difference between content that's being hit right now and content that's just sitting there, and it needs to make that call constantly, not just once when the file is uploaded.

This gets harder the bigger the catalog gets. A small library can afford to keep most things fast. A large one can't, because the cost of fast storage scales with how much you keep on it, and most of a video catalog at any given moment is not being watched by anyone. The interesting problem isn't storing video, it's figuring out which small slice of the catalog deserves to be fast right now, and being wrong about that as rarely as possible.

The approach

Storage is split into two real tiers plus a routing layer between them, rather than one flat cache sitting in front of a single storage backend.

Edge cache — the fastest tier, closest to the user. Reserved for content getting hit hard right now. This is deliberately kept small relative to the catalog, since keeping it small is what makes it fast and cheap to run.

NVMe — a middle caching tier for content that's warm but not viral. Fast enough to serve without a real latency penalty, and it takes pressure off origin without needing to keep everything on edge. This tier does most of the actual work day to day, since most requests aren't for the single hottest video on the platform, they're for the few thousand videos that are moderately popular at any given time.

HDD — this is origin, not cache. It's the primary storage layer holding the full catalog, everything that's ever been uploaded, whether it gets watched once a year or never again. Cheap per terabyte, which matters a lot when the catalog is large and most of it is cold.

Between edge and NVMe sits a proxy layer that handles routing. It decides where a given request actually gets served from, and if a faster tier is under load or doesn't have the content, it fails over to the next tier down instead of the request just failing. This layer is what makes the tiering actually work in practice rather than just in theory, since without something making real-time routing decisions, you'd need every client to know the storage topology, which doesn't scale.

Content gets promoted toward the faster tiers based on access patterns rather than sitting wherever it was first stored. Something that suddenly spikes in popularity, a video that goes viral overnight for example, doesn't need to hit HDD on every single request while it's trending. It gets pulled up to NVMe and potentially edge as demand shows up, and it drifts back down once that demand fades. The goal is for the fast tiers to always be full of whatever's actually hot right now, not whatever was hot last week.

What it looks like in production

Here's a snapshot from a 24-hour window on this system:

  • 297.38k unique visitors
  • 173.06M total requests
  • 67.68% of requests served from cache
  • 230 TB of data served
  • 156 TB of that served from cache instead of origin
Analytics dashboard showing a 24-hour production traffic window
24-hour window, production traffic

That cache hit rate is the number that actually matters. Two out of every three requests never touch HDD at all, which is the entire point of the tiering and routing setup. Origin storage is there to hold everything, not to serve everything, and the fact that it barely needs to serve most requests directly is a sign the promotion logic is doing its job.

Conclusion

Tiering works when it reflects the way people actually watch video. Most of the catalog can stay durable and inexpensive, while the active slice moves closer to users as demand appears. The routing layer makes that movement useful by keeping requests available even when a faster tier is missing content or under pressure.