Media & Streaming

Designing a Global Video Streaming Platform (YouTube / Netflix) (2 Billion DAU • 1M+ Video Uploads/Day • 500 Gbps Ingress)

Staff / Principal

Architect a globally distributed video streaming system handling petabyte-scale storage, asynchronous chunked transcoding (HLS / DASH), CDN edge caching, and adaptive bitrate streaming.

Production Scale: 2 Billion DAU • 1M+ Video Uploads/Day • 500 Gbps Ingress

Functional Requirements

  • •Upload video files up to 10GB
  • •Adaptive Bitrate (ABR) video playback across resolutions (360p to 4K)
  • •Global search and view count tracking

Non-Functional Requirements

  • •High availability (99.99%)
  • •Ultra-low buffering latency (<200ms TTFB)
  • •Zero data loss for uploaded media

Capacity & Scale Estimation

Daily Active Users (DAU)2 Billion Users
Daily Video Uploads1,000,000 videos / day (~500 TB/day)
Streaming Bandwidth20 Petabytes / hour (Peak CDN egress)
Storage Required (5 Yrs)~900 Petabytes of replicated blob storage

Core Architectural Components

1Upload Service & Chunking

Receives multi-part file chunks, uploads to temporary S3 staging bucket, and triggers async transcoding event to Kafka.

2Transcoding Pipeline (DAG)

Distributed worker pool (FFmpeg) splits video into 6-second segments, encodes into multiple resolutions (AV1/H.265), and outputs HLS/DASH manifest (.m3u8).

3Global Edge CDN (Cloudflare/CloudFront)

Caches video segments at regional edge points of presence (PoPs) to serve 95%+ of video bytes without hitting origin servers.

4Metadata & Search Cluster

PostgreSQL for transactional metadata + Elasticsearch cluster for sub-second video search indexing.

Architectural FAQs & Interview Deep Dives

What are the core functional and non-functional requirements for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Functional requirements define user-facing capabilities, while non-functional requirements mandate high availability (99.99%), sub-100ms p99 latency, horizontal scalability, and data durability.

How do you calculate QPS, storage, and bandwidth capacity estimates for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Estimate daily active users (DAU), read/write ratio (e.g. 100:1), average payload size (e.g. 2KB), and calculate peak QPS (2–3x average) and 5-year storage projections.

How do you design the high-level API schema (REST / gRPC) for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Expose idempotent endpoints with explicit authentication headers, rate-limiting metadata, pagination cursors, and structured JSON / Protobuf error schemas.

What database paradigm (Relational SQL vs NoSQL vs Graph) is optimal for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Relational SQL (PostgreSQL) is chosen for ACID transactions and structured queries, NoSQL (Cassandra/DynamoDB) for high-write key-values, and Vector/Graph DBs for specialized relationships.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) implement database sharding and partitioning?

By using consistent hashing on user/entity IDs with virtual nodes, distributing partition keys uniformly across shards while preventing hot partitions.

How do you prevent cache stampedes (thundering herds) in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Use probabilistic early expiration (XFetch algorithm), distributed mutex locks, or pre-warming background worker threads before keys expire.

What caching strategy (Cache-Aside, Write-Through, Write-Behind) is best for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Cache-Aside is standard for read-heavy workloads, while Write-Through guarantees consistency at the cost of write latency.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) guarantee idempotency for financial transactions and mutations?

Clients send unique idempotency keys in request headers, which are stored in Redis/PostgreSQL with unique constraints to reject duplicate execution.

How do you handle distributed transactions and consistency across microservices in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Implement the Saga Pattern (choreography or orchestration) with compensating transactions to ensure eventual consistency without two-phase commit locks.

What message broker (Kafka vs RabbitMQ vs NATS) should be selected for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Apache Kafka is optimal for high-throughput event replay and partitioning, RabbitMQ for complex AMQP routing, and NATS JetStream for ultra-low latency messaging.

How do you handle message deduplication and out-of-order delivery in event streams for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Use monotonic event sequence numbers, store processed event IDs in transactional tables, and ensure consumer handlers are strictly idempotent.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) implement rate limiting at the API gateway layer?

Employ the Token Bucket or Sliding Window Log algorithm implemented with Redis Lua scripts to enforce client IP and account tier limits.

How do you design leader election and distributed consensus for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Use Raft or Paxos consensus engines (etcd, Consul, ZooKeeper) to maintain deterministic leader election and distributed lock state.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) scale WebSocket and real-time bidirectional connections?

Stateless WebSocket gateway servers maintain open TCP connections, backed by Redis Pub/Sub or Kafka to route messages across cluster nodes.

How do you handle multi-region active-active database replication in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Use CRDTs (Conflict-Free Replicated Data Types) or Last-Write-Wins timestamps with vector clocks to resolve cross-region concurrent write conflicts.

What is the Disaster Recovery (DR) strategy and RPO/RTO targets for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Define Recovery Point Objective (RPO < 1 min) and Recovery Time Objective (RTO < 5 min) supported by automated cross-region DNS failover (Route 53).

How do you prevent single points of failure (SPOF) across Designing a Global Video Streaming Platform (YouTube / Netflix) architecture?

Ensure every component (load balancers, web servers, databases, queues) runs with at least N+1 redundancy across independent cloud Availability Zones.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) isolate noisy neighbors in multi-tenant environments?

Implement separate worker pools, per-tenant rate limits, dedicated database schemas, and fair-share queue scheduling.

How do you handle large file uploads and streaming media in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Generate presigned S3/GCS upload URLs allowing clients to upload directly to object storage with multipart chunking and background event processing.

How do you design full-text search and filtering capabilities for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Stream database CDC (Change Data Capture) events via Debezium and Kafka to Elasticsearch, OpenSearch, or ClickHouse for sub-second analytical search.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) manage connection pooling under high concurrency?

Deploy database proxy layers (PgBouncer, ProxySQL) in transaction pooling mode to multiplex thousands of client connections over a compact server pool.

How do you execute zero-downtime database schema migrations for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Follow the Expand-Contract pattern: 1) Add new column/table, 2) Dual-write to both old and new schemas, 3) Backfill historical data, 4) Read from new schema, 5) Drop old column.

How do you implement distributed tracing across microservices in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Propagate W3C Trace Context headers (`traceparent`) across HTTP/gRPC boundaries and send spans to Jaeger or Grafana Tempo via OpenTelemetry collectors.

What metrics are most critical on the primary dashboard for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Monitor the Four Golden Signals: Latency (p50, p95, p99), Traffic (QPS), Errors (5xx rate), and Saturation (CPU, RAM, connection pool usage).

How do you design graceful degradation and circuit breakers in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Implement Resilience4j / Polly circuit breakers that open when upstream failure rates exceed 50%, serving cached fallback data rather than blocking.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) secure sensitive data at rest and in transit?

Enforce TLS 1.3 encryption in transit and AES-256-GCM / KMS envelope encryption at rest, with automated key rotation.

How do you implement Role-Based Access Control (RBAC) and ABAC in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Evaluate fine-grained permissions using Open Policy Agent (OPA) or Zanzibar-style relation graphs (Ory Keto) for low-latency authorization checks.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) prevent Distributed Denial of Service (DDoS) attacks?

Deploy edge CDNs with anycast routing (Cloudflare, AWS CloudFront), implement SYN flood protection, and enforce IP reputation rate limits.

How do you handle asynchronous background jobs and retry dead-letter queues in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Enqueue tasks to background worker pools (Celery, Sidekiq, Temporal) with exponential backoff retries and route permanently failing tasks to Dead Letter Queues (DLQ).

How does Designing a Global Video Streaming Platform (YouTube / Netflix) optimize cloud infrastructure costs (FinOps)?

Use auto-scaling spot instances for stateless workloads, purchase reserved instances for baseline database capacity, and implement S3 object storage lifecycle policies.

How do you manage DNS routing and traffic distribution for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Use latency-based or geolocation DNS routing with health-checked weighted target endpoints.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) handle distributed locking across multiple instances?

Utilize Redis Redlock algorithm or database advisory locks with explicit TTL lease renewal.

What serialization format is best for inter-service communication in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Protobuf over gRPC for internal low-latency microservices and JSON over HTTPS for external public APIs.

How do you structure database read replicas and replica lag mitigation for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Route read-only queries to asynchronous replicas while directing time-sensitive writes to the primary database.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) prevent memory exhaustion under unconstrained pagination?

Enforce keyset / cursor-based pagination with strict maximum limit caps (e.g. max 100 items per request).

What is the best strategy for handling hot keys in the caching tier of Designing a Global Video Streaming Platform (YouTube / Netflix)?

Replicate hot keys across multiple cache shards using random suffix salts or local in-memory L1 cache buffers.

How do you implement change data capture (CDC) for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Read database write-ahead logs (Postgres WAL / MySQL binlog) using Debezium to stream clean entity change events.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) handle graceful server restarts during rolling deployments?

Intercept SIGTERM signals, pause inbound health checks, drain active connection pools, and exit within 30 seconds.

What strategy is used for database connection multiplexing in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Deploy transaction-level connection poolers like PgBouncer to support 10,000+ client connections without database backend thrashing.

How do you enforce security headers and CSRF protection in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Set Content-Security-Policy (CSP), Strict-Transport-Security (HSTS), and use SameSite=Lax HttpOnly cookies.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) perform point-in-time recovery (PITR) for databases?

Archive continuous WAL segments to immutable cloud storage alongside nightly base backups.

What is the recommended logging format for Designing a Global Video Streaming Platform (YouTube / Netflix) in Kubernetes?

Emit structured JSON logs containing timestamp, level, trace_id, span_id, and service name to standard output.

How do you isolate blast radius during catastrophic service failures in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Implement bulkhead isolation patterns separating mission-critical billing workers from non-essential notification workers.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) validate semantic schema evolution in API contracts?

Use OpenAPI schemas and automated backward-compatibility linters (Spectral, Buf) in CI/CD pull request checks.

What strategy ensures zero data loss during message consumer crashes in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Disable auto-commit and acknowledge message offsets only after successful business logic execution.

How do you design multi-AZ failover for database clusters in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Configure synchronous replication to a standby instance in an adjacent Availability Zone with automated failover.

How does Designing a Global Video Streaming Platform (YouTube / Netflix) handle clock skew across distributed servers?

Synchronize all nodes using NTP / AWS Time Sync Service and avoid relying on physical timestamps for distributed ordering.

What approach prevents cascading failures when dependent third-party APIs slow down in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Wrap external HTTP calls in aggressive timeouts (max 2000ms) with circuit breakers and fallback responses.

How do you test system resilience against unpredictable infrastructure outages in Designing a Global Video Streaming Platform (YouTube / Netflix)?

Conduct automated Chaos Engineering experiments (Chaos Mesh, Gremlin) injecting simulated pod terminations and network latency.

What is the most fundamental architecture principle for Designing a Global Video Streaming Platform (YouTube / Netflix)?

Keep services stateless, design every operation to be idempotent, and decouple storage and compute independently.