Skip to main content
Toggle menu
Cracking
Walnuts
Interviews
Interview Roadmap
Start here
Pick your level, get your full prep plan
Case Studies
Full systems, worked end to end
HLD Playbooks
A full problem, step by step
Design Drills
One decision at a time, drilled
Compare Technologies
Which cache, queue or database, and why
LLD Problems
Object-oriented design rounds
DSA Templates
Patterns & templates in 5 languages
Cheat Sheets
Quick-ref for interviews
Behavioral
Fresher to Principal
Company Playbooks
MAANG and AI labs, by level
Staff+ Interviews
Staff/principal rounds
Forward Deployed Engineer
The FDE role and its decomposition round
AI
AI Engineering
LLM fundamentals, internals, inference, RAG
Math for AI
The math behind LLMs, with small numbers
Security
Security
AppSec to cloud to architecture, attacker-first
AI Security
LLM, RAG, agent, and MCP security
Engineering
API Design
REST, HTTP semantics, OpenAPI & versioning
Concurrency
Java, Python & Go
Performance Eng
Latency, percentiles, profiling & labs
Networking & Protocols
TCP, HTTP, TLS & more
Linux Internals
Kernel & system internals
Distributed Systems
Consensus & replication
Infrastructure
Production infra & ops
Key Technologies
Core tech deep dives
Database Internals
How every engine works inside
Cryptography
How HTTPS, TLS, passkeys, and PQC work
DevOps Internals
How Kubernetes, Terraform, and the stack work
Blockchain Engineering
Bitcoin, Ethereum, consensus, L2s and ZK
Leadership
Arch Decisions
ADRs & trade-offs
Platform Eng
Internal dev platforms
Incident Patterns
Postmortems & SRE
Metrics & FinOps
DORA, cost & KPIs
Compliance
Security & governance
Leadership
Tech lead & influence
Org Design
Team topologies & orgs
More
Posts
System design deep-dive topics
Visualizers & Calcs
Interactive system design tools
Whitepapers
Academic & industry papers
Books
Beyond-the-code reading picks
Search...
Ctrl+K
Go Premium
Log in
C
r
a
c
k
i
n
g
W
a
l
n
u
t
s
Home
System Design
System Design
(27 posts)
Durable Object Runtime
One name. One owner. One epoch.
room-101
Resident on Node A · epoch 42
SQLite
sockets
timers
Node A
CAS won · lease alive
If-Match: abc ✓
owns
Node C
fenced · precondition failed
If-Match: abc ✗
coordination + durability
Shared object store — the root of authority
objects/room-101/owner.json
node
"Node-A"
epoch
42
ETag
"abc123"
nodes/Node-A/lease.json
addr
a.internal:8080
expires
T
one lease per node, not per object
objects/room-101/data/
e42/
e43/
the epoch in the key is the fence
local SQLite · ETag compare-and-swap · node leases · epoch fencing · route caches · lazy failover
Make coordination expensive at the ownership boundary, then make execution local
Premium
Design a Durable Object Runtime Like Cloudflare Durable Objects
System Design
Aug 26, 2026 · 53 min
LLM inference serving platform architecture
An LLM request passes through a policy gateway, pinned tokenizer, atomic admission, and KV-aware router into a four-GPU tensor-parallel replica group. A continuous-batch scheduler coordinates prefill and repeated decode while paged KV memory supports both phases and an optional privacy-scoped prefix cache can accelerate eligible prefills. Tokens pass through a bounded streamer. A separate control plane manages immutable releases, group placement, canaries, and continuous group autoscaling.
LLM Inference Serving Platform
A long-lived request: prefill once → decode repeatedly → stream tokens
REGIONAL DATA PLANE
Policy Gateway
auth · tier · timeout
Tokenizer
pinned chat template
Atomic Admission
tokens · KV · spend
deadline feasibility
KV-Aware Router
prefix · queue · load
one complete group
TP=4 Replica Group
one schedulable worker
one failure + scaling unit
Streamer
backpressure · cancel
token events → client
AGGREGATED vLLM WORKER · DENSE 70B GQA · BF16 · FOUR 80GB GPUs
CONTINUOUS-BATCH SCHEDULER
chooses prefill or decode work each iteration
Prefill Once
1,500 prompt tokens · chunk long prompts
Decode Repeatedly
one next token / step · 247 actual tokens
Paged KV Memory
active state supports both phases
Optional Prefix Cache
privacy-scoped reuse feeds prefill
Token Stream
delta · terminal usage
release KV + reconcile spend
TENSOR-PARALLEL GROUP
load · run · drain · fail together
GPU Rank 0 · shard
GPU Rank 1 · shard
GPU Rank 2 · shard
GPU Rank 3 · shard
fast link
for example, NVLink
DEPLOYMENT CONTROL PLANE
Immutable Release
weights + tokenizer
Gang Placement
whole TP=4 group
Canary + Quality
health then samples
Continuous Autoscaling
tokens · KV · zone reserve
ZONE RESILIENCE
Zone A
Zone B
Zone C
survivors hold protected peak load
10K avg · 25K peak RPS
70B · TP=4
TTFT < 2s p95
TPOT < 80ms p95
cost target: $0.005/request
Premium
System Design: LLM Inference Serving
System Design
AI
Aug 22, 2026 · 62 min
Premium
System Design: LLM Safety Pipeline
System Design
AI
Aug 22, 2026 · 36 min
Premium
System Design: LLM Evaluation Platform
System Design
AI
Aug 21, 2026 · 39 min
Distributed training of a 70B model
A worked design for training a 70-billion-parameter model on 1024 GPUs using tensor parallelism four, context parallelism two, pipeline parallelism four, and sharded data parallelism thirty-two. An illustrative BF16 AdamW policy uses sixteen bytes per parameter before activations. FSDP2 gathers each parameter group on first use, keeps it resident across sixteen accumulated microbatches, then reduce-scatters gradients and updates local optimizer shards. At forty-five percent active-run MFU and ninety-six percent training availability, the compute estimate becomes about thirty-three calendar days. Checkpoint policy is planned from an illustrative unplanned-interruption rate rather than treating every interruption as hardware failure.
Training a 70B Model on 1024 GPUs
TP4 × CP2 per node · PP4 across four nodes · sharded-DP32 · about 33 calendar days
ILLUSTRATIVE BF16 ADAMW POLICY · 16 BYTES PER PARAMETER BEFORE ACTIVATIONS
weights
2B bf16 · 140 GB
grads
2B bf16 · 140 GB
master
4B fp32 · 280 GB
momentum
4B fp32 · 280 GB
variance
4B fp32 · 280 GB
16 B/param
= 1.12 TB
the model
optimizer + master copy · 12 of the 16 bytes
ONE FSDP2 PARAMETER GROUP · ACCUMULATION-AWARE RESHARD POLICY
Gather on first use
microbatch 1
Forward + backward
keep parameters resident
Reuse × 15
no repeated gathers
Final backward +
reduce-scatter
keep local gradient shard
Sharded
optimizer
after accumulation
retain stage parameters across 16 microbatches · spend HBM to avoid repeated inter-node gathers
actual wire bytes depend on group size, dtype, model partitioning, collective algorithm, and reshard policy
SCHEDULE = COMPUTE / (PEAK × ACTIVE MFU × AVAILABILITY)
compute needed
6 × 70e9 params × 3e12 tokens
= 1.26e24 FLOPs
peak available
1024 × 989 TFLOP/s dense bf16
= 1.01e18 FLOP/s
at 45% MFU
0.45 × 1.01e18
= 4.56e17 FLOP/s
1.26e24 / 4.56e17
= 32 active days
at 96% training availability: 32 / 0.96 ≈ 33.3 calendar days
PLANNING REFERENCE · ONE UNPLANNED INTERRUPTION PER ~50,000 GPU-HOURS
day 0
day 32
│ illustrative interruption marks · not every interruption is a hardware fault
Checkpoint every 30 min
shards + bounded storage writers
Lose T/2 on average
15 min × 15 failures ≈ 4 h
watchdog detects → fence node → use warm spare → restore the last complete checkpoint
restore the next global sample IDs, topology and RNG state before training resumes
Premium
System Design: Distributed 70B Training
System Design
AI
Aug 20, 2026 · 42 min
ML model serving platform architecture
A predictive inference request passes through a gateway, contract validation and admission, stable experiment and immutable version selection, optional input preparation, and model-aware routing to a ready CPU or GPU replica. The runtime may batch compatible requests and returns one bounded prediction. A deployment pipeline manages immutable releases and canaries, while quality evaluation, autoscaling, and observability run as parallel feedback loops.
ML Model Serving Platform
Predictive inference: one request in → one bounded prediction out
LIVE DATA PLANE
API Gateway
auth · limits · deadline
Validate + Gateway Admit
contract · tenant · deadline
coarse rate + concurrency
Experiment + Version
stable A/B assignment
pin immutable release
Model Router
ready version
queue · health · load
Ready Model Replica
exact tensor → bounded queue → runtime
CPU or GPU
dynamic batch
Prediction
postprocess · result
one bounded response
OPTIONAL INPUT PREPARATION
entity IDs → online features → versioned transform
media or text → validated preprocessing
prepared tensors bypass feature/media preparation
THE SAME PLATFORM, DIFFERENT BOUNDED MODELS
ResNet · image classifier
image tensor → model →
“cat”
XGBoost · fraud model
transaction features →
0.92
Ranking model
user + candidates →
A › C › B
CONTROL PLANE + FEEDBACK
DEPLOYMENT FLOW
Model Registry
artifact · schema · lineage
Validate + Warm
probe · place · ready
Canary + Publish
observe · promote · rollback
PARALLEL FEEDBACK LOOPS
Quality + A/B
outcomes · drift · slices
Autoscaling
queue · latency · reserve
Observability
p99 · errors · cost
50K inference RPS
ranker p99 < 50 ms
100+ models
CPU · GPU · multiple frameworks
Premium
System Design: ML Model Serving Platform
System Design
AI
Aug 19, 2026 · 71 min
ChatGPT
End-to-end: APIs, inference, memory, RAG, tools & streaming
API Gateway
auth + rate limit
Input Safety
< 50ms BERT
Context Builder
assemble prompt
Model Router
7B / 34B / 70B
GPU Scheduler
admission + priority
Conversation DB
recent + summary
User Memory
cross-conversation
RAG / Vector DB
web + files + embeds
Inference Cluster · tensor-parallel 2x A100
Prefill
compute-bound
KV-Cache
PagedAttention
Decode
memory-bound
Sample
top-p + temp
Agent / Tool Loop
web search
code sandbox
external APIs
Speculative 1.5-2x
Continuous Batching
INT4 / FP16 / FP8
3-Tier Context
System ~500
Recent ~2000
Summary ~500
Safety: 3 Layers
1. Input Classifier
2. System Prompt
3. Output Classifier
jailbreak + injection defense
Alignment Pipeline
SFT · 50K demos
Reward Model · pref pairs
DPO / PPO · optimize policy
Shadow Eval · 5% traffic
Output Safety
scan each chunk
SSE Streaming
tokens + tools + cites
User
token-by-token
persist
100M DAU
35K peak QPS
2.3T tokens / day
~1000 GPUs
P50 TTFT 1.5s
Gateway → Safety → Context Builder → Route → Schedule → Tools → Inference → Stream
Premium
System Design: ChatGPT End-to-End (Inference, Memory, RAG, Streaming, RLHF)
System Design
Jun 4, 2026 · 49 min
Maps Platform
Navigation, Routing, and Map Rendering
A
B
23 min
12.4 km via Main St
Contraction CH
S2 Cell Index
Vector Tiles
5B routes/day
50M GPS/sec
Route <1ms (CH)
500M nodes
1.2B edges
~500 nodes/query
Kafka + Flink
HMM map match
30s traffic refresh
N
Premium
System Design: Maps Platform (Navigation, Routing, and Map Rendering)
System Design
Apr 27, 2026 · 29 min
GitHub
200M repos, pull requests & CI/CD
Open
feat: add distributed cache layer
#42 opened 2 hours ago by
@kim
main
←
feat/cache-layer
All checks passed
(4/4)
build
23s
test
1m 47s
lint
12s
security-scan
34s
alice
approved
1h ago
bob
approved
20m ago
Files changed
8 files
+342
-89
Squash and merge
cache/distributed.go
+ func NewCache(opts ...Option) *Cache {
+ return &Cache{shards: 16}
+ }
// Existing code
- func oldMethod() {
- // deprecated
@alice: Consider 32 shards
for better distribution at this scale.
@bob: Add cache-hit metrics
before this lands on main.
Repository
main
feat/cache
1.2K commits
42 contributors
Source
1B git ops/day
Push -> PR -> Review -> CI -> Merge
Premium
System Design: GitHub (200M Repos, Git Object Storage, Sparse Trigram Code Search, Per-Job VM CI)
System Design
Apr 23, 2026 · 57 min
Online Auction
50K bids/sec, effectively-once settlement
Live Bid Feed
Vintage Rolex Submariner · ends in 1m 47s
j***n · $12,450
seq #847 · just now
LEADING
m***k · $12,400
seq #846 · +2m extend
EXTENDED
a***i · $12,300
seq #845 · 12s ago
OUTBID
r***l · $12,250
stale expected_price
REJECTED
CAS · Lua atomic
Bid Processor
★
Effectively-Once
Settled · Charged
Anti-Sniping
Fencing Tokens
10M auctions
1M watchers
Bids: 50K/s
Watchers: 1M
Sold today: 87K
Premium
System Design: Online Auction (50K Bids/sec, Effectively-Once Settlement, Anti-Sniping)
System Design
Apr 17, 2026 · 48 min
Job Scheduler
10M jobs/day, effectively-once execution
Task Pipeline
DAG execution with priority queues
data-pipeline-etl
attempt 1 · 25s
RUNNING
user-sync-nightly
attempt 2 · 12s
RUNNING
report-monthly
QUEUED
analytics-daily
DONE 2m ago
*/5 * * * *
Scheduler
Effectively-Once
Delivered
DAG Execution
Priority Queues
10M jobs/day
Auto Retry
Queued: 12K
Running: 4K
Done: 8.2M
Premium
System Design: Job Scheduler (10M Jobs/day, DAG Dependencies, Effectively-Once Execution)
System Design
Apr 15, 2026 · 35 min
LeetCode
50M submissions/day, sandboxed code execution
solution.py
class
Solution
:
def
twoSum
(self, nums, target):
seen = {}
for
i, num
in
enumerate
(nums):
comp = target - num
if
comp
in
seen:
return
[seen[comp], i]
seen[num] = i
1
2
3
4
5
6
7
8
Output
Accepted
Runtime: 48ms (beats 95%)
Memory: 17.2MB (beats 88%)
Test Cases: 57/57 passed
Submit
Run
Easy
1. Two Sum
Array
Hash Map
Problem
Firecracker
microVM Isolation
50M subs/day
20+ languages
CRIU Warm Pool
Anti-Cheat
Submit -> Kafka -> Firecracker Sandbox -> Judge -> WebSocket
Premium
System Design: LeetCode (Code Sandbox, Container Isolation, Real-Time Contests)
System Design
Apr 12, 2026 · 58 min
URL Shortener
10B URLs, 100K redirects/sec
Shorten your URL
https://example.com/very/long/path/to/resource?param=value&q=2
Shorten
Your shortened URL:
sho.rt/Ab3xK9f
Copy
QR
142K clicks
42 countries
68% mobile
Expires: 30d
ID Generation
Counter: 5000042
Bijective shuffle
Ab3xK9f
Encode
302 Redirect
CDN -> Valkey -> Scylla
p99 < 5ms (cache hit)
10B URLs
100K redirects/s
Base62 / 7 chars
ClickHouse
Create -> Encode -> Store -> Redirect -> Track
Premium
System Design: URL Shortener (10B Short URLs, 100K Redirects/sec)
System Design
Apr 11, 2026 · 34 min
Ad Click Aggregator
10B clicks/day, Lambda architecture & fraud detection
Click Analytics Dashboard
Last 24 hours - All campaigns
Total Clicks
847.2M
CTR
2.34%
Revenue
$4.2M
Click Volume (hourly)
0
6
12
18
24
Fraud: 2.1%
Stream (Flink)
Batch (Iceberg)
Click Events
115K/sec
10B clicks/day
Ingest
Exactly-Once
Flink checkpoints
10B clicks/day
<1min freshness
Lambda Arch
Fraud Detection
Click -> Dedup -> Aggregate -> Bill -> Reconcile
Premium
System Design: Ad Click Aggregator (10B Clicks/day, Lambda Architecture, Fraud Detection)
System Design
Data Engineering
Apr 10, 2026 · 46 min
Ad Exchange
10M auctions/sec, RTB & sub-100ms
Ad Exchange Overview
Auctions/sec
10M
DSPs
20+
Impressions
1B/day
Revenue / yr
$10B
DSP: Google DV360
Bid: $2.45 CPM
Winner
DSP: The Trade Desk
Bid: $2.12 CPM
Runner-up
DSP: MediaMath
No bid (budget cap)
Skipped
Request -> Fan-out -> Auction -> Serve
Bid Request
SSP to DSPs
Header bidding
10M auctions/s
Sub-100ms
OpenRTB 2.6
First-Price
Premium
System Design: Ad Exchange (Real-Time Bidding, Sub-100ms Auctions, DSP/SSP, Impression Serving)
System Design
Apr 10, 2026 · 44 min
Flash Sale
10M users, limited coupons, one-per-user
FLASH SALE - Up to 70% OFF
Limited Stock! Ends in:
02:34:12
Wireless Earbuds
$99
$29
-70%
12 left!
Smart Watch
$149
$59
-60%
58 left
Bluetooth Speaker
SOLD OUT
USB-C Hub
$79
$39
-50%
142 left
SAVE20
Extra 20% off
Apply
One coupon per user - 500K available
Buy Now - $29
Queue
#2,431
of 10M users
~4 min wait
Waiting Room
10M users
500K coupons
Lua DECR
SET NX 1/user
Queue -> Browse -> Coupon -> Pay -> Confirm
System Design: E-Commerce Flash Sales (10M Users, Coupon System, One-Per-User Enforcement)
System Design
Apr 5, 2026 · 86 min
News Aggregator
100K sources, dedup & personalization
News Aggregator Overview
Sources
100K
Articles/day
5M
Unique
2M
Users
50M
reuters.com
Earthquake in Japan
Breaking
bbc.co.uk
Election results update
Trending
techcrunch.com
AI startup raises $50M
Tech
Crawl -> Dedup -> Rank -> Personalize
RSS Polling
Adaptive frequency
5min to 6hr
100K sources
MinHash LSH
Adaptive Poll
Decay Ranking
Premium
System Design: News Aggregator (100K Sources, Dedup, Personalized Ranking)
System Design
Apr 5, 2026 · 64 min
Dropbox
File sync, chunking & deduplication
report.xlsx · 100 MB · 8 chunks
C1
sha:a3f2..
C2
sha:b7e1..
C3'
sha:NEW!
C4
sha:d9c4..
C5
sha:e2a8..
C6
sha:f1b3..
C7
sha:0c7d..
C8
sha:4e9f..
upload 4 MB
7 chunks unchanged · skip · saved 96 MB
Edit
!
Chunk + Delta
✓
✓
✓
Sync
500M
users
100 PB
storage
<10s
cross-device sync
50%
dedup savings
Edit file. Upload only what changed. Sync everywhere.
presigned URLs · bytes skip the server · CDC or rsync
Premium
System Design: Dropbox (File Sync, Chunking, and Deduplication)
System Design
Apr 2, 2026 · 60 min
Object Storage
Store anything, retrieve anytime, lose nothing
11 Nines Durability
1.4x Overhead
Strong Consistency
100T+ Objects
Exabyte Scale
350K req/sec
PUT anything. GET anytime. Never lose a byte.
photos/2024/beach.jpg
Premium
System Design: Object Storage (Erasure Coding, Flat Namespace, and Exabyte Scale)
System Design
Mar 31, 2026 · 49 min
AI Software Engineer
From autocomplete to autonomous app builder
L1: AUTOCOMPLETE
300ms
12
async
function
getUser
(id) {
13
const user =
await db.find(id)
14
if (!user) throw
Tab ↵
Context + FIM + Ranking
90M completions/day
L2: CODEBASE AGENT
45 sec
search_files("authenticate")
✓ 8 files
read_file("src/auth/session.ts")
✓ 142 lines
edit_file("src/middleware/auth.ts")
✓ applied
run_command("npm test")
✓ 48 pass
12 files edited, all tests pass
Think → Act → Observe → Repeat
1.5M agent sessions/day
L3: AI ENGINEER
4 hours
DB Schema
Done
Auth Module
Done
Kanban Board
In Progress
Stripe Billing
Queued
Deploy to Vercel
Queued
Step 120 / 200 (60%)
Checkpoint cp-120
Memory: CLAUDE.md
Spec → Build → Test → Deploy
50K build sessions/day
Model 50% | System 50%
Model 25% | System 75%
Model 10% | System 90%
L1: Autocomplete
L2: Agent
L3: Autonomous
Keystroke → Context → Model → Verify → Ship
Premium
System Design: AI Software Engineer (From Autocomplete to Autonomous App Builder)
System Design
AI
Mar 25, 2026 · 81 min
RAG & LLM Platform
From Documents to Accurate, Cited Answers at Scale
Ingest
2M+ documents
Chunk & Embed
10M vectors
Retrieve
Hybrid + Re-rank
Generate
Routed + Cited
Evaluate
LLM-as-Judge
Continuous Improvement Loop
RAG and LLM Platform at Scale: Ingestion, Retrieval, Generation, and Evaluation for 10M Queries/Day
System Design
AI
Mar 22, 2026 · 56 min
AI Agent Platform
Restaurant operations intelligence at scale
Detect
Investigate
Act
DATA IN
POS Systems
Toast · Square · Clover
Delivery
DoorDash · Uber Eats · Grubhub
Payments
Stripe · Adyen
Inventory
MarketMan · BlueCart
Agent
Agent
Agent
Agent
ACTIONS OUT
Alerts
SMS · Email · Dashboard
Reports
Root cause analysis
Auto-Fix
Disputes · Reorders · Pauses
Insights
Trends · ROI · Recommendations
KAFKA · FLINK · CLICKHOUSE · TEMPORAL · LLM AGENTS
10K investigations/day · Multi-tenant · Real-time anomaly detection
Premium
Building a Multi-Tenant AI Agent Platform for Restaurant Intelligence
System Design
AI
Mar 17, 2026 · 99 min
Collaborative Editor
Real-time editing with CRDTs
B
I
U
Alice
Bob
Carol
A
Alice
B
Bob
C
Carol
CRDT (Yjs)
Offline-first
100 editors/doc
<100ms sync
originLeft
originRight
merge(A,B) = merge(B,A)
WebSocket
bidirectional
binary sync
System Design: Real-Time Collaborative Editor
System Design
Mar 14, 2026 · 72 min
Observability Platform
Metrics, Traces, Logs, and Profiles at Scale
SLO 99.9%
ML
eBPF + OTel → Pipeline → Kafka → Processing → Storage → Query
Traces
Profiles
Logs
Grafana
500M metrics/s
200M spans/s
50M log lines/s
50K profiles/s
p99 < 200ms
4 signal types
tail sampling
eBPF + OTel SDK
Premium
Observability Platform at Scale - Metrics, Traces, Logs, and Profiles
System Design
Mar 12, 2026 · 87 min
Notification System
100M/sec across every channel
Web
Mobile
Wearable
p99 < 500ms
99.99% up
100M connections
3 channels
Premium
System Design: Notification Platform, 100M Notifications/Second
System Design
Mar 6, 2026 · 66 min
Top-K Streaming
What's trending right now?
1
2
3
4
5
1.5M
980K
720K
410K
290K
LIVE
updated 12s ago
Premium
Building a Fail-Safe, Scalable Top-K Streaming System
System Design
Data Engineering
Feb 27, 2026 · 44 min
Tagging System
Multi-tenant, globally consistent
Spanner
US
Redis + Kafka + ES
EU
Redis + Kafka + ES
Asia
Redis + Kafka + ES
#marketing
#design
#api
#docs
#sales
#launch
#v2
#infra
CONSISTENT
across 3 regions
Premium
Designing a Multi-Tenant, Globally Scalable Tagging System
System Design
May 25, 2025 · 21 min