Skip to main content
Toggle menu
Cracking
Walnuts
Interviews
Interview Roadmap
Start here
Pick your level, get your full prep plan
Case Studies
Full systems, worked end to end
HLD Playbooks
A full problem, step by step
Design Drills
One decision at a time, drilled
LLD Problems
Object-oriented design rounds
DSA Templates
Patterns & templates in 5 languages
Cheat Sheets
Quick-ref for interviews
Behavioral
Fresher to Principal
Company Playbooks
MAANG and AI labs, by level
Staff+ Interviews
Staff/principal rounds
Forward Deployed Engineer
The FDE role and its decomposition round
AI
AI Engineering
LLM fundamentals, internals, inference, RAG
Math for AI
The math behind LLMs, with small numbers
Security
Security
AppSec to cloud to architecture, attacker-first
AI Security
LLM, RAG, agent, and MCP security
Engineering
API Design
REST, HTTP semantics, OpenAPI & versioning
Concurrency
Java, Python & Go
Performance Eng
Latency, percentiles, profiling & labs
Networking & Protocols
TCP, HTTP, TLS & more
Linux Internals
Kernel & system internals
Distributed Systems
Consensus & replication
Infrastructure
Production infra & ops
Key Technologies
Core tech deep dives
Database Internals
How every engine works inside
Cryptography
How HTTPS, TLS, passkeys, and PQC work
DevOps Internals
How Kubernetes, Terraform, and the stack work
Leadership
Arch Decisions
ADRs & trade-offs
Platform Eng
Internal dev platforms
Incident Patterns
Postmortems & SRE
Metrics & FinOps
DORA, cost & KPIs
Compliance
Security & governance
Leadership
Tech lead & influence
Org Design
Team topologies & orgs
More
Posts
System design deep-dive topics
Visualizers & Calcs
Interactive system design tools
Whitepapers
Academic & industry papers
Books
Beyond-the-code reading picks
Search...
Ctrl+K
Go Premium
Log in
C
r
a
c
k
i
n
g
W
a
l
n
u
t
s
Insights
Deep analysis and thought pieces on AI, blockchain, and emerging technology trends.
How a 284B Model Can Run on a Consumer GPU
Mixture of Experts plus a bandwidth-aware runtime, not a compression trick
SYSTEM RAM
complete routed-expert pool Ā· source of truth
complete expert pool stays in system RAM
PCIe
measured bandwidth
GPU VRAM
reused expert weights Ā· eviction discards cached copies
plus non-expert weights, KV state and runtime buffers
CPU EXECUTION
executes selected experts whose weights stay in system RAM
q* = m Ć B_PCIe / B_Host
Router selects experts
for this token and layer
Cache hits run on GPU
weights already in VRAM
Selected experts absent from VRAM
copy for GPU execution or run on CPU
How a 284B Model Can Run on a Consumer GPU
System Design
AI
Aug 25, 2026 Ā· 51 min
How an LLM actually runs in Kubernetes
A request passes an API gateway that resolves tenant, quota and cache salt, then an llm-d Router and Endpoint Picker that selects a replica using prefix locality and load, then a vLLM server that schedules, batches, prefills and decodes. One model replica spans four GPUs joined by NCCL collectives. Kubernetes runs the GPU pods; the inference runtime manages token work, KV blocks and tensor parallelism.
How an LLM Actually Runs in Kubernetes
Kubernetes runs GPU pods. vLLM turns the model into an inference server.
THE REQUEST PATH
Client
one prompt
API Gateway
tenant Ā· quota Ā· cache_salt
reject before the GPU
llm-d Router / EPP
prefix locality Ā· load
filter Ā· score Ā· pick endpoint
vLLM
schedule Ā· batch Ā· prefill
Ā· decode
4 GPUs, one replica
NCCL collectives
THE TWO CACHES PEOPLE CONFUSE
KV cache
attention state for tokens already processed
Prefix cache
lets later requests reuse matching KV blocks
prefix caching reuses KV blocks; neither is your conversation database
WHO OWNS WHAT
Kubernetes
GPUs, pods, network, rollout
llm-d
which replica gets this request
vLLM
KV cache, batching, tensor parallel
NCCL / CUDA
NCCL: rank communication Ā· CUDA: kernels
Kubernetes never runs an AllReduce, and it never sees a token
WHY ONE REPLICA NEEDS FOUR GPUs
70B in BF16 ā 140 GB of weights
before any KV cache at all
One 80 GB GPU is not enough
so the model is sharded
Tensor parallel = 4
one logical replica, four GPUs, not four copies
Scaling out is the other axis: replicas multiply throughput, tensor parallelism makes one model fit.
TP 4 Ć 2 replicas = 8 GPUs
A GPU failure takes the whole four-GPU group with it, so that group is the unit of failure and of scaling.
How an LLM Actually Runs in Kubernetes
AI
Aug 23, 2026 Ā· 48 min
LLM inference serving platform architecture
An LLM request passes through a policy gateway, pinned tokenizer, atomic admission, and KV-aware router into a four-GPU tensor-parallel replica group. A continuous-batch scheduler coordinates prefill and repeated decode while paged KV memory supports both phases and an optional privacy-scoped prefix cache can accelerate eligible prefills. Tokens pass through a bounded streamer. A separate control plane manages immutable releases, group placement, canaries, and continuous group autoscaling.
LLM Inference Serving Platform
A long-lived request: prefill once ā decode repeatedly ā stream tokens
REGIONAL DATA PLANE
Policy Gateway
auth Ā· tier Ā· timeout
Tokenizer
pinned chat template
Atomic Admission
tokens Ā· KV Ā· spend
deadline feasibility
KV-Aware Router
prefix Ā· queue Ā· load
one complete group
TP=4 Replica Group
one schedulable worker
one failure + scaling unit
Streamer
backpressure Ā· cancel
token events ā client
AGGREGATED vLLM WORKER Ā· DENSE 70B GQA Ā· BF16 Ā· FOUR 80GB GPUs
CONTINUOUS-BATCH SCHEDULER
chooses prefill or decode work each iteration
Prefill Once
1,500 prompt tokens Ā· chunk long prompts
Decode Repeatedly
one next token / step Ā· 247 actual tokens
Paged KV Memory
active state supports both phases
Optional Prefix Cache
privacy-scoped reuse feeds prefill
Token Stream
delta Ā· terminal usage
release KV + reconcile spend
TENSOR-PARALLEL GROUP
load Ā· run Ā· drain Ā· fail together
GPU Rank 0 Ā· shard
GPU Rank 1 Ā· shard
GPU Rank 2 Ā· shard
GPU Rank 3 Ā· shard
fast link
for example, NVLink
DEPLOYMENT CONTROL PLANE
Immutable Release
weights + tokenizer
Gang Placement
whole TP=4 group
Canary + Quality
health then samples
Continuous Autoscaling
tokens Ā· KV Ā· zone reserve
ZONE RESILIENCE
Zone A
Zone B
Zone C
survivors hold protected peak load
10K avg Ā· 25K peak RPS
70B Ā· TP=4
TTFT < 2s p95
TPOT < 80ms p95
cost target: $0.005/request
Premium
System Design: LLM Inference Serving
System Design
AI
Aug 22, 2026 Ā· 62 min
Premium
System Design: LLM Safety Pipeline
System Design
AI
Aug 22, 2026 Ā· 36 min
Premium
System Design: LLM Evaluation Platform
System Design
AI
Aug 21, 2026 Ā· 39 min
Distributed training of a 70B model
A worked design for training a 70-billion-parameter model on 1024 GPUs using tensor parallelism four, context parallelism two, pipeline parallelism four, and sharded data parallelism thirty-two. An illustrative BF16 AdamW policy uses sixteen bytes per parameter before activations. FSDP2 gathers each parameter group on first use, keeps it resident across sixteen accumulated microbatches, then reduce-scatters gradients and updates local optimizer shards. At forty-five percent active-run MFU and ninety-six percent training availability, the compute estimate becomes about thirty-three calendar days. Checkpoint policy is planned from an illustrative unplanned-interruption rate rather than treating every interruption as hardware failure.
Training a 70B Model on 1024 GPUs
TP4 Ć CP2 per node Ā· PP4 across four nodes Ā· sharded-DP32 Ā· about 33 calendar days
ILLUSTRATIVE BF16 ADAMW POLICY Ā· 16 BYTES PER PARAMETER BEFORE ACTIVATIONS
weights
2B bf16 Ā· 140 GB
grads
2B bf16 Ā· 140 GB
master
4B fp32 Ā· 280 GB
momentum
4B fp32 Ā· 280 GB
variance
4B fp32 Ā· 280 GB
16 B/param
= 1.12 TB
the model
optimizer + master copy Ā· 12 of the 16 bytes
ONE FSDP2 PARAMETER GROUP Ā· ACCUMULATION-AWARE RESHARD POLICY
Gather on first use
microbatch 1
Forward + backward
keep parameters resident
Reuse Ć 15
no repeated gathers
Final backward +
reduce-scatter
keep local gradient shard
Sharded
optimizer
after accumulation
retain stage parameters across 16 microbatches Ā· spend HBM to avoid repeated inter-node gathers
actual wire bytes depend on group size, dtype, model partitioning, collective algorithm, and reshard policy
SCHEDULE = COMPUTE / (PEAK Ć ACTIVE MFU Ć AVAILABILITY)
compute needed
6 Ć 70e9 params Ć 3e12 tokens
= 1.26e24 FLOPs
peak available
1024 Ć 989 TFLOP/s dense bf16
= 1.01e18 FLOP/s
at 45% MFU
0.45 Ć 1.01e18
= 4.56e17 FLOP/s
1.26e24 / 4.56e17
= 32 active days
at 96% training availability: 32 / 0.96 ā 33.3 calendar days
PLANNING REFERENCE Ā· ONE UNPLANNED INTERRUPTION PER ~50,000 GPU-HOURS
day 0
day 32
ā illustrative interruption marks Ā· not every interruption is a hardware fault
Checkpoint every 30 min
shards + bounded storage writers
Lose T/2 on average
15 min Ć 15 failures ā 4 h
watchdog detects ā fence node ā use warm spare ā restore the last complete checkpoint
restore the next global sample IDs, topology and RNG state before training resumes
Premium
System Design: Distributed 70B Training
System Design
AI
Aug 20, 2026 Ā· 42 min
ML model serving platform architecture
A predictive inference request passes through a gateway, contract validation and admission, stable experiment and immutable version selection, optional input preparation, and model-aware routing to a ready CPU or GPU replica. The runtime may batch compatible requests and returns one bounded prediction. A deployment pipeline manages immutable releases and canaries, while quality evaluation, autoscaling, and observability run as parallel feedback loops.
ML Model Serving Platform
Predictive inference: one request in ā one bounded prediction out
LIVE DATA PLANE
API Gateway
auth Ā· limits Ā· deadline
Validate + Gateway Admit
contract Ā· tenant Ā· deadline
coarse rate + concurrency
Experiment + Version
stable A/B assignment
pin immutable release
Model Router
ready version
queue Ā· health Ā· load
Ready Model Replica
exact tensor ā bounded queue ā runtime
CPU or GPU
dynamic batch
Prediction
postprocess Ā· result
one bounded response
OPTIONAL INPUT PREPARATION
entity IDs ā online features ā versioned transform
media or text ā validated preprocessing
prepared tensors bypass feature/media preparation
THE SAME PLATFORM, DIFFERENT BOUNDED MODELS
ResNet Ā· image classifier
image tensor ā model ā
ācatā
XGBoost Ā· fraud model
transaction features ā
0.92
Ranking model
user + candidates ā
A āŗ C āŗ B
CONTROL PLANE + FEEDBACK
DEPLOYMENT FLOW
Model Registry
artifact Ā· schema Ā· lineage
Validate + Warm
probe Ā· place Ā· ready
Canary + Publish
observe Ā· promote Ā· rollback
PARALLEL FEEDBACK LOOPS
Quality + A/B
outcomes Ā· drift Ā· slices
Autoscaling
queue Ā· latency Ā· reserve
Observability
p99 Ā· errors Ā· cost
50K inference RPS
ranker p99 < 50 ms
100+ models
CPU Ā· GPU Ā· multiple frameworks
Premium
System Design: ML Model Serving Platform
System Design
AI
Aug 19, 2026 Ā· 71 min
DEEPSEEK HARNESS + CORDIS
Everything is a plugin.
The paper that explains how live components can leave cleanly.
PROVIDERS
CONSUMER
MODEL PLUGIN
TOOLS PLUGIN
SESSION PLUGIN
CORDIS CONTEXT
effects + dependencies + lifecycle
AGENT LOOP
provide
require
TEMPORAL
Undo what the plugin changed
register tool → remove tool
open resource → close resource
SPATIAL
React when a dependency changes
service appears → activate
service disappears → deactivate
A Programming Paradigm for Spatiotemporal Composability
DeepSeek Harness Says Everything Is a Plugin. Cordis Explains How That Can Work
AI
Software Architecture
Deep Dives
Aug 18, 2026 Ā· 25 min
How Can an Invisible Watermark Live Inside AI-Generated Text?
AI
Security
Deep Dives
Aug 11, 2026 Ā· 15 min
INSIDE THE HTTP TERMINATOR
138 documents in.
30,000 attacks out.
How the AI generated them, and why it was never allowed to say which ones worked.
WHAT THE MACHINE DID
138
documents
15,000
fragments
THE AI
reads each one
30,000
attack ideas
30,000
sites tested
~700
actually broke
WHO DECIDED WHAT
THE AI
invents the odd requests
allowed to be wrong
THE CODE
decides pass or fail
never argues, never guesses
THE PERSON
says which result matters
and why it matters
AI generates the possibilities. Code decides what survives. A person decides what any of it means.
HTTP Terminator Ā· PortSwigger Research, August 2026
How AI Helped Find New HTTP Attacks, and What It Actually Did
AI
Security
Deep Dives
Aug 7, 2026 Ā· 9 min
Proving You Are Over 18 Without Revealing Your Birth Date
Security
Deep Dives
Jul 30, 2026 Ā· 22 min
CRUSH
Controlled Replication Under Scalable Hashing
OBJECT
customer-123/
profile.jpg
PLACEMENT GROUP
PG 7.1d
1 of 128
CRUSH
cluster map + rule
+ device weights
OSD SET, IN ORDER
osd.1
rack-a / node-a
primary
osd.2
rack-a / node-b
replica
osd.5
rack-b / node-c
replica
no lookup table
topology aware
weighted by capacity
A Billion Objects, No Index: How Ceph Finds Anything
Deep Dives
Database
Jul 29, 2026 Ā· 32 min
AI Software Engineer
From autocomplete to autonomous app builder
L1: AUTOCOMPLETE
300ms
12
async
function
getUser
(id) {
13
const user =
await db.find(id)
14
if (!user) throw
Tab āµ
Context + FIM + Ranking
90M completions/day
L2: CODEBASE AGENT
45 sec
search_files("authenticate")
ā 8 files
read_file("src/auth/session.ts")
ā 142 lines
edit_file("src/middleware/auth.ts")
ā applied
run_command("npm test")
ā 48 pass
12 files edited, all tests pass
Think ā Act ā Observe ā Repeat
1.5M agent sessions/day
L3: AI ENGINEER
4 hours
DB Schema
Done
Auth Module
Done
Kanban Board
In Progress
Stripe Billing
Queued
Deploy to Vercel
Queued
Step 120 / 200 (60%)
Checkpoint cp-120
Memory: CLAUDE.md
Spec ā Build ā Test ā Deploy
50K build sessions/day
Model 50% | System 50%
Model 25% | System 75%
Model 10% | System 90%
L1: Autocomplete
L2: Agent
L3: Autonomous
Keystroke ā Context ā Model ā Verify ā Ship
Premium
System Design: AI Software Engineer (From Autocomplete to Autonomous App Builder)
System Design
AI
Mar 25, 2026 Ā· 81 min
RAG & LLM Platform
From Documents to Accurate, Cited Answers at Scale
Ingest
2M+ documents
Chunk & Embed
10M vectors
Retrieve
Hybrid + Re-rank
Generate
Routed + Cited
Evaluate
LLM-as-Judge
Continuous Improvement Loop
RAG and LLM Platform at Scale: Ingestion, Retrieval, Generation, and Evaluation for 10M Queries/Day
System Design
AI
Mar 22, 2026 Ā· 56 min
AI Agent Platform
Restaurant operations intelligence at scale
Detect
Investigate
Act
DATA IN
POS Systems
Toast Ā· Square Ā· Clover
Delivery
DoorDash Ā· Uber Eats Ā· Grubhub
Payments
Stripe Ā· Adyen
Inventory
MarketMan Ā· BlueCart
Agent
Agent
Agent
Agent
ACTIONS OUT
Alerts
SMS Ā· Email Ā· Dashboard
Reports
Root cause analysis
Auto-Fix
Disputes Ā· Reorders Ā· Pauses
Insights
Trends Ā· ROI Ā· Recommendations
KAFKA Ā· FLINK Ā· CLICKHOUSE Ā· TEMPORAL Ā· LLM AGENTS
10K investigations/day Ā· Multi-tenant Ā· Real-time anomaly detection
Premium
Building a Multi-Tenant AI Agent Platform for Restaurant Intelligence
System Design
AI
Mar 17, 2026 Ā· 99 min
x402: Can HTTP 402 Power Payments for AI Agents?
AI
Blockchain
Deep Dives
Feb 21, 2026 Ā· 18 min
Updated
No Humans Allowed: Inside Moltbook, the AI-Only Social Network. Build and Deploy Your Own Autonomous Agent
AI
Jan 31, 2026 Ā· 18 min
Passwordless Authentication's Silent Choice: Trust vs Privacy
Security
Deep Dives
Aug 10, 2025 Ā· 8 min
Premium
End-to-End Encryption in Chat Applications
Security
Deep Dives
Jun 19, 2025 Ā· 13 min
BGP: The Backbone Protocol That Powers Global DNS and Content Delivery
Networking
Deep Dives
May 1, 2025 Ā· 12 min
QUIC and WebTransport: Rebuilding the Internet for the Real World
Networking
Deep Dives
Apr 29, 2025 Ā· 8 min
How Distributed Databases Handle Conflicts: Vector Clocks, Syncing, and Conflict Resolution
Database
Deep Dives
Apr 26, 2025 Ā· 4 min
ClickHouse Capabilities: A Quick Overview
Database
Deep Dives
Apr 9, 2025 Ā· 5 min
How Apache Iceberg powers the Data Lake and Trino Makes It Explorable
Data Engineering
Deep Dives
Apr 8, 2025 Ā· 7 min
How Apache Pinot Achieves Ultra-Low Latency Analytics for User-Facing Applications
Data Engineering
Deep Dives
Apr 8, 2025 Ā· 5 min
CockroachDB vs Google Spanner: A Deep Dive Beyond the Basics
Deep Dives
Database
Apr 8, 2025 Ā· 2 min
Securing Access: The Power of RBAC, ABAC, and ReBAC
Security
Deep Dives
Apr 8, 2025 Ā· 6 min
How Google Spanner Achieves Global Consistency
Deep Dives
Database
Apr 8, 2025 Ā· 6 min