System Design: Object Storage (Erasure Coding, Flat Namespace, and Exabyte Scale)
Goal: Build a distributed object storage platform storing 100+ trillion objects across exabytes of raw data.
Support PutObject, GetObject, DeleteObject, ListObjectsV2, multipart uploads up to 5TB, versioning, storage class transitions, and presigned URLs.
Target 11 nines of durability (99.999999999%) using Reed-Solomon erasure coding (10+4) combined with fast repair, multi-AZ isolation, and failure independence.
Strong read-after-write consistency. Flat namespace.
CDN-accelerated reads. Clear control plane / data plane separation.
1. The Core Problems
Problem: Durability at 11 nines. A typical data center runs around 100K drives. At a 2% annual failure rate, roughly 5.5 drives fail daily.
Failures are constant, not rare. Storage costs about $0.02/GB/month (~$20M per EB).
Triple replication stores 3 EB per EB of logical data, costing ~$60M/month. Reed-Solomon (10+4) stores 1.4 EB, costing ~$28M/month.
That saves ~$32M/month and tolerates up to 4 shard failures instead of 2.
Problem: Flat namespace at trillion scale. No directories. The key "photos/2024/vacation/beach.jpg" is opaque.
LIST with prefix=photos/2024/ and delimiter=/ simulates directories via efficient range scans over sorted, partitioned keys. This avoids metadata hotspots and distributed locking entirely.
Problem: Strong consistency. PUT then GET must always return the latest version. This requires Raft consensus at every metadata write.
Problem: Partial updates on immutable blobs. There is no PATCH, no append. Editing 10MB in a 1GB file means re-uploading the entire object.
The workaround is a manifest-based chunking layer built on top of the storage system. This is an application concern, not a storage concern.
Scale: 100T+ objects, exabytes of data, 350K req/sec peak, 5TB max object size.
2. Requirements
Functional Requirements
| ID | Requirement | Priority |
|---|---|---|
| FR-1 | PutObject: upload objects up to 5TB | P0 |
| FR-2 | GetObject: retrieve objects with byte-range reads | P0 |
| FR-3 | DeleteObject: soft delete (versioning) or hard delete (version-id) | P0 |
| FR-4 | ListObjectsV2: prefix + delimiter + pagination over trillions of objects | P0 |
| FR-5 | Multipart upload: create, upload-part (up to 10,000 parts), complete, abort | P0 |
| FR-6 | Presigned URLs: time-limited signed URLs for direct client-to-storage transfers | P0 |
| FR-7 | Bucket versioning: every PUT creates a new version, recoverable deletes | P1 |
| FR-8 | Storage class transitions: Standard, IA, Glacier Instant/Flexible, Deep Archive, Intelligent-Tiering | P1 |
| FR-9 | Lifecycle policies: automatic transition and expiration | P1 |
| FR-10 | CopyObject: server-side copy without re-upload | P1 |
| FR-11 | Object tagging for lifecycle rules and access policies | P1 |
| FR-12 | Bucket policies and IAM: resource-level and user-level access control | P0 |
| FR-13 | Server-side encryption: SSE-Default, SSE-KMS, SSE-C | P0 |
| FR-14 | Event notifications on object create/delete/transition | P1 |
| FR-15 | Cross-region replication: event-driven async replication with version ordering | P2 |
| FR-16 | CDN integration: edge caching via CDN-signed URLs, WAF/DDoS protection | P1 |
| FR-17 | Automatic rebalancing and repair on node addition/removal/failure | P0 |
Non-Functional Requirements
| Property | Target |
|---|---|
| Durability | 99.999999999% (11 nines) |
| Availability | 99.99% (52.6 min downtime/year) |
| PUT latency (p99, <1MB) | < 200ms |
| GET latency (p99, first byte, <1MB) | < 100ms |
| LIST latency (p99, 1000 results) | < 500ms |
| Total objects | 100T+ |
| Total logical storage | Exabytes |
| Maximum object size | 5TB |
| Peak request rate | 350K/sec |
| Storage overhead | < 1.5x (erasure coding) |
| Consistency | Strong read-after-write |
| CDN cache hit ratio | >= 60% for read-heavy workloads |
3. Design Principles
🔒 Premium section
4. Erasure Coding
🔒 Premium section
5. Object Placement: Placement Groups and CRUSH
🔒 Premium section
6. Metadata Architecture
🔒 Premium section
7. Storage Classes
🔒 Premium section
8. CDN and Edge Layer
🔒 Premium section
9. Control Plane vs Data Plane
🔒 Premium section
10. Architecture Overview
🔒 Premium section
11. Scale Estimation
🔒 Premium section
12. Data Model
🔒 Premium section
13. API Design
🔒 Premium section
14. End-to-End: From Upload to Recovery
🔒 Premium section
15. The PUT Path
🔒 Premium section
16. The GET Path
🔒 Premium section
17. Presigned URLs: Storage vs CDN
🔒 Premium section
18. Multipart Upload
🔒 Premium section
19. Small Object Packing: Extent Files
🔒 Premium section
20. Prefix Listing at Scale
🔒 Premium section
21. Garbage Collection
🔒 Premium section
22. Rebalancing and Repair
🔒 Premium section
23. Versioning
🔒 Premium section
24. Hot Objects and Caching
🔒 Premium section
25. Storage Layer Details
🔒 Premium section
26. User Experience Scenarios
🔒 Premium section
27. The Partial Update Layer
🔒 Premium section
28. What Breaks First at Scale
🔒 Premium section
29. Bottlenecks and Mitigations
🔒 Premium section
30. Failure Scenarios
🔒 Premium section
31. Deployment and Operations
🔒 Premium section
32. Beyond This Design: Real-World Evolution
🔒 Premium section
33. Explore the Technologies
🔒 Premium section