System Design: Dropbox (File Sync, Chunking, and Deduplication)
Goal: Build a cloud file sync platform serving 500M registered users, 50M DAU, syncing 10B files across 100PB of storage. Support chunked uploads, resumable transfers, delta sync (upload only changed bytes), cross-device sync via WebSocket, content deduplication across users, peer-to-peer LAN sync, and file versioning. Files available on other devices within 10 seconds for files under 10MB on broadband.
Mental model, four ideas that make everything else click:
- A file = ordered list of chunk hashes. Chunks are immutable, content-addressed blobs in object storage.
- Editing a file means producing a new list of hashes. Most stay the same. Only new chunks get uploaded.
- Dedup = identical content produces the same SHA-256 hash, stored once, referenced many times.
- Sync = push metadata via WebSocket, other devices download only missing chunks.
1. The Three Problems
File sync looks simple until the math gets real. Three problems shape every decision in this design.
Problem 1: Efficient transfer. The platform processes 1M file uploads per minute at peak, average file size 2MB.
A user opens a 100MB spreadsheet, changes one cell (maybe 1KB of actual data), and saves. Re-uploading the entire 100MB wastes bandwidth.
Multiply that by millions of users and the cost becomes unsustainable. The system must detect what actually changed and upload only those bytes.
Problem 2: Multi-device sync. 50M daily active users, each with 2-3 devices on average. That's 100-150M devices that need to stay in sync.
When a user edits a file on their laptop, their phone and work desktop need to know within seconds. And "know about it" means downloading exactly the changed portions, not the whole file.
Polling is out of the question, 150M devices hitting the server every 5 seconds asking "anything new?" is 30M requests/second of pure waste. The server must push changes.
Problem 3: Storage efficiency. 100PB of total storage. Without deduplication, every copy of the same file takes separate space.
A company onboards 500 employees who all receive the same 50MB employee handbook PDF. That's 500 copies, 25GB, for one file.
With content-addressable deduplication, the system stores one copy and 500 pointers.
Scale numbers:
- 500M registered users
- 50M DAU
- 10B files under management
- 100PB total storage
- 1M uploads/minute at peak (~16,667/sec)
- Average file size: 2MB
2. Requirements
Functional Requirements
| ID | Requirement | Priority |
|---|---|---|
| FR-01 | Upload files from desktop, mobile, and web clients | P0 |
| FR-02 | Download files to any connected device | P0 |
| FR-03 | Resume interrupted uploads and downloads from the point of failure | P0 |
| FR-04 | Multi-device sync: file changes propagate to all user devices automatically | P0 |
| FR-05 | Delta sync: upload only changed portions of a file, not the entire file | P0 |
| FR-06 | Content deduplication: identical content stored once across all users | P0 |
| FR-07 | File versioning: retain previous versions for rollback | P1 |
| FR-08 | Conflict resolution: handle concurrent edits from multiple devices | P0 |
| FR-09 | Selective sync: choose which folders sync to which devices | P1 |
| FR-10 | LAN sync: transfer files directly between devices on the same network | P1 |
| FR-11 | File and folder sharing with permissions (view, edit) | P1 |
| FR-12 | Full-text search by filename, file type, and content metadata | P2 |
| FR-13 | Bandwidth throttling: user-configurable upload/download limits | P2 |
| FR-14 | File preview: thumbnails and previews for images, documents, videos | P2 |
| FR-15 | Team workspaces: shared folders with role-based access control | P1 |
Non-Functional Requirements
| ID | Requirement | Target |
|---|---|---|
| NFR-01 | Sync latency (metadata propagation) | < 1 second |
| NFR-02 | File availability on other devices | < 10 seconds for files under 10MB on broadband |
| NFR-03 | Resumable upload gap tolerance | Up to 7 days between pause and resume |
| NFR-04 | Deduplication ratio | 50% storage savings across all users |
| NFR-05 | Data durability | 99.999999999% (11 nines, standard tier) |
| NFR-06 | Service availability | 99.99% (52 minutes downtime/year) |
| NFR-07 | Maximum file size | 50GB |
| NFR-08 | Concurrent sync connections | 150M (50M DAU * ~3 devices * 30% concurrency) |
| NFR-09 | Version retention | 180 days or 100 versions, whichever comes first |
3. Why Naive Approaches Fail
🔒 Premium section
4. The Two-Hash Foundation
🔒 Premium section
5. Content-Defined Chunking (CDC)
🔒 Premium section
6. rsync-Style Delta Sync
🔒 Premium section
7. CDC vs rsync: Choosing the Right Model
🔒 Premium section
8. Chunk Size Tuning
🔒 Premium section
9. Technology Selection
🔒 Premium section
10. Architecture Overview
🔒 Premium section
11. Scale Estimation
🔒 Premium section
12. Data Model
🔒 Premium section
13. API Design
🔒 Premium section
14. The Upload Pipeline
🔒 Premium section
15. Resumable Uploads
🔒 Premium section
16. Content Deduplication
🔒 Premium section
17. Integrity Verification
🔒 Premium section
18. The Sync Protocol
🔒 Premium section
19. Conflict Resolution
🔒 Premium section
20. Delta Sync in Practice
🔒 Premium section
21. LAN Sync
🔒 Premium section
22. File Versioning
🔒 Premium section
23. Download Path and Caching
🔒 Premium section
24. Bottlenecks and Mitigations
🔒 Premium section
25. Failure Scenarios
🔒 Premium section
26. Deployment and Operations
🔒 Premium section
27. Beyond This Design: Real-World Evolution
🔒 Premium section
28. Explore the Technologies
🔒 Premium section