"Design Dropbox for 100 million users: files sync across devices in seconds, editing one slide in a 2 GB deck doesn't re-upload 2 GB, and storing everyone's files doesn't bankrupt us."
"When I hit save, the whole file uploads again." Not in any competent sync system. A 2 GB video with 3 changed seconds re-uploads ~12 MB (3 changed 4 MB blocks), not 2 GB. The client diffs block hashes against the server's manifest and ships only the deltas. If your interview answer has the client uploading whole files, you've designed FTP with a nicer logo. Try the widget — edit a file and count the blocks that actually move.
| Assumption | Value |
|---|---|
| Users | 100,000,000 |
| Avg stored per user | 10 GB → 1 exabyte logical |
| Dedup ratio (shared memes, OS files, templates) | ~30% → 700 PB physical |
| Block size | 4 MB → 250 blocks per 1 GB file |
| Daily changed data | 1% of stored/day = 10 PB logical… but only changed blocks cross the wire |
| Edit 1 slide in a 2 GB deck | ~3 blocks change → 12 MB uploaded (vs 2,000 MB — 167× less) |
| Metadata | 700 PB ÷ 4 MB ≈ 175B block hashes × 64 B ≈ 11 TB — sharded metadata DB |
The math that matters: wire bytes per edit (changed blocks only) and dedup savings (the single biggest cost lever — 300 PB you don't buy disks for). Everything else is plumbing.
Smaller blocks → finer deltas (less re-upload per edit) but more metadata and more per-block overhead. Bigger blocks → fewer hashes to track but a 1-byte edit re-uploads the whole block. 4 MB is the industry sweet spot (Dropbox, S3 multipart, and rsync-family tools all land nearby). Bonus: 4 MB × a few parallel connections saturates most uplinks without head-of-line blocking.
The centerpiece: metadata and blocks take completely different paths.
flowchart TD
A[Desktop / mobile client
watches filesystem] --> B[Sync engine
chunk + hash + diff]
B --> C[Metadata service
block manifests]
B --> D[Block store
content-addressed: hash → bytes]
C --> E[(Metadata DB
sharded by user)]
D --> F[(Object storage
S3-style)]
G[Notify service
long-poll / push] --> A
C --> G
H[Second device] --> C
H --> D
Think of it like a library with a card catalog: the metadata service is the catalog — "War and Peace, edition 7, consists of cards #881…#1402." The block store is the stacks — card #881 exists exactly once no matter how many catalog entries reference it. Editing a chapter replaces a few cards; the catalog entry gets a new edition number atomically.
sequenceDiagram
participant C as Client
participant M as Metadata service
participant B as Block store
C->>C: file changed → split 4MB → SHA-256 each
C->>M: "I have manifest [h1,h2,h3',h4]"
M-->>C: "I need h3' only"
C->>B: PUT blocks/h3' (12 MB, not 2 GB)
B-->>C: stored (dedup: h1,h2,h4 already here)
C->>M: commit manifest v8 (atomic)
M-->>C: 200 — all devices notified
The manifest commit is the atomic moment: until v8 commits, every device still sees v7. Readers never observe a half-uploaded file. This is the same trick databases use with write-ahead logs.
sequenceDiagram
participant P as Phone
participant M as Metadata service
participant B as Block store
P->>M: long-poll: anything new?
M-->>P: file F now v8: [h1,h2,h3',h4]
P->>P: I have h1,h2,h4 cached — need h3'
P->>B: GET blocks/h3'
B-->>P: 4 MB
P->>P: reassemble + verify hashes
The phone downloads one block, not the file — because it diffed its local cache against the manifest first. Same protocol both directions; the client is just a smart cache.
POST /v1/files/{path}/manifest
{ baseVersion: 7, blocks: ["h1","h2","h3'","h4"] }
→ { needed: ["h3'"], version: 8 } // only missing blocks requested
PUT /v1/blocks/{sha256} // content-addressed, idempotent
GET /v1/files/{path} → { version, blocks:[...], size }
GET /v1/changes?cursor=… // long-poll for other devices
POST /v1/share/{path} → { link }
SHA-256(content) — the hash is the address. Identical content from any user maps to one stored object: dedup falls out of the addressing scheme.(user_id, path) → {version, [block hashes], mtime}. Version increments atomically; conflicts (two devices edited offline) fork versions and surface a "conflicted copy."| Decision | Option A | Option B | Pick |
|---|---|---|---|
| Block size | Fixed 4 MB | Content-defined (rolling hash) | A for v1 — simple, predictable; B later for files with insertions (a byte inserted at the top shifts every fixed block — content-defined boundaries survive that) |
| Dedup scope | Per user | Global | Global — the savings are in cross-user duplicates (OS installers, viral files); hash addressing makes it free |
| Encryption | Server-side | Client-side (convergent) | Server-side at rest for v1 — true client-side encryption kills global dedup (same file → different ciphertext); name the trade-off explicitly |
| Conflict handling | Last-writer-wins | Conflicted copies | Conflicted copies — silent data loss is the one unforgivable sync sin |
Client in Rust or Go (filesystem watchers + hashing are CPU-hot; you want the speed). Metadata service: Go + Postgres sharded by user id, manifests as JSONB. Block store: S3 with the SHA-256 as the object key — content addressing over S3 gives dedup, 11 nines, and zero block-server code. Notify: long-polling (simple, firewall-friendly) with push as an optimization. Delta encoding upgrade later: rsync-style rolling checksums inside changed blocks for sub-block deltas. Sharing: signed URLs over the block store. Honestly, the v1 is "content-addressed S3 + a manifest database + a good client" — and that's most of Dropbox's architecture too.
Dropbox famously added LAN sync: if your coworker's laptop already has block h3', fetch it over the local network instead of the internet. Same manifest protocol, different peer. It's a beautiful example of the block abstraction paying off — the source of a block is irrelevant to the protocol.
File A is 32 MB = 8 blocks of 4 MB. Hit edit — a few bytes change, the file re-splits, and only changed blocks upload. Then add an identical file B and watch dedup do its magic.
changed uploading synced deduped — 0 bytes sent