How do you design a scalable file storage system like Google Drive or Dropbox?
- Technology:
- System Design
- Experience:
- Senior
Quick answer
Store file contents as chunks in object storage and file metadata in a database; clients upload chunks directly with pre-signed URLs, a sync service tracks versions and changes, and a CDN serves downloads.
Problem
Design a service where users upload, download, sync and share files across devices.
Functional requirements
- Upload and download files of any size, including resuming interrupted uploads.
- Organise files in folders and keep a version history.
- Sync changes across a user鈥檚 devices.
- Share files and folders with other users with permissions.
Non-functional requirements
- Scalability: petabytes of data and many concurrent uploads.
- Availability: files can be read even during partial outages.
- Performance: fast downloads worldwide, and only changed parts re-uploaded.
- Reliability: files must never be lost or corrupted.
High-level architecture
- Clients split files into chunks (for example 4 MB) and hash each one. The client asks the metadata service which chunks are new, then uploads only those directly to object storage using pre-signed URLs.
- Once all chunks are uploaded, the client commits a new file version listing its chunk hashes. The metadata service records it in a transaction.
- A sync service publishes change events; other devices of the same user receive them over a long-lived connection or by polling and download the changed chunks.
- Downloads are served through a CDN in front of object storage, with access checked by issuing short-lived signed URLs.
Main components
- API / metadata service
- Folders, files, versions, sharing and permission checks.
- Object storage
- Durable, replicated storage of immutable chunks addressed by hash.
- Database
- Relational store for metadata, sharded by user or workspace.
- Cache
- Hot metadata such as folder listings and permission lookups.
- Message queue
- Change events for sync, thumbnail generation and virus scanning.
- CDN
- Caches and serves popular file downloads close to users.
Database design
- files(id, owner_id, parent_folder_id, name, current_version_id, deleted_at)
- file_versions(id, file_id, size, created_at, created_by)
- version_chunks(version_id, position, chunk_hash)
- chunks(hash PRIMARY KEY, size, storage_key, ref_count)
- shares(file_id, grantee_id, permission)
Scaling
- Content-addressed chunks deduplicate identical data across versions and users.
- Uploading directly to object storage keeps large transfers off the application servers.
- Shard metadata by owner or workspace; most queries stay within one shard.
- Move old versions and rarely used files to cheaper storage tiers.
Trade-offs
- Chunking enables resumable uploads and delta sync but adds metadata and client complexity.
- Deduplication saves storage but requires reference counting and careful deletion.
- Conflict handling: when two devices edit offline, keeping both copies is simpler and safer than automatic merging.
Discussion
A file storage service must store large files reliably, sync changes across devices and let users share files, while keeping storage and bandwidth costs under control.
Separating metadata from content is the central decision: metadata (names, folders, versions, permissions) is small and needs transactions; content is large, immutable once written and fits object storage.
Key points
- Separate metadata (database) from content (object storage)
- Chunk files and deduplicate by content hash
- Upload directly to storage with pre-signed URLs
- CDN for downloads, change events for sync