How to Design WhatsApp — System Design Interview Question & Answer
1. Interview Question
Design a messaging system like WhatsApp.
The interviewer expects you to discuss:
- 1-to-1 messaging
- Group messaging
- Online/offline users
- Message delivery and acknowledgements
- Read receipts
- Push notifications
- Message ordering
- Media/file sharing
- Multi-device support
- End-to-end encryption
- Scalability
- High availability
- Failure handling
A good interview answer should not start with databases or Kafka. Start by clarifying requirements and then progressively build the architecture.
2. Clarify Requirements
I would begin the interview by asking:
Functional requirements
- Can users send 1-to-1 messages?
- Do we support group chats?
- Do messages need delivery and read receipts?
- Should users see online/offline presence?
- Do we support images, videos, documents, and voice messages?
- Should users receive messages when offline?
- Should messages synchronize across multiple devices?
- Do we need message search?
- Do we need voice/video calls?
For a 45-minute interview, I’d explicitly scope out voice/video calling unless the interviewer asks for it.
Non-functional requirements
We want:
- Low latency — messages should arrive within milliseconds to a few seconds.
- High availability — messaging should continue even if some servers fail.
- Massive scalability — potentially hundreds of millions of concurrent connections.
- Durability — messages shouldn’t disappear because a server crashes.
- Ordering — messages in a conversation should appear in a sensible order.
- Security — messages should be encrypted end-to-end.
3. High-Level Architecture
A reasonable architecture looks like this:
┌─────────────────────┐
│ Clients │
│ iOS / Android / Web │
└──────────┬──────────┘
│
WebSocket / HTTPS
│
▼
┌─────────────────────┐
│ Load Balancer │
└──────────┬──────────┘
│
┌─────────────────┼─────────────────┐
│ │ │
▼ ▼ ▼
┌───────────┐ ┌────────────┐ ┌─────────────┐
│Connection │ │ Message │ │ Presence │
│ Service │ │ Service │ │ Service │
└─────┬─────┘ └─────┬──────┘ └──────┬──────┘
│ │ │
│ ▼ │
│ ┌─────────────┐ │
│ │ Message Bus │ │
│ │ Kafka │ │
│ └──────┬──────┘ │
│ │ │
│ ┌────────┴─────────┐ │
│ ▼ ▼ │
│ ┌─────────────┐ ┌─────────────┐ │
│ │ Message DB │ │ Notification│ │
│ │ │ │ Service │ │
│ └─────────────┘ └──────┬──────┘ │
│ │ │
│ ▼ │
│ APNs / FCM │
│ │
└──────────────┬─────────────────────┘
│
▼
┌──────────────┐
│ Redis / KV │
│ Presence │
│ Connections │
└──────────────┘
Media:
Client → Object Storage → CDN
The important design principle is:
Separate persistent messaging from real-time connection management.
4. Client Connection
WhatsApp needs persistent connections because users expect messages to arrive immediately.
The client establishes a:
Client
│
│ WebSocket
▼
Connection Gateway
The WebSocket remains open while the application is active.
For example:
Alice
│
│ WebSocket
▼
Connection Server #17
The connection server maintains:
userId → connectionId
However, this mapping cannot live only in local memory.
We need distributed connection state:
Redis
alice → server-17
bob → server-42
john → server-8
When Alice sends a message to Bob:
Alice
│
▼
Connection Server 17
│
▼
Message Service
│
▼
Find Bob's connection
│
▼
Connection Server 42
│
▼
Bob
5. Sending a Message
Suppose Alice sends:
“Hello Bob”
The client generates a message ID:
messageId = UUID
conversationId = C123
senderId = Alice
receiverId = Bob
timestamp = ...
The request might look like:
{
"messageId": "m123",
"conversationId": "c456",
"senderId": "alice",
"type": "TEXT",
"ciphertext": "encrypted-content"
}
Notice something important:
The server doesn’t necessarily see the plaintext.
For true end-to-end encryption, the client encrypts the message before sending it.
6. Message Processing
The Message Service receives the message.
Alice
│
▼
Message Service
│
├── Validate
├── Authenticate
├── Authorize
├── Assign ordering
├── Persist
│
▼
Kafka
│
├── Delivery
├── Notification
└── Analytics
The message should be persisted before acknowledging successful acceptance.
This prevents:
Client → Server → ACK
X
crash
where the client thinks the message was accepted but the message was actually lost.
7. Message Ordering
Ordering is one of the most important interview topics.
Imagine:
Alice sends:
M1: "Are you free?"
M2: "Let's meet at 7."
M3: "At the restaurant."
The recipient should normally see:
M1
M2
M3
not:
M2
M1
M3
Approach
Partition messages by conversation:
Kafka
Partition 0 → Conversation A
Partition 1 → Conversation B
Partition 2 → Conversation C
Use:
partitionKey = conversationId
This ensures messages from the same conversation are processed sequentially by the same partition.
However, be careful:
Kafka ordering is guaranteed within a partition, not globally.
That’s actually exactly what we want.
We need per-conversation ordering, not global ordering across WhatsApp.
8. Delivery States
A WhatsApp-style system typically has several states:
SENT
↓
DELIVERED
↓
READ
For example:
Alice Bob
M1 ────────────────────►
SENT
◄──────────────────
DELIVERED
◄──────────────────
READ
We can represent:
messageId = M123
senderStatus:
SENT
receiverStatus:
DELIVERED
readStatus:
READ
9. What Happens If Bob Is Offline?
This is a classic interview question.
Suppose:
Alice → Bob
but Bob has no active WebSocket connection.
The Message Service checks:
Presence Service
Bob → OFFLINE
The message remains persisted.
Then:
Notification Service
│
├── iOS → APNs
│
└── Android → FCM
Bob receives:
You have a new message.
When Bob opens WhatsApp:
Bob
│
▼
WebSocket
│
▼
Message Service
│
▼
Fetch undelivered messages
│
▼
M1
M2
M3
Then the client sends acknowledgements.
10. Message Delivery Is Usually At-Least-Once
This is another important interview point.
Suppose Bob receives:
M123
but the ACK is lost.
The server may resend:
M123
Therefore the client must tolerate duplicates.
Use:
messageId
as an idempotency key.
The client can maintain:
processedMessageIds
and ignore duplicates.
Therefore:
At-least-once delivery + idempotent processing is often more practical than trying to guarantee exactly-once delivery across a distributed system.
11. Database Design
A possible message model:
Message
---------------------------
message_id
conversation_id
sender_id
sequence_number
timestamp
message_type
encrypted_payload
status
For example:
conversation_id | sequence | sender | message
------------------------------------------------
C100 | 101 | Alice | Hello
C100 | 102 | Bob | Hi
C100 | 103 | Alice | How are you?
Why conversation_id?
Because almost every query is:
WHERE conversation_id = ?
ORDER BY sequence_number
rather than:
SQL
WHERE message_id = ?
12. Choosing the Database
For a massive messaging system, a distributed NoSQL database is attractive.
Potential choices:
- Cassandra
- ScyllaDB
- DynamoDB
- Bigtable
For example:
Partition Key:
conversation_id
Sort Key:
sequence_number
This gives efficient:
Get conversation history
operations.
A relational database can work at smaller scale, but eventually horizontal partitioning becomes critical.
13. Database Sharding
Suppose we have billions of messages.
We cannot put everything on one database.
Partition by:
hash(conversationId) % N
Example:
Conversation A → Shard 1
Conversation B → Shard 8
Conversation C → Shard 4
Conversation D → Shard 2
The key is:
Keep messages from the same conversation on the same logical partition whenever possible.
That makes ordering and conversation-history queries much easier.
14. Group Messaging
Now suppose we have:
Group G1
Alice
Bob
Charlie
David
Emma
Alice sends:
“Meeting at 7.”
There are two major approaches.
Fan-out on write
Immediately create delivery records:
M1 → Bob
M1 → Charlie
M1 → David
M1 → Emma
Advantages:
- Fast reads
- Easy unread counts
Disadvantages:
- Expensive for huge groups
Fan-out on read
Store:
Group G1
│
└── M1
Members retrieve messages when they open the group.
Advantages:
- Much cheaper for huge groups
Disadvantages:
- More work during reads
A production system can use a hybrid strategy.
For example:
Small group → fan-out on write
Large group → fan-out on read
15. Presence Service
Users expect:
Bob
Online
or:
Bob
Last seen 5 minutes ago
Presence is usually an ephemeral state.
A good architecture:
Client
│
│ heartbeat
▼
Presence Service
│
▼
Redis
Example:
userId → {
status: ONLINE,
lastSeen: 2026-08-09T20:00:00
}
Use TTL:
Bob → ONLINE
TTL = 30 seconds
If Bob stops sending heartbeats:
TTL expires
↓
Bob → OFFLINE
This avoids requiring expensive database writes for every presence update.
16. Media Messages
Images and videos should not go through the messaging service.
Bad architecture:
Client
│
▼
Message Server
│
▼
Image
This consumes huge amounts of application-server bandwidth.
Instead:
Client
│
│ Upload
▼
Object Storage
│
▼
CDN
Examples:
S3
GCS
Azure Blob Storage
The message contains metadata:
{
"type": "IMAGE",
"mediaId": "img123",
"url": "...",
"size": 2048000
}
For secure messaging, the media itself should also be encrypted appropriately.
17. CDN
Media downloads are highly cacheable.
Instead of:
Bob → WhatsApp Server → Storage
use:
Bob
│
▼
CDN
│
├── Cache HIT → image
│
└── Cache MISS
│
▼
Object Storage
This dramatically reduces origin traffic.
18. Multi-Device Support
A user might have:
Alice
├── iPhone
├── Mac
├── iPad
└── Web browser
The system therefore needs device-level identity.
Instead of:
userId → connection
use:
userId
│
├── device-A
├── device-B
├── device-C
└── device-D
Each device has:
deviceId
publicKey
connection
lastSeen
Messages may need to be delivered to multiple devices.
This also becomes important for end-to-end encryption.
19. End-to-End Encryption
For a WhatsApp-like system, security is fundamental.
Conceptually:
Alice
│
│ Encrypt
▼
Ciphertext
│
▼
WhatsApp Servers
│
│ Cannot decrypt
▼
Bob
│
│ Decrypt
▼
Plaintext
The server should not need access to plaintext messages.
A modern design can use protocols based on the Signal Protocol concepts:
Identity Keys
Prekeys
Session Keys
Ratchets
The important interview statement is:
The messaging infrastructure routes and stores encrypted payloads, while cryptographic keys remain under client control.
20. WhatsApp-Style Architecture
Putting everything together:
┌───────────────────┐
│ Clients │
│ iOS / Android/Web │
└─────────┬─────────┘
│
TLS / WebSocket
│
▼
┌───────────────────┐
│ Load Balancer │
└─────────┬─────────┘
│
┌──────────────┼──────────────┐
│ │ │
▼ ▼ ▼
┌────────────┐ ┌────────────┐ ┌────────────┐
│ Connection │ │ Message │ │ Presence │
│ Gateway │ │ Service │ │ Service │
└─────┬──────┘ └──────┬─────┘ └──────┬─────┘
│ │ │
│ ▼ ▼
│ ┌──────────────┐ ┌───────┐
│ │ Kafka │ │ Redis │
│ └──────┬───────┘ └───────┘
│ │
│ ┌───────┴────────┐
│ │ │
▼ ▼ ▼
┌─────────┐ ┌────────────┐ ┌─────────────┐
│ Redis │ │ Message DB │ │ Notification│
│ Routing │ │ Cassandra/ │ │ Service │
│ │ │ DynamoDB │ └──────┬──────┘
└─────────┘ └────────────┘ │
▼
┌───────────┐
│ APNs/FCM │
└───────────┘
Media Pipeline
───────────────
Client ──► Object Storage ──► CDN ──► Client
21. Scaling the System
Let’s assume:
500M daily active users
100M concurrent connections
50B messages/day
The architecture must scale horizontally.
Instead of:
1 Message Server
we have:
Message Server
Message Server
Message Server
...
Message Server
Stateless services can scale using:
Kubernetes
+
Load Balancer
+
Auto Scaling
Connection servers are slightly different because WebSocket connections are long-lived.
We therefore need:
Connection Gateway
+
Connection Registry
to know which server owns each connection.
22. Handling Hot Conversations
A major scalability problem occurs when a single group becomes extremely popular.
For example:
Group:
100,000 members
If one message arrives:
1 message
↓
100,000 deliveries
This creates a fan-out storm.
Solutions include:
- asynchronous delivery
- batching
- partitioned fan-out
- backpressure
- rate limiting
- fan-out-on-read for very large groups
- dedicated partitions for hot conversations
This is a great point to mention in an interview because it demonstrates that you’re thinking beyond the happy path.
23. Failure Handling
What happens if Kafka is temporarily unavailable?
The Message Service should avoid simply dropping messages.
Possible approach:
Client
│
▼
Message Service
│
├── Primary message store
│
└── Kafka
Use retries with:
exponential backoff
+
dead-letter queue
For duplicate requests:
messageId
provides idempotency.
For database failures:
Replication
+
Multi-AZ
+
Automatic failover
For regional failures:
Region A
│
├── replicated data
│
▼
Region B
24. Consistency vs Availability
This is a classic distributed-systems trade-off.
For messaging:
Strong consistency is useful for:
- message ordering within a conversation
- message IDs
- delivery state transitions
Eventual consistency is acceptable for:
- online/offline status
- last seen
- unread counters
- some analytics
For example, if Bob’s status changes:
ONLINE → OFFLINE
being delayed by 1–2 seconds is generally acceptable.
But:
M1 → M2 → M3
being displayed as:
M2 → M1 → M3
is much more problematic.
25. API Design
A simplified API might look like:
Send message
http
POST /conversations/{conversationId}/messages
{
"messageId": "m123",
"type": "TEXT",
"ciphertext": "..."
}
Get messages
http
GET /conversations/{conversationId}/messages?before=m123&limit=50
Mark as read
http
POST /messages/{messageId}/read
WebSocket events
MESSAGE_RECEIVED
MESSAGE_DELIVERED
MESSAGE_READ
TYPING_STARTED
TYPING_STOPPED
PRESENCE_CHANGED
26. Security
A production system should consider:
- TLS
- end-to-end encryption
- device authentication
- key rotation
- rate limiting
- abuse detection
- spam prevention
- account takeover protection
- encrypted media
- secure push notifications
- metadata protection
- device revocation
One subtle point:
E2E encryption protects message content, but it doesn’t automatically hide all metadata.
The system may still know things such as:
sender
recipient
timestamp
message size
IP/device information
depending on the design.
27. Monitoring
For a production WhatsApp-like platform, I’d monitor:
Infrastructure
CPU
Memory
Network
Disk
Kafka lag
Database latency
Messaging
Message throughput
Message delivery latency
Delivery failure rate
Duplicate rate
Retry rate
Queue depth
User experience
P50 delivery latency
P95 delivery latency
P99 delivery latency
Offline message delay
Connection success rate
The most important business-facing metric could be:
End-to-end message delivery latency.
28. The Key Interview Trade-offs
If the interviewer pushes deeper, focus on these:
| Problem | Design Decision |
|---|---|
| Real-time messaging | WebSocket |
| Offline users | Persistent message store + push notification |
| Message ordering | Partition by conversation |
| Massive scale | Horizontal scaling + sharding |
| Duplicate delivery | Idempotent message IDs |
| Presence | Redis + TTL |
| Message streaming | Kafka |
| Message storage | Cassandra/DynamoDB/ScyllaDB |
| Media | Object storage + CDN |
| Large groups | Hybrid fan-out |
| Multi-device | Device-level sessions |
| Security | End-to-end encryption |
| Failure recovery | Replication + retry + DLQ |
| High availability | Multi-AZ / multi-region |
| Observability | Metrics + tracing + logs |
29. A Strong 2-Minute Interview Answer
If the interviewer says “Give me the high-level design”, don’t spend 20 minutes explaining every component.
A strong answer would be:
“I’d design WhatsApp around a persistent WebSocket connection for real-time communication, a stateless messaging layer for message ingestion, and a durable distributed message store for offline delivery.
Clients establish WebSocket connections through connection gateways. A distributed presence/connection registry tracks which gateway currently owns each user’s connection.
When Alice sends Bob a message, the message service authenticates the request, validates the conversation, assigns a conversation-level sequence number, persists the encrypted message, and publishes an event to a partitioned message bus such as Kafka. The conversation ID is the partition key so messages within a conversation maintain ordering.
If Bob is online, the delivery service routes the message to the gateway holding Bob’s WebSocket connection. If he’s offline, the message remains in durable storage and a notification service sends a push notification through APNs or FCM. When Bob reconnects, the client fetches or receives all undelivered messages.
I’d use idempotent message IDs because distributed delivery is typically at-least-once, and the client/server must tolerate duplicate delivery.
For storage, I’d use a horizontally scalable database such as Cassandra, ScyllaDB, or DynamoDB, partitioned primarily by conversation ID. Presence would live in Redis with TTL because it is ephemeral.
Images and videos would bypass the messaging servers and go directly to object storage, with a CDN handling downloads.
For large groups, I’d use a hybrid fan-out strategy: fan-out-on-write for small groups and fan-out-on-read or partitioned asynchronous fan-out for very large groups.
Finally, I’d design the system for multi-AZ and potentially multi-region availability, with end-to-end encryption so the messaging infrastructure transports encrypted payloads without requiring access to plaintext.”
That answer demonstrates distributed systems thinking, rather than simply listing technologies.
30. The Most Important Interview Insight
The best answer isn’t:
“Use Kafka + Redis + Cassandra + Kubernetes.”
That’s technology-first thinking.
The stronger answer is:
Requirements → communication model → data flow → consistency → partitioning → failure handling → scaling → security.
For a senior/lead-level system design interview, the interviewer is usually evaluating whether you can explain why each architectural decision exists.
A useful mental model is:
┌───────────────────┐
│ WhatsApp │
└─────────┬─────────┘
│
┌──────────────┼──────────────┐
▼ ▼ ▼
CONNECT STORE DELIVER
│ │ │
WebSocket Database Push/WebSocket
│ │ │
└──────────────┼──────────────┘
▼
SCALE
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Sharding Partitioning Replication
│
▼
TRUST
│
Encryption + Security
This is the structure I’d use in a Senior/Lead Java + System Design interview because it gives you a clear path from requirements all the way to distributed-systems trade-offs.