Skip to content
Bill Liao
Go back

WhatsApp System Design Interview Question and Answer

Edit page

How to Design WhatsApp — System Design Interview Question & Answer

1. Interview Question

Design a messaging system like WhatsApp.

The interviewer expects you to discuss:

A good interview answer should not start with databases or Kafka. Start by clarifying requirements and then progressively build the architecture.


2. Clarify Requirements

I would begin the interview by asking:

Functional requirements

  1. Can users send 1-to-1 messages?
  2. Do we support group chats?
  3. Do messages need delivery and read receipts?
  4. Should users see online/offline presence?
  5. Do we support images, videos, documents, and voice messages?
  6. Should users receive messages when offline?
  7. Should messages synchronize across multiple devices?
  8. Do we need message search?
  9. Do we need voice/video calls?

For a 45-minute interview, I’d explicitly scope out voice/video calling unless the interviewer asks for it.

Non-functional requirements

We want:


3. High-Level Architecture

A reasonable architecture looks like this:

                         ┌─────────────────────┐
                         │      Clients        │
                         │ iOS / Android / Web │
                         └──────────┬──────────┘
                                    │
                         WebSocket / HTTPS
                                    │
                                    ▼
                         ┌─────────────────────┐
                         │   Load Balancer     │
                         └──────────┬──────────┘
                                    │
                  ┌─────────────────┼─────────────────┐
                  │                 │                 │
                  ▼                 ▼                 ▼
           ┌───────────┐     ┌────────────┐    ┌─────────────┐
           │Connection │     │ Message    │    │   Presence  │
           │  Service  │     │  Service   │    │   Service   │
           └─────┬─────┘     └─────┬──────┘    └──────┬──────┘
                 │                 │                  │
                 │                 ▼                  │
                 │          ┌─────────────┐            │
                 │          │ Message Bus │            │
                 │          │   Kafka     │            │
                 │          └──────┬──────┘            │
                 │                 │                   │
                 │        ┌────────┴─────────┐         │
                 │        ▼                  ▼         │
                 │ ┌─────────────┐   ┌─────────────┐  │
                 │ │ Message DB  │   │ Notification│  │
                 │ │             │   │   Service   │  │
                 │ └─────────────┘   └──────┬──────┘  │
                 │                          │         │
                 │                          ▼         │
                 │                   APNs / FCM      │
                 │                                    │
                 └──────────────┬─────────────────────┘
                                │
                                ▼
                         ┌──────────────┐
                         │ Redis / KV   │
                         │ Presence     │
                         │ Connections  │
                         └──────────────┘

Media:
Client → Object Storage → CDN

The important design principle is:

Separate persistent messaging from real-time connection management.


4. Client Connection

WhatsApp needs persistent connections because users expect messages to arrive immediately.

The client establishes a:

Client
   │
   │ WebSocket
   ▼
Connection Gateway

The WebSocket remains open while the application is active.

For example:

Alice
  │
  │ WebSocket
  ▼
Connection Server #17

The connection server maintains:

userId → connectionId

However, this mapping cannot live only in local memory.

We need distributed connection state:

Redis

alice → server-17
bob   → server-42
john  → server-8

When Alice sends a message to Bob:

Alice
  │
  ▼
Connection Server 17
  │
  ▼
Message Service
  │
  ▼
Find Bob's connection
  │
  ▼
Connection Server 42
  │
  ▼
Bob

5. Sending a Message

Suppose Alice sends:

“Hello Bob”

The client generates a message ID:

messageId = UUID
conversationId = C123
senderId = Alice
receiverId = Bob
timestamp = ...

The request might look like:

{
  "messageId": "m123",
  "conversationId": "c456",
  "senderId": "alice",
  "type": "TEXT",
  "ciphertext": "encrypted-content"
}

Notice something important:

The server doesn’t necessarily see the plaintext.

For true end-to-end encryption, the client encrypts the message before sending it.


6. Message Processing

The Message Service receives the message.

Alice
  │
  ▼
Message Service
  │
  ├── Validate
  ├── Authenticate
  ├── Authorize
  ├── Assign ordering
  ├── Persist
  │
  ▼
Kafka
  │
  ├── Delivery
  ├── Notification
  └── Analytics

The message should be persisted before acknowledging successful acceptance.

This prevents:

Client → Server → ACK
                 X
              crash

where the client thinks the message was accepted but the message was actually lost.


7. Message Ordering

Ordering is one of the most important interview topics.

Imagine:

Alice sends:

M1: "Are you free?"
M2: "Let's meet at 7."
M3: "At the restaurant."

The recipient should normally see:

M1
M2
M3

not:

M2
M1
M3

Approach

Partition messages by conversation:

Kafka

Partition 0 → Conversation A
Partition 1 → Conversation B
Partition 2 → Conversation C

Use:

partitionKey = conversationId

This ensures messages from the same conversation are processed sequentially by the same partition.

However, be careful:

Kafka ordering is guaranteed within a partition, not globally.

That’s actually exactly what we want.

We need per-conversation ordering, not global ordering across WhatsApp.


8. Delivery States

A WhatsApp-style system typically has several states:

SENT
  ↓
DELIVERED
  ↓
READ

For example:

Alice                    Bob

  M1 ────────────────────►
       SENT

       ◄──────────────────
          DELIVERED

       ◄──────────────────
             READ

We can represent:

messageId = M123

senderStatus:
    SENT

receiverStatus:
    DELIVERED

readStatus:
    READ

9. What Happens If Bob Is Offline?

This is a classic interview question.

Suppose:

Alice → Bob

but Bob has no active WebSocket connection.

The Message Service checks:

Presence Service

Bob → OFFLINE

The message remains persisted.

Then:

Notification Service
        │
        ├── iOS → APNs
        │
        └── Android → FCM

Bob receives:

You have a new message.

When Bob opens WhatsApp:

Bob
 │
 ▼
WebSocket
 │
 ▼
Message Service
 │
 ▼
Fetch undelivered messages
 │
 ▼
M1
M2
M3

Then the client sends acknowledgements.


10. Message Delivery Is Usually At-Least-Once

This is another important interview point.

Suppose Bob receives:

M123

but the ACK is lost.

The server may resend:

M123

Therefore the client must tolerate duplicates.

Use:

messageId

as an idempotency key.

The client can maintain:

processedMessageIds

and ignore duplicates.

Therefore:

At-least-once delivery + idempotent processing is often more practical than trying to guarantee exactly-once delivery across a distributed system.


11. Database Design

A possible message model:

Message
---------------------------
message_id
conversation_id
sender_id
sequence_number
timestamp
message_type
encrypted_payload
status

For example:

conversation_id | sequence | sender | message
------------------------------------------------
C100            | 101      | Alice  | Hello
C100            | 102      | Bob    | Hi
C100            | 103      | Alice  | How are you?

Why conversation_id?

Because almost every query is:

WHERE conversation_id = ?
ORDER BY sequence_number

rather than:

SQL

WHERE message_id = ?

12. Choosing the Database

For a massive messaging system, a distributed NoSQL database is attractive.

Potential choices:

For example:

Partition Key:
conversation_id

Sort Key:
sequence_number

This gives efficient:

Get conversation history

operations.

A relational database can work at smaller scale, but eventually horizontal partitioning becomes critical.


13. Database Sharding

Suppose we have billions of messages.

We cannot put everything on one database.

Partition by:

hash(conversationId) % N

Example:

Conversation A → Shard 1
Conversation B → Shard 8
Conversation C → Shard 4
Conversation D → Shard 2

The key is:

Keep messages from the same conversation on the same logical partition whenever possible.

That makes ordering and conversation-history queries much easier.


14. Group Messaging

Now suppose we have:

Group G1

Alice
Bob
Charlie
David
Emma

Alice sends:

“Meeting at 7.”

There are two major approaches.

Fan-out on write

Immediately create delivery records:

M1 → Bob
M1 → Charlie
M1 → David
M1 → Emma

Advantages:

Disadvantages:


Fan-out on read

Store:

Group G1
   │
   └── M1

Members retrieve messages when they open the group.

Advantages:

Disadvantages:

A production system can use a hybrid strategy.

For example:

Small group → fan-out on write

Large group → fan-out on read

15. Presence Service

Users expect:

Bob
Online

or:

Bob
Last seen 5 minutes ago

Presence is usually an ephemeral state.

A good architecture:

Client
  │
  │ heartbeat
  ▼
Presence Service
  │
  ▼
Redis

Example:

userId → {
    status: ONLINE,
    lastSeen: 2026-08-09T20:00:00
}

Use TTL:

Bob → ONLINE
TTL = 30 seconds

If Bob stops sending heartbeats:

TTL expires
     ↓
Bob → OFFLINE

This avoids requiring expensive database writes for every presence update.


16. Media Messages

Images and videos should not go through the messaging service.

Bad architecture:

Client
  │
  ▼
Message Server
  │
  ▼
Image

This consumes huge amounts of application-server bandwidth.

Instead:

Client
  │
  │ Upload
  ▼
Object Storage
  │
  ▼
CDN

Examples:

S3
GCS
Azure Blob Storage

The message contains metadata:

{
  "type": "IMAGE",
  "mediaId": "img123",
  "url": "...",
  "size": 2048000
}

For secure messaging, the media itself should also be encrypted appropriately.


17. CDN

Media downloads are highly cacheable.

Instead of:

Bob → WhatsApp Server → Storage

use:

Bob
 │
 ▼
CDN
 │
 ├── Cache HIT → image
 │
 └── Cache MISS
        │
        ▼
     Object Storage

This dramatically reduces origin traffic.


18. Multi-Device Support

A user might have:

Alice
 ├── iPhone
 ├── Mac
 ├── iPad
 └── Web browser

The system therefore needs device-level identity.

Instead of:

userId → connection

use:

userId
   │
   ├── device-A
   ├── device-B
   ├── device-C
   └── device-D

Each device has:

deviceId
publicKey
connection
lastSeen

Messages may need to be delivered to multiple devices.

This also becomes important for end-to-end encryption.


19. End-to-End Encryption

For a WhatsApp-like system, security is fundamental.

Conceptually:

Alice
  │
  │ Encrypt
  ▼
Ciphertext
  │
  ▼
WhatsApp Servers
  │
  │ Cannot decrypt
  ▼
Bob
  │
  │ Decrypt
  ▼
Plaintext

The server should not need access to plaintext messages.

A modern design can use protocols based on the Signal Protocol concepts:

Identity Keys
Prekeys
Session Keys
Ratchets

The important interview statement is:

The messaging infrastructure routes and stores encrypted payloads, while cryptographic keys remain under client control.


20. WhatsApp-Style Architecture

Putting everything together:

                         ┌───────────────────┐
                         │      Clients      │
                         │ iOS / Android/Web │
                         └─────────┬─────────┘
                                   │
                          TLS / WebSocket
                                   │
                                   ▼
                         ┌───────────────────┐
                         │  Load Balancer    │
                         └─────────┬─────────┘
                                   │
                    ┌──────────────┼──────────────┐
                    │              │              │
                    ▼              ▼              ▼
             ┌────────────┐ ┌────────────┐ ┌────────────┐
             │ Connection │ │  Message   │ │  Presence  │
             │  Gateway   │ │  Service   │ │  Service   │
             └─────┬──────┘ └──────┬─────┘ └──────┬─────┘
                   │               │              │
                   │               ▼              ▼
                   │        ┌──────────────┐   ┌───────┐
                   │        │    Kafka     │   │ Redis │
                   │        └──────┬───────┘   └───────┘
                   │               │
                   │       ┌───────┴────────┐
                   │       │                │
                   ▼       ▼                ▼
             ┌─────────┐ ┌────────────┐ ┌─────────────┐
             │ Redis   │ │ Message DB │ │ Notification│
             │ Routing │ │ Cassandra/ │ │   Service   │
             │         │ │ DynamoDB   │ └──────┬──────┘
             └─────────┘ └────────────┘        │
                                               ▼
                                         ┌───────────┐
                                         │ APNs/FCM  │
                                         └───────────┘

                    Media Pipeline
                    ───────────────

Client ──► Object Storage ──► CDN ──► Client

21. Scaling the System

Let’s assume:

500M daily active users
100M concurrent connections
50B messages/day

The architecture must scale horizontally.

Instead of:

1 Message Server

we have:

Message Server
Message Server
Message Server
...
Message Server

Stateless services can scale using:

Kubernetes
+
Load Balancer
+
Auto Scaling

Connection servers are slightly different because WebSocket connections are long-lived.

We therefore need:

Connection Gateway
      +
Connection Registry

to know which server owns each connection.


22. Handling Hot Conversations

A major scalability problem occurs when a single group becomes extremely popular.

For example:

Group:
100,000 members

If one message arrives:

1 message
     ↓
100,000 deliveries

This creates a fan-out storm.

Solutions include:

This is a great point to mention in an interview because it demonstrates that you’re thinking beyond the happy path.


23. Failure Handling

What happens if Kafka is temporarily unavailable?

The Message Service should avoid simply dropping messages.

Possible approach:

Client
  │
  ▼
Message Service
  │
  ├── Primary message store
  │
  └── Kafka

Use retries with:

exponential backoff
+
dead-letter queue

For duplicate requests:

messageId

provides idempotency.

For database failures:

Replication
+
Multi-AZ
+
Automatic failover

For regional failures:

Region A
   │
   ├── replicated data
   │
   ▼
Region B

24. Consistency vs Availability

This is a classic distributed-systems trade-off.

For messaging:

Strong consistency is useful for:

Eventual consistency is acceptable for:

For example, if Bob’s status changes:

ONLINE → OFFLINE

being delayed by 1–2 seconds is generally acceptable.

But:

M1 → M2 → M3

being displayed as:

M2 → M1 → M3

is much more problematic.


25. API Design

A simplified API might look like:

Send message

http

POST /conversations/{conversationId}/messages
{
  "messageId": "m123",
  "type": "TEXT",
  "ciphertext": "..."
}

Get messages

http

GET /conversations/{conversationId}/messages?before=m123&limit=50

Mark as read

http

POST /messages/{messageId}/read

WebSocket events

MESSAGE_RECEIVED
MESSAGE_DELIVERED
MESSAGE_READ
TYPING_STARTED
TYPING_STOPPED
PRESENCE_CHANGED

26. Security

A production system should consider:

One subtle point:

E2E encryption protects message content, but it doesn’t automatically hide all metadata.

The system may still know things such as:

sender
recipient
timestamp
message size
IP/device information

depending on the design.


27. Monitoring

For a production WhatsApp-like platform, I’d monitor:

Infrastructure

CPU
Memory
Network
Disk
Kafka lag
Database latency

Messaging

Message throughput
Message delivery latency
Delivery failure rate
Duplicate rate
Retry rate
Queue depth

User experience

P50 delivery latency
P95 delivery latency
P99 delivery latency
Offline message delay
Connection success rate

The most important business-facing metric could be:

End-to-end message delivery latency.


28. The Key Interview Trade-offs

If the interviewer pushes deeper, focus on these:

ProblemDesign Decision
Real-time messagingWebSocket
Offline usersPersistent message store + push notification
Message orderingPartition by conversation
Massive scaleHorizontal scaling + sharding
Duplicate deliveryIdempotent message IDs
PresenceRedis + TTL
Message streamingKafka
Message storageCassandra/DynamoDB/ScyllaDB
MediaObject storage + CDN
Large groupsHybrid fan-out
Multi-deviceDevice-level sessions
SecurityEnd-to-end encryption
Failure recoveryReplication + retry + DLQ
High availabilityMulti-AZ / multi-region
ObservabilityMetrics + tracing + logs

29. A Strong 2-Minute Interview Answer

If the interviewer says “Give me the high-level design”, don’t spend 20 minutes explaining every component.

A strong answer would be:

“I’d design WhatsApp around a persistent WebSocket connection for real-time communication, a stateless messaging layer for message ingestion, and a durable distributed message store for offline delivery.

Clients establish WebSocket connections through connection gateways. A distributed presence/connection registry tracks which gateway currently owns each user’s connection.

When Alice sends Bob a message, the message service authenticates the request, validates the conversation, assigns a conversation-level sequence number, persists the encrypted message, and publishes an event to a partitioned message bus such as Kafka. The conversation ID is the partition key so messages within a conversation maintain ordering.

If Bob is online, the delivery service routes the message to the gateway holding Bob’s WebSocket connection. If he’s offline, the message remains in durable storage and a notification service sends a push notification through APNs or FCM. When Bob reconnects, the client fetches or receives all undelivered messages.

I’d use idempotent message IDs because distributed delivery is typically at-least-once, and the client/server must tolerate duplicate delivery.

For storage, I’d use a horizontally scalable database such as Cassandra, ScyllaDB, or DynamoDB, partitioned primarily by conversation ID. Presence would live in Redis with TTL because it is ephemeral.

Images and videos would bypass the messaging servers and go directly to object storage, with a CDN handling downloads.

For large groups, I’d use a hybrid fan-out strategy: fan-out-on-write for small groups and fan-out-on-read or partitioned asynchronous fan-out for very large groups.

Finally, I’d design the system for multi-AZ and potentially multi-region availability, with end-to-end encryption so the messaging infrastructure transports encrypted payloads without requiring access to plaintext.”

That answer demonstrates distributed systems thinking, rather than simply listing technologies.


30. The Most Important Interview Insight

The best answer isn’t:

“Use Kafka + Redis + Cassandra + Kubernetes.”

That’s technology-first thinking.

The stronger answer is:

Requirements → communication model → data flow → consistency → partitioning → failure handling → scaling → security.

For a senior/lead-level system design interview, the interviewer is usually evaluating whether you can explain why each architectural decision exists.

A useful mental model is:

                  ┌───────────────────┐
                  │     WhatsApp      │
                  └─────────┬─────────┘
                            │
             ┌──────────────┼──────────────┐
             ▼              ▼              ▼
         CONNECT          STORE          DELIVER
             │              │              │
          WebSocket       Database       Push/WebSocket
             │              │              │
             └──────────────┼──────────────┘
                            ▼
                         SCALE
                            │
              ┌─────────────┼─────────────┐
              ▼             ▼             ▼
           Sharding      Partitioning   Replication
                            │
                            ▼
                         TRUST
                            │
                    Encryption + Security

This is the structure I’d use in a Senior/Lead Java + System Design interview because it gives you a clear path from requirements all the way to distributed-systems trade-offs.


Edit page
Share this post:

Previous Post
Airbnb System Design Interview Question and Answer
Next Post
100 Kafka interview Questions and Answers