100 Most Frequently Asked Apache Kafka Interview Questions & Answers
Below is a comprehensive Kafka interview guide, progressing from fundamentals to advanced architecture, performance, Kafka Streams, security, operations, and system-design questions.
1. Kafka Fundamentals
1. What is Apache Kafka?
Answer:
Apache Kafka is a distributed event-streaming platform designed to handle high-throughput, fault-tolerant, scalable streams of data.
Kafka is commonly used for:
- Event-driven architectures
- Microservice communication
- Log aggregation
- Real-time analytics
- Data pipelines
- Event sourcing
- Stream processing
Kafka stores events durably and allows multiple consumers to independently process the same events.
2. What are the main components of Kafka?
Answer:
The core Kafka components are:
- Producer — publishes records
- Consumer — reads records
- Broker — Kafka server that stores and serves records
- Topic — logical category of events
- Partition — ordered, append-only log within a topic
- Consumer Group — group of consumers cooperating to process partitions
- Controller — manages cluster metadata and partition leadership
- Kafka Streams — library for stream processing
Modern Kafka clusters use KRaft for metadata management instead of ZooKeeper.
3. What is a Kafka topic?
Answer:
A topic is a logical stream/category to which producers publish records and from which consumers read records.
For example:
customer-events
payment-events
order-events
A topic is divided into one or more partitions for scalability.
4. What is a Kafka partition?
Answer:
A partition is an ordered, append-only sequence of records.
For example:
orders
├── partition-0
├── partition-1
└── partition-2
Partitions allow Kafka to distribute data across brokers and process records in parallel.
5. Why does Kafka use partitions?
Answer:
Partitions provide:
- Horizontal scalability
- Parallel processing
- Higher throughput
- Fault tolerance through replication
If a topic has 10 partitions, multiple consumers can process those partitions concurrently.
6. What is a Kafka offset?
Answer:
An offset is a monotonically increasing number identifying a record’s position within a partition.
Example:
Partition 0
Offset Event
0 Order A
1 Order B
2 Order C
3 Order D
Offsets are maintained per partition.
7. Are Kafka offsets globally unique?
Answer:
No.
Offsets are only unique within a partition.
For example:
Partition 0 → offset 10
Partition 1 → offset 10
Both are valid because offsets belong to different partitions.
8. Does Kafka guarantee message ordering?
Answer:
Kafka guarantees ordering within a partition, not across an entire topic.
If ordering matters for a customer:
key = customerId
Kafka will normally route all events for that key to the same partition, preserving their order.
9. What is a Kafka broker?
Answer:
A broker is a Kafka server responsible for storing partitions and serving producer and consumer requests.
A Kafka cluster normally contains multiple brokers:
Kafka Cluster
├── Broker 1
├── Broker 2
└── Broker 3
10. What is a Kafka cluster?
Answer:
A Kafka cluster is a collection of Kafka brokers working together to provide:
- Scalability
- Availability
- Fault tolerance
- Data replication
- Distributed processing
2. Producers
11. What is a Kafka producer?
Answer:
A producer is an application that publishes records to Kafka topics.
For example:
producer.send(
new ProducerRecord<>("orders", orderId, order)
);
The producer determines the destination topic and, directly or indirectly, the partition.
12. How does a Kafka producer choose a partition?
Answer:
If a key is provided, Kafka normally hashes the key to determine the partition.
Conceptually:
partition = hash(key) % numberOfPartitions
Without a key, the producer can distribute records across partitions according to the producer’s partitioning strategy.
13. Why should Kafka messages have keys?
Answer:
Keys are useful when related events must remain ordered.
For example:
customerId = 123
Events:
CustomerCreated
CustomerUpdated
CustomerDeleted
can be routed to the same partition.
14. What is acks in Kafka?
Answer:
acks controls how much acknowledgement the producer requires from Kafka.
Common values:
acks=0
acks=1
acks=all
acks=all provides the strongest durability guarantee when combined with appropriate replication settings.
15. Difference between acks=0, acks=1, and acks=all?
Answer:
| Setting | Meaning | Durability |
|---|---|---|
0 | No broker acknowledgement | Lowest |
1 | Leader acknowledges | Medium |
all | Required in-sync replicas acknowledge | Highest |
For important financial events, acks=all is usually preferred.
16. What is min.insync.replicas?
Answer:
min.insync.replicas specifies the minimum number of replicas that must be in sync for a producer using acks=all to successfully write.
Example:
replication.factor = 3
min.insync.replicas = 2
acks = all
At least two replicas must be in sync.
17. What is producer batching?
Answer:
Kafka producers can accumulate multiple records into batches before sending them.
Important settings include:
batch.size
linger.ms
Larger batches can improve throughput but may increase latency.
18. What does linger.ms do?
Answer:
linger.ms determines how long the producer may wait for additional records before sending a batch.
Example:
properties
linger.ms=5
The producer may wait up to approximately 5 ms to build a larger batch.
19. What does batch.size control?
Answer:
batch.size controls the maximum size of a producer batch per partition.
It does not mean the producer waits until the batch is completely full. linger.ms determines whether it waits for more records.
20. How can you improve Kafka producer throughput?
Answer:
Common approaches:
- Use batching
- Tune
linger.ms - Increase
batch.size - Enable compression
- Use asynchronous sends
- Avoid unnecessarily large messages
- Increase partition count when appropriate
- Use efficient serialization
- Tune producer buffer memory
For example:
compression.type=zstd
linger.ms=5
batch.size=65536
acks=all
The exact values should be benchmarked rather than blindly copied.
3. Consumers
21. What is a Kafka consumer?
Answer:
A consumer reads records from Kafka topics.
Consumers typically belong to consumer groups and track their offsets.
22. What is a consumer group?
Answer:
A consumer group is a collection of consumers cooperating to consume a topic.
Within one consumer group:
Each partition is assigned to at most one consumer at a time.
Example:
Topic: 6 partitions
Consumer Group
├── Consumer A → P0, P1
├── Consumer B → P2, P3
└── Consumer C → P4, P5
23. Can two consumers in the same group read the same partition simultaneously?
Answer:
Normally, no.
A partition is assigned to only one consumer within a consumer group.
However, consumers in different groups can independently consume the same partition.
24. Can multiple consumer groups consume the same topic?
Answer:
Yes.
For example:
orders topic
|
┌───┼────────────┐
↓ ↓ ↓
Billing Analytics Fraud
Group Group Group
Each group maintains its own offsets.
25. What happens if there are more consumers than partitions?
Answer:
Some consumers will remain idle.
For example:
3 partitions
5 consumers
Only three consumers can actively consume partitions at that time.
Therefore:
Maximum parallelism within a consumer group is bounded by the number of partitions.
26. What happens when a consumer crashes?
Answer:
Kafka detects the consumer failure and triggers a rebalance.
Its partitions are reassigned to other consumers in the group.
27. What is consumer rebalancing?
Answer:
Rebalancing is the process of redistributing partitions among consumers in a consumer group.
It can happen when:
- A consumer joins
- A consumer leaves
- A consumer crashes
- Topic partitions change
- Group membership changes
28. What is a consumer offset?
Answer:
A consumer offset represents the position a consumer has reached in a partition.
Kafka stores committed consumer-group offsets in the internal topic:
__consumer_offsets
29. What is auto offset commit?
Answer:
With:
properties
enable.auto.commit=true
the Kafka consumer automatically commits offsets periodically.
The default behavior can be convenient, but it may lead to undesirable processing semantics if the application crashes between processing and committing.
30. What is manual offset commit?
Answer:
The application explicitly commits offsets.
For example:
Java
consumer.commitSync();
This provides more control over when an event is considered successfully processed.
4. Delivery Semantics
31. What is at-most-once delivery?
Answer:
Read
↓
Commit
↓
Process
If the application crashes after committing but before processing, the message may be lost.
So:
At-most-once = no duplicates, but messages may be lost.
32. What is at-least-once delivery?
Answer:
Read
↓
Process
↓
Commit
If the application crashes after processing but before committing, Kafka may deliver the message again.
Therefore:
At-least-once = messages are not intentionally lost, but duplicates are possible.
33. What is exactly-once processing?
Answer:
Exactly-once semantics aim to ensure that a Kafka processing pipeline does not produce duplicate effects from retries.
Kafka supports exactly-once processing using mechanisms such as:
- Transactions
- Idempotent producers
- Transactional offsets
- Kafka Streams exactly-once processing
However, exactly-once is not automatically guaranteed across arbitrary external systems such as a traditional database or REST API.
34. What is idempotency?
Answer:
An operation is idempotent if executing it multiple times produces the same effective result as executing it once.
For example:
SET account_balance = 100
is naturally more idempotent than:
account_balance += 100
Kafka consumers should often be designed to tolerate duplicate delivery.
35. How do you make a Kafka consumer idempotent?
Answer:
Common approaches include:
- Unique event IDs
- Database unique constraints
- Idempotency tables
- Upsert operations
- State checks
- Transactional processing
Example:
event_id = 8f72...
Store processed event IDs so duplicates can be ignored.
5. Replication & Fault Tolerance
36. What is Kafka replication?
Answer:
Replication means Kafka maintains multiple copies of a partition across brokers.
Example:
Partition 0
Broker 1 → Leader
Broker 2 → Follower
Broker 3 → Follower
If the leader fails, another replica can become leader.
37. What is replication factor?
Answer:
Replication factor specifies how many copies of each partition exist.
Example:
replication.factor=3
means three replicas exist.
38. What is a leader replica?
Answer:
The leader handles reads and writes for a partition.
Followers replicate data from the leader.
39. What is a follower replica?
Answer:
A follower replica maintains a copy of a partition and replicates records from the leader.
Followers can become leaders if the current leader fails.
40. What is ISR?
Answer:
ISR means In-Sync Replicas.
It is the set of replicas considered sufficiently caught up with the partition leader.
Example:
Replication Factor = 3
Broker 1 → Leader
Broker 2 → ISR
Broker 3 → ISR
41. What happens if a broker containing a partition leader fails?
Answer:
Kafka elects another eligible replica as the new leader.
Clients refresh metadata and begin communicating with the new leader.
42. What is an unclean leader election?
Answer:
An unclean leader election allows a replica that is not in the ISR to become leader.
This can reduce data-loss protection.
Typically:
properties
unclean.leader.election.enable=false
is preferred when durability is more important than availability.
43. What is rack awareness in Kafka?
Answer:
Rack awareness allows Kafka to distribute replicas across different failure domains.
For example:
Rack A → Broker 1
Rack B → Broker 2
Rack C → Broker 3
This protects against an entire rack or availability zone failing.
6. Kafka Storage
44. How does Kafka store messages?
Answer:
Kafka stores records in append-only log files called log segments.
Conceptually:
Partition
├── segment-0001
├── segment-0002
└── segment-0003
This sequential I/O model contributes to Kafka’s high throughput.
45. What is a log segment?
Answer:
A log segment is a physical file containing records for a partition.
Kafka periodically rolls segments based on configuration such as:
segment.bytes
segment.ms
46. What is log retention?
Answer:
Retention determines how long Kafka keeps records.
Common policies include:
retention.ms
retention.bytes
Kafka can therefore retain data for:
- Hours
- Days
- Months
- Longer, depending on storage requirements
47. Does Kafka delete messages after they are consumed?
Answer:
No.
This is a major difference from traditional queues.
Kafka retains records according to the topic’s retention policy, regardless of whether consumers have already read them.
48. Can a consumer reread old Kafka messages?
Answer:
Yes.
A consumer can reset its offset and replay historical records as long as those records are still retained.
For example:
kafka-consumer-groups \
--reset-offsets \
--to-earliest
49. What is log compaction?
Answer:
Log compaction retains the latest record for each key.
Example:
key=A value=10
key=B value=20
key=A value=15
After compaction, Kafka can retain:
A → 15
B → 20
Compaction is useful for maintaining the latest state of entities.
50. Difference between retention and compaction?
Answer:
| Retention | Compaction |
|---|---|
| Removes old records based on time/size | Removes older records for the same key |
| Good for event history | Good for latest state |
retention.ms / retention.bytes | cleanup.policy=compact |
A topic can also use:
properties
cleanup.policy=compact,delete
7. Kafka Performance
51. Why is Kafka so fast?
Answer:
Kafka achieves high throughput through several design choices:
- Sequential disk I/O
- Append-only logs
- Batching
- Compression
- Partition parallelism
- Efficient network transfer
- Zero-copy techniques
- Page cache
- Distributed architecture
52. What is zero-copy?
Answer:
Zero-copy allows data to move from the file system/page cache to the network without unnecessary user-space copying.
This reduces CPU and memory overhead.
53. How can Kafka consumer throughput be increased?
Answer:
Possible strategies:
- Increase partition count
- Increase consumers up to partition count
- Increase fetch sizes
- Optimize processing
- Batch processing
- Tune
fetch.min.bytes - Tune
fetch.max.wait.ms - Increase parallelism
- Reduce expensive downstream operations
54. What is fetch.min.bytes?
Answer:
It specifies the minimum amount of data the broker should return to a consumer request when possible.
Larger values can improve throughput but potentially increase latency.
55. What is fetch.max.wait.ms?
Answer:
It controls how long the broker may wait to accumulate enough data to satisfy fetch.min.bytes.
56. What is max.poll.records?
Answer:
It limits the number of records returned by one poll() call.
Example:
properties
max.poll.records=500
This can help control the amount of work processed per poll.
57. What is max.poll.interval.ms?
Answer:
It defines the maximum interval between successful calls to poll() before Kafka considers the consumer unable to keep up.
If processing takes too long, the consumer may leave the group and trigger a rebalance.
58. What is consumer lag?
Answer:
Consumer lag measures how far a consumer is behind the latest available records.
Conceptually:
Latest Offset - Consumer Offset = Lag
High lag indicates that consumers are not keeping up with producers.
59. How do you troubleshoot high consumer lag?
Answer:
Check:
- Consumer processing latency
- Number of partitions
- Number of consumers
- CPU/memory utilization
- Downstream database/API latency
max.poll.interval.ms- Fetch configuration
- Broker performance
- Network latency
- Consumer errors/rebalances
60. Can adding consumers always reduce consumer lag?
Answer:
No.
If the topic has only 3 partitions:
3 partitions
10 consumers
only three consumers can actively process partitions.
To increase parallelism, the topic may need more partitions.
8. Kafka Architecture
61. What is KRaft?
Answer:
KRaft is Kafka’s metadata management architecture based on Kafka’s own Raft-based consensus mechanism.
It replaces the historical dependency on ZooKeeper.
Modern Kafka deployments can therefore operate without ZooKeeper.
62. What was ZooKeeper used for?
Answer:
Older Kafka architectures used ZooKeeper for:
- Controller election
- Cluster metadata
- Broker registration
- Configuration coordination
- Partition leadership management
Modern Kafka deployments use KRaft instead.
63. What is a Kafka controller?
Answer:
The controller manages cluster-level metadata operations such as:
- Partition leadership
- Replica state
- Broker membership
- Metadata changes
In KRaft mode, controller nodes participate in the metadata quorum.
64. What is a Kafka metadata quorum?
Answer:
In KRaft, Kafka metadata is replicated among controller nodes using a Raft-based quorum.
This allows Kafka to manage cluster metadata without ZooKeeper.
65. What is the difference between broker and controller roles?
Answer:
A broker handles:
- Client requests
- Partition storage
- Replication
A controller handles:
- Cluster metadata
- Leadership changes
- Metadata quorum operations
Kafka deployments can use combined or separated roles depending on architecture and scale.
9. Kafka Transactions
66. What is a Kafka transaction?
Answer:
A Kafka transaction allows a producer to atomically write multiple records, potentially across multiple partitions and topics.
Example:
Topic A
Topic B
Topic C
can be written as one transactional unit.
67. What is a transactional producer?
Answer:
A producer configured with a unique transactional.id can participate in Kafka transactions.
Example:
enable.idempotence=true
transactional.id=payment-service-1
68. What is transactional.id?
Answer:
transactional.id uniquely identifies a transactional producer.
It allows Kafka to detect producer restarts and support transactional semantics.
69. What is read_committed?
Answer:
A consumer configured with:
properties
isolation.level=read_committed
reads only committed transactional records.
It will not expose aborted transactional records.
70. What is read_uncommitted?
Answer:
It allows consumers to see records regardless of whether their transaction ultimately commits or aborts.
It is the default isolation level.
10. Kafka Streams
71. What is Kafka Streams?
Answer:
Kafka Streams is a Java library for processing Kafka data streams.
It supports:
- Filtering
- Mapping
- Aggregation
- Joins
- Windowing
- Stateful processing
- Exactly-once processing
72. What is a KStream?
Answer:
A KStream represents an unbounded stream of records.
Example:
KStream<String, Order> orders =
builder.stream("orders");
Each record represents an event.
73. What is a KTable?
Answer:
A KTable represents a changelog stream interpreted as the latest state for each key.
For example:
customer-1 → ACTIVE
customer-1 → SUSPENDED
The current state becomes:
customer-1 → SUSPENDED
74. Difference between KStream and KTable?
Answer:
| KStream | KTable |
|---|---|
| Event stream | Current state |
| Every event matters | Latest value per key |
| Event-oriented | State-oriented |
75. What is a GlobalKTable?
Answer:
A GlobalKTable is replicated to every Kafka Streams instance.
It is useful for small reference datasets that need to be joined with streams without repartitioning the stream.
76. What is a Kafka Streams state store?
Answer:
A state store is local storage used by Kafka Streams to maintain state required for operations such as:
- Aggregations
- Joins
- Windowing
- Deduplication
State stores can be backed by changelog topics for fault recovery.
77. What is a repartition topic?
Answer:
A repartition topic is an internal Kafka topic used when records need to be redistributed according to a new key.
For example:
Original key
↓
selectKey()
↓
Repartition
↓
Join/Aggregation
78. What is a Kafka Streams topology?
Answer:
A topology is the processing graph defining how Kafka Streams consumes, transforms, and produces records.
Example:
Input Topic
↓
Filter
↓
Map
↓
Aggregate
↓
Output Topic
11. Kafka Connect
79. What is Kafka Connect?
Answer:
Kafka Connect is a framework for moving data between Kafka and external systems.
Examples:
Database → Kafka
Kafka → Elasticsearch
Kafka → Data Warehouse
80. What is a source connector?
Answer:
A source connector imports data into Kafka.
Example:
PostgreSQL
↓
Kafka Connect
↓
Kafka
81. What is a sink connector?
Answer:
A sink connector exports data from Kafka.
Example:
Kafka
↓
Kafka Connect
↓
Elasticsearch
82. What is a Kafka Connect worker?
Answer:
A worker is a Kafka Connect process responsible for executing connectors and their tasks.
Connect can run in:
- Standalone mode
- Distributed mode
83. What is the difference between a connector and a task?
Answer:
A connector defines the integration configuration and determines how work is divided.
Tasks perform the actual data movement.
One connector can create multiple tasks for parallelism.
12. Kafka Security
84. How do you secure Kafka?
Answer:
Kafka security commonly includes:
- TLS/SSL encryption
- SASL authentication
- ACL authorization
- Network isolation
- Secret management
- Client authentication
- Encryption in transit
85. What is SASL?
Answer:
SASL provides an authentication framework for Kafka clients and brokers.
Common mechanisms include:
SASL/SCRAM
SASL/GSSAPI
SASL/OAUTHBEARER
86. What are Kafka ACLs?
Answer:
ACLs control which principals can perform operations on Kafka resources.
For example:
Service A → READ orders
Service B → WRITE orders
Service C → READ payments
This follows the principle of least privilege.
87. How does TLS work with Kafka?
Answer:
TLS can encrypt communication between:
- Producer and broker
- Consumer and broker
- Broker and broker
It protects data in transit and can also support certificate-based authentication.
13. Kafka Schema & Serialization
88. Why is schema management important in Kafka?
Answer:
Kafka messages are often consumed by many independent services.
Schema management helps prevent one producer change from unexpectedly breaking consumers.
Typical concerns include:
- Backward compatibility
- Forward compatibility
- Schema evolution
- Data validation
89. What is Avro?
Answer:
Avro is a compact binary serialization format commonly used with Kafka.
It provides:
- Strong schemas
- Compact messages
- Schema evolution
- Language interoperability
90. What is a Schema Registry?
Answer:
A Schema Registry stores and manages schemas used by Kafka applications.
It can enforce compatibility rules when schemas evolve.
Typical flow:
Producer
↓
Schema Registry
↓
Kafka
↓
Consumer
91. JSON vs Avro vs Protobuf in Kafka?
Answer:
| Format | Advantages | Disadvantages |
|---|---|---|
| JSON | Simple, human-readable | Larger payload |
| Avro | Compact, schema evolution | Requires schema tooling |
| Protobuf | Efficient, strongly typed | Schema tooling required |
For large enterprise event platforms, Avro or Protobuf is often preferable to unrestricted JSON.
14. Error Handling & Reliability
92. How do you handle failed Kafka messages?
Answer:
Common patterns include:
Main Topic
↓
Consumer
↓
Processing Failure
↓
Retry Topic
↓
Retry
↓
DLQ
A Dead Letter Topic/Queue can hold messages that repeatedly fail.
93. What is a Dead Letter Topic?
Answer:
A Dead Letter Topic (DLT/DLQ) is a Kafka topic where messages that cannot be successfully processed are redirected.
It helps:
- Prevent poison messages from blocking processing
- Preserve failed events
- Enable later investigation/reprocessing
94. How would you implement retries in Kafka?
Answer:
Possible approaches include:
- Application-level retry
- Retry topics
- Delayed retry topics
- Exponential backoff
- Dead Letter Topics
- Kafka Streams error handlers
For long delays, dedicated retry topics are often preferable to blocking a consumer thread.
95. What is a poison message?
Answer:
A poison message is an event that repeatedly causes processing failures.
Examples:
- Invalid schema
- Corrupt payload
- Unexpected business data
- Unsupported event version
A retry loop can repeatedly process the same message unless failure handling is designed carefully.
15. Advanced System Design Questions
96. How would you design Kafka for millions of events per second?
Answer:
I would consider:
┌── Broker 1
Producers → Kafka ──┼── Broker 2
├── Broker 3
├── ...
└── Broker N
↓
Consumer Groups
Key design considerations:
- Sufficient partition count
- Multiple brokers
- Appropriate replication factor
- SSD/storage throughput
- Network capacity
- Producer batching
- Compression
- Consumer parallelism
- Capacity planning
- Monitoring consumer lag
- Failure-domain-aware replica placement
- Retention/storage planning
I would benchmark the workload rather than selecting partition counts purely from theoretical calculations.
97. How would you design an order-processing system using Kafka?
Answer:
A possible architecture:
┌──────────────┐
│ Order API │
└──────┬───────┘
│
▼
┌──────────────┐
│ orders topic │
└──────┬───────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Payment Group Inventory Group Notification
│ │ │
▼ ▼ ▼
Payment DB Inventory DB Email/SMS
Use orderId as the Kafka key when ordering per order is important.
98. How would you guarantee ordering for customer transactions?
Answer:
Use the customer ID as the message key:
key = customerId
All events for the same customer should then be routed to the same partition.
However, this creates a trade-off:
Ordering is achieved per key, but a hot key can create partition imbalance.
99. How would you migrate a monolith to Kafka-based microservices?
Answer:
I would use an incremental migration rather than rewriting everything.
For example:
┌───────────────┐
│ Monolith │
└───────┬───────┘
│
Domain Events
│
▼
┌─────────┐
│ Kafka │
└────┬────┘
│
┌───────────┼───────────┐
▼ ▼ ▼
Service A Service B Service C
Key considerations:
- Define domain events
- Establish event ownership
- Avoid dual-write inconsistencies
- Consider the transactional outbox pattern
- Introduce consumers incrementally
- Monitor event lag and failures
- Establish schema compatibility rules
100. What is the transactional outbox pattern, and why is it useful with Kafka?
Answer:
The transactional outbox pattern solves the dual-write problem.
Without it:
Database Transaction
+
Kafka Publish
The database operation could succeed while Kafka publishing fails, or vice versa.
With an outbox:
Application
↓
Database Transaction
┌───────────────┐
│ Business Data │
│ Outbox Event │
└───────────────┘
↓
Outbox Publisher / CDC
↓
Kafka
The business data and outbox event are committed atomically in the same database transaction.
A separate process then publishes the outbox event to Kafka.
This pattern is particularly valuable in microservices where database state and Kafka events must remain consistent.
Bonus: 10 Kafka Scenario-Based Interview Questions
These are especially useful for Senior Java / Spring Boot / Microservices interviews.
101. A consumer is constantly rebalancing. What would you investigate?
Check:
max.poll.interval.ms- Long-running processing
session.timeout.ms- Consumer crashes
- Network instability
- GC pauses
- Excessive processing per poll
- Consumer group membership
102. Kafka consumer lag suddenly increases. What could cause it?
Possible causes:
- Consumer slowdown
- Database slowdown
- External API latency
- CPU saturation
- Network problems
- Broker performance issues
- Too few partitions
- Consumer crashes
- Rebalancing
- Producer traffic spike
103. How do you prevent duplicate processing?
Use:
- Idempotent consumers
- Unique event IDs
- Database constraints
- Upserts
- Transactional processing
- Kafka transactions where appropriate
The key principle is:
Design for at-least-once delivery and make business operations idempotent.
104. How do you handle out-of-order events?
Possible approaches:
- Partition by entity key
- Include event timestamps/version numbers
- Maintain entity sequence numbers
- Buffer events temporarily
- Reject stale versions
- Use Kafka Streams stateful processing
105. What happens if one Kafka partition becomes a hot partition?
The affected partition can become a throughput bottleneck.
Investigate:
key distribution
partitioner
producer traffic
consumer processing
Solutions may include:
- Better key distribution
- Composite keys
- Increasing partitions
- Redesigning partitioning strategy
But simply increasing partitions does not fix a poor key that always maps related traffic to one partition.
106. How would you handle a database failure while consuming Kafka?
Avoid committing the Kafka offset until the database operation succeeds.
Conceptually:
Kafka
↓
Consumer
↓
Database
↓ success
Commit Offset
If the database fails:
Kafka
↓
Consumer
↓
Database ❌
↓
Do NOT commit
The record can then be retried.
The database operation should ideally be idempotent.
107. How would you safely process financial transactions with Kafka?
I would focus on:
- Strong delivery semantics
acks=all- Appropriate replication
min.insync.replicas- Idempotent producers
- Transactions where appropriate
- Idempotent consumers
- Event IDs
- Immutable event records
- Audit trails
- Schema governance
- Encryption
- Authentication/authorization
- Monitoring and alerting
I would avoid claiming that Kafka alone guarantees end-to-end exactly-once behavior across external systems.
108. How do you choose the number of Kafka partitions?
Consider:
Required throughput
↓
Consumer parallelism
↓
Expected growth
↓
Ordering requirements
↓
Key distribution
↓
Broker capacity
A useful starting calculation is:
Required partitions ≈
max(
producer throughput requirement / throughput per partition,
consumer throughput requirement / throughput per partition
)
Then validate the result through load testing.
109. When would you choose Kafka over RabbitMQ?
Kafka is generally preferable when you need:
- Very high throughput
- Durable event history
- Event replay
- Multiple independent consumers
- Stream processing
- Large-scale event pipelines
RabbitMQ may be preferable when you need:
- Traditional work queues
- Complex routing
- Request/task-oriented messaging
- Lower-scale messaging patterns
- Broker-managed delivery semantics suited to queue workloads
The choice should be based on workload rather than simply asking which technology is “faster.”
110. What are the most important Kafka production metrics?
Monitor at least:
Broker metrics
- Request latency
- Network throughput
- Disk utilization
- CPU
- Under-replicated partitions
- Offline partitions
- ISR changes
Producer metrics
- Record send rate
- Error rate
- Request latency
- Batch size
- Record retries
Consumer metrics
- Consumer lag
- Records consumed
- Poll latency
- Rebalances
- Processing latency
- Commit failures
A particularly important production metric is:
Consumer lag by consumer group and partition.
Kafka Interview Cheat Sheet
| Concept | Key Point |
|---|---|
| Topic | Logical event stream |
| Partition | Ordered unit of parallelism |
| Offset | Position within a partition |
| Producer | Writes events |
| Consumer | Reads events |
| Consumer Group | Parallel consumption |
| Broker | Stores/serves partitions |
| Replication Factor | Number of partition copies |
| ISR | In-sync replicas |
acks=all | Strong producer acknowledgement |
min.insync.replicas | Minimum ISR required for writes |
| KRaft | Kafka metadata quorum architecture |
| Retention | Controls how long records remain |
| Compaction | Keeps latest value per key |
| Consumer Lag | Producer position minus consumer position |
| KStream | Event stream |
| KTable | Current state by key |
| Kafka Connect | External-system integration |
| Schema Registry | Schema management |
| Transaction | Atomic Kafka writes |
| Idempotency | Safe repeated processing |
| DLT | Failed-message destination |
| Outbox | Solves database/Kafka dual-write problem |