Skip to content
Bill Liao
Go back

100 Kafka interview Questions and Answers

Edit page

100 Most Frequently Asked Apache Kafka Interview Questions & Answers

Below is a comprehensive Kafka interview guide, progressing from fundamentals to advanced architecture, performance, Kafka Streams, security, operations, and system-design questions.


1. Kafka Fundamentals

1. What is Apache Kafka?

Answer:
Apache Kafka is a distributed event-streaming platform designed to handle high-throughput, fault-tolerant, scalable streams of data.

Kafka is commonly used for:

Kafka stores events durably and allows multiple consumers to independently process the same events.


2. What are the main components of Kafka?

Answer:

The core Kafka components are:

Modern Kafka clusters use KRaft for metadata management instead of ZooKeeper.


3. What is a Kafka topic?

Answer:
A topic is a logical stream/category to which producers publish records and from which consumers read records.

For example:

customer-events
payment-events
order-events

A topic is divided into one or more partitions for scalability.


4. What is a Kafka partition?

Answer:
A partition is an ordered, append-only sequence of records.

For example:

orders
 ├── partition-0
 ├── partition-1
 └── partition-2

Partitions allow Kafka to distribute data across brokers and process records in parallel.


5. Why does Kafka use partitions?

Answer:
Partitions provide:

  1. Horizontal scalability
  2. Parallel processing
  3. Higher throughput
  4. Fault tolerance through replication

If a topic has 10 partitions, multiple consumers can process those partitions concurrently.


6. What is a Kafka offset?

Answer:
An offset is a monotonically increasing number identifying a record’s position within a partition.

Example:

Partition 0

Offset   Event
0        Order A
1        Order B
2        Order C
3        Order D

Offsets are maintained per partition.


7. Are Kafka offsets globally unique?

Answer:
No.

Offsets are only unique within a partition.

For example:

Partition 0 → offset 10
Partition 1 → offset 10

Both are valid because offsets belong to different partitions.


8. Does Kafka guarantee message ordering?

Answer:
Kafka guarantees ordering within a partition, not across an entire topic.

If ordering matters for a customer:

key = customerId

Kafka will normally route all events for that key to the same partition, preserving their order.


9. What is a Kafka broker?

Answer:
A broker is a Kafka server responsible for storing partitions and serving producer and consumer requests.

A Kafka cluster normally contains multiple brokers:

Kafka Cluster
 ├── Broker 1
 ├── Broker 2
 └── Broker 3

10. What is a Kafka cluster?

Answer:
A Kafka cluster is a collection of Kafka brokers working together to provide:


2. Producers

11. What is a Kafka producer?

Answer:
A producer is an application that publishes records to Kafka topics.

For example:

producer.send(
    new ProducerRecord<>("orders", orderId, order)
);

The producer determines the destination topic and, directly or indirectly, the partition.


12. How does a Kafka producer choose a partition?

Answer:
If a key is provided, Kafka normally hashes the key to determine the partition.

Conceptually:

partition = hash(key) % numberOfPartitions

Without a key, the producer can distribute records across partitions according to the producer’s partitioning strategy.


13. Why should Kafka messages have keys?

Answer:
Keys are useful when related events must remain ordered.

For example:

customerId = 123

Events:

CustomerCreated
CustomerUpdated
CustomerDeleted

can be routed to the same partition.


14. What is acks in Kafka?

Answer:
acks controls how much acknowledgement the producer requires from Kafka.

Common values:

acks=0
acks=1
acks=all

acks=all provides the strongest durability guarantee when combined with appropriate replication settings.


15. Difference between acks=0, acks=1, and acks=all?

Answer:

SettingMeaningDurability
0No broker acknowledgementLowest
1Leader acknowledgesMedium
allRequired in-sync replicas acknowledgeHighest

For important financial events, acks=all is usually preferred.


16. What is min.insync.replicas?

Answer:
min.insync.replicas specifies the minimum number of replicas that must be in sync for a producer using acks=all to successfully write.

Example:

replication.factor = 3
min.insync.replicas = 2
acks = all

At least two replicas must be in sync.


17. What is producer batching?

Answer:
Kafka producers can accumulate multiple records into batches before sending them.

Important settings include:

batch.size
linger.ms

Larger batches can improve throughput but may increase latency.


18. What does linger.ms do?

Answer:
linger.ms determines how long the producer may wait for additional records before sending a batch.

Example:

properties

linger.ms=5

The producer may wait up to approximately 5 ms to build a larger batch.


19. What does batch.size control?

Answer:
batch.size controls the maximum size of a producer batch per partition.

It does not mean the producer waits until the batch is completely full. linger.ms determines whether it waits for more records.


20. How can you improve Kafka producer throughput?

Answer:

Common approaches:

For example:

compression.type=zstd
linger.ms=5
batch.size=65536
acks=all

The exact values should be benchmarked rather than blindly copied.


3. Consumers

21. What is a Kafka consumer?

Answer:
A consumer reads records from Kafka topics.

Consumers typically belong to consumer groups and track their offsets.


22. What is a consumer group?

Answer:
A consumer group is a collection of consumers cooperating to consume a topic.

Within one consumer group:

Each partition is assigned to at most one consumer at a time.

Example:

Topic: 6 partitions

Consumer Group
 ├── Consumer A → P0, P1
 ├── Consumer B → P2, P3
 └── Consumer C → P4, P5

23. Can two consumers in the same group read the same partition simultaneously?

Answer:
Normally, no.

A partition is assigned to only one consumer within a consumer group.

However, consumers in different groups can independently consume the same partition.


24. Can multiple consumer groups consume the same topic?

Answer:
Yes.

For example:

orders topic
     |
 ┌───┼────────────┐
 ↓   ↓            ↓
Billing Analytics Fraud
Group   Group     Group

Each group maintains its own offsets.


25. What happens if there are more consumers than partitions?

Answer:
Some consumers will remain idle.

For example:

3 partitions
5 consumers

Only three consumers can actively consume partitions at that time.

Therefore:

Maximum parallelism within a consumer group is bounded by the number of partitions.


26. What happens when a consumer crashes?

Answer:
Kafka detects the consumer failure and triggers a rebalance.

Its partitions are reassigned to other consumers in the group.


27. What is consumer rebalancing?

Answer:
Rebalancing is the process of redistributing partitions among consumers in a consumer group.

It can happen when:


28. What is a consumer offset?

Answer:
A consumer offset represents the position a consumer has reached in a partition.

Kafka stores committed consumer-group offsets in the internal topic:

__consumer_offsets

29. What is auto offset commit?

Answer:
With:

properties

enable.auto.commit=true

the Kafka consumer automatically commits offsets periodically.

The default behavior can be convenient, but it may lead to undesirable processing semantics if the application crashes between processing and committing.


30. What is manual offset commit?

Answer:
The application explicitly commits offsets.

For example:

Java

consumer.commitSync();

This provides more control over when an event is considered successfully processed.


4. Delivery Semantics

31. What is at-most-once delivery?

Answer:

Read
 ↓
Commit
 ↓
Process

If the application crashes after committing but before processing, the message may be lost.

So:

At-most-once = no duplicates, but messages may be lost.


32. What is at-least-once delivery?

Answer:

Read
 ↓
Process
 ↓
Commit

If the application crashes after processing but before committing, Kafka may deliver the message again.

Therefore:

At-least-once = messages are not intentionally lost, but duplicates are possible.


33. What is exactly-once processing?

Answer:
Exactly-once semantics aim to ensure that a Kafka processing pipeline does not produce duplicate effects from retries.

Kafka supports exactly-once processing using mechanisms such as:

However, exactly-once is not automatically guaranteed across arbitrary external systems such as a traditional database or REST API.


34. What is idempotency?

Answer:
An operation is idempotent if executing it multiple times produces the same effective result as executing it once.

For example:

SET account_balance = 100

is naturally more idempotent than:

account_balance += 100

Kafka consumers should often be designed to tolerate duplicate delivery.


35. How do you make a Kafka consumer idempotent?

Answer:

Common approaches include:

Example:

event_id = 8f72...

Store processed event IDs so duplicates can be ignored.


5. Replication & Fault Tolerance

36. What is Kafka replication?

Answer:
Replication means Kafka maintains multiple copies of a partition across brokers.

Example:

Partition 0

Broker 1 → Leader
Broker 2 → Follower
Broker 3 → Follower

If the leader fails, another replica can become leader.


37. What is replication factor?

Answer:
Replication factor specifies how many copies of each partition exist.

Example:

replication.factor=3

means three replicas exist.


38. What is a leader replica?

Answer:
The leader handles reads and writes for a partition.

Followers replicate data from the leader.


39. What is a follower replica?

Answer:
A follower replica maintains a copy of a partition and replicates records from the leader.

Followers can become leaders if the current leader fails.


40. What is ISR?

Answer:
ISR means In-Sync Replicas.

It is the set of replicas considered sufficiently caught up with the partition leader.

Example:

Replication Factor = 3

Broker 1 → Leader
Broker 2 → ISR
Broker 3 → ISR

41. What happens if a broker containing a partition leader fails?

Answer:
Kafka elects another eligible replica as the new leader.

Clients refresh metadata and begin communicating with the new leader.


42. What is an unclean leader election?

Answer:
An unclean leader election allows a replica that is not in the ISR to become leader.

This can reduce data-loss protection.

Typically:

properties

unclean.leader.election.enable=false

is preferred when durability is more important than availability.


43. What is rack awareness in Kafka?

Answer:
Rack awareness allows Kafka to distribute replicas across different failure domains.

For example:

Rack A → Broker 1
Rack B → Broker 2
Rack C → Broker 3

This protects against an entire rack or availability zone failing.


6. Kafka Storage

44. How does Kafka store messages?

Answer:
Kafka stores records in append-only log files called log segments.

Conceptually:

Partition
 ├── segment-0001
 ├── segment-0002
 └── segment-0003

This sequential I/O model contributes to Kafka’s high throughput.


45. What is a log segment?

Answer:
A log segment is a physical file containing records for a partition.

Kafka periodically rolls segments based on configuration such as:

segment.bytes
segment.ms

46. What is log retention?

Answer:
Retention determines how long Kafka keeps records.

Common policies include:

retention.ms
retention.bytes

Kafka can therefore retain data for:


47. Does Kafka delete messages after they are consumed?

Answer:
No.

This is a major difference from traditional queues.

Kafka retains records according to the topic’s retention policy, regardless of whether consumers have already read them.


48. Can a consumer reread old Kafka messages?

Answer:
Yes.

A consumer can reset its offset and replay historical records as long as those records are still retained.

For example:

kafka-consumer-groups \
  --reset-offsets \
  --to-earliest

49. What is log compaction?

Answer:
Log compaction retains the latest record for each key.

Example:

key=A value=10
key=B value=20
key=A value=15

After compaction, Kafka can retain:

A → 15
B → 20

Compaction is useful for maintaining the latest state of entities.


50. Difference between retention and compaction?

Answer:

RetentionCompaction
Removes old records based on time/sizeRemoves older records for the same key
Good for event historyGood for latest state
retention.ms / retention.bytescleanup.policy=compact

A topic can also use:

properties

cleanup.policy=compact,delete

7. Kafka Performance

51. Why is Kafka so fast?

Answer:
Kafka achieves high throughput through several design choices:


52. What is zero-copy?

Answer:
Zero-copy allows data to move from the file system/page cache to the network without unnecessary user-space copying.

This reduces CPU and memory overhead.


53. How can Kafka consumer throughput be increased?

Answer:

Possible strategies:


54. What is fetch.min.bytes?

Answer:
It specifies the minimum amount of data the broker should return to a consumer request when possible.

Larger values can improve throughput but potentially increase latency.


55. What is fetch.max.wait.ms?

Answer:
It controls how long the broker may wait to accumulate enough data to satisfy fetch.min.bytes.


56. What is max.poll.records?

Answer:
It limits the number of records returned by one poll() call.

Example:

properties

max.poll.records=500

This can help control the amount of work processed per poll.


57. What is max.poll.interval.ms?

Answer:
It defines the maximum interval between successful calls to poll() before Kafka considers the consumer unable to keep up.

If processing takes too long, the consumer may leave the group and trigger a rebalance.


58. What is consumer lag?

Answer:
Consumer lag measures how far a consumer is behind the latest available records.

Conceptually:

Latest Offset - Consumer Offset = Lag

High lag indicates that consumers are not keeping up with producers.


59. How do you troubleshoot high consumer lag?

Answer:

Check:

  1. Consumer processing latency
  2. Number of partitions
  3. Number of consumers
  4. CPU/memory utilization
  5. Downstream database/API latency
  6. max.poll.interval.ms
  7. Fetch configuration
  8. Broker performance
  9. Network latency
  10. Consumer errors/rebalances

60. Can adding consumers always reduce consumer lag?

Answer:
No.

If the topic has only 3 partitions:

3 partitions
10 consumers

only three consumers can actively process partitions.

To increase parallelism, the topic may need more partitions.


8. Kafka Architecture

61. What is KRaft?

Answer:
KRaft is Kafka’s metadata management architecture based on Kafka’s own Raft-based consensus mechanism.

It replaces the historical dependency on ZooKeeper.

Modern Kafka deployments can therefore operate without ZooKeeper.


62. What was ZooKeeper used for?

Answer:
Older Kafka architectures used ZooKeeper for:

Modern Kafka deployments use KRaft instead.


63. What is a Kafka controller?

Answer:
The controller manages cluster-level metadata operations such as:

In KRaft mode, controller nodes participate in the metadata quorum.


64. What is a Kafka metadata quorum?

Answer:
In KRaft, Kafka metadata is replicated among controller nodes using a Raft-based quorum.

This allows Kafka to manage cluster metadata without ZooKeeper.


65. What is the difference between broker and controller roles?

Answer:

A broker handles:

A controller handles:

Kafka deployments can use combined or separated roles depending on architecture and scale.


9. Kafka Transactions

66. What is a Kafka transaction?

Answer:
A Kafka transaction allows a producer to atomically write multiple records, potentially across multiple partitions and topics.

Example:

Topic A
Topic B
Topic C

can be written as one transactional unit.


67. What is a transactional producer?

Answer:
A producer configured with a unique transactional.id can participate in Kafka transactions.

Example:

enable.idempotence=true
transactional.id=payment-service-1

68. What is transactional.id?

Answer:
transactional.id uniquely identifies a transactional producer.

It allows Kafka to detect producer restarts and support transactional semantics.


69. What is read_committed?

Answer:
A consumer configured with:

properties

isolation.level=read_committed

reads only committed transactional records.

It will not expose aborted transactional records.


70. What is read_uncommitted?

Answer:
It allows consumers to see records regardless of whether their transaction ultimately commits or aborts.

It is the default isolation level.


10. Kafka Streams

71. What is Kafka Streams?

Answer:
Kafka Streams is a Java library for processing Kafka data streams.

It supports:


72. What is a KStream?

Answer:
A KStream represents an unbounded stream of records.

Example:

KStream<String, Order> orders =
    builder.stream("orders");

Each record represents an event.


73. What is a KTable?

Answer:
A KTable represents a changelog stream interpreted as the latest state for each key.

For example:

customer-1 → ACTIVE
customer-1 → SUSPENDED

The current state becomes:

customer-1 → SUSPENDED

74. Difference between KStream and KTable?

Answer:

KStreamKTable
Event streamCurrent state
Every event mattersLatest value per key
Event-orientedState-oriented

75. What is a GlobalKTable?

Answer:
A GlobalKTable is replicated to every Kafka Streams instance.

It is useful for small reference datasets that need to be joined with streams without repartitioning the stream.


76. What is a Kafka Streams state store?

Answer:
A state store is local storage used by Kafka Streams to maintain state required for operations such as:

State stores can be backed by changelog topics for fault recovery.


77. What is a repartition topic?

Answer:
A repartition topic is an internal Kafka topic used when records need to be redistributed according to a new key.

For example:

Original key
     ↓
selectKey()
     ↓
Repartition
     ↓
Join/Aggregation

78. What is a Kafka Streams topology?

Answer:
A topology is the processing graph defining how Kafka Streams consumes, transforms, and produces records.

Example:

Input Topic
    ↓
Filter
    ↓
Map
    ↓
Aggregate
    ↓
Output Topic

11. Kafka Connect

79. What is Kafka Connect?

Answer:
Kafka Connect is a framework for moving data between Kafka and external systems.

Examples:

Database → Kafka
Kafka → Elasticsearch
Kafka → Data Warehouse

80. What is a source connector?

Answer:
A source connector imports data into Kafka.

Example:

PostgreSQL
    ↓
Kafka Connect
    ↓
Kafka

81. What is a sink connector?

Answer:
A sink connector exports data from Kafka.

Example:

Kafka
  ↓
Kafka Connect
  ↓
Elasticsearch

82. What is a Kafka Connect worker?

Answer:
A worker is a Kafka Connect process responsible for executing connectors and their tasks.

Connect can run in:


83. What is the difference between a connector and a task?

Answer:
A connector defines the integration configuration and determines how work is divided.

Tasks perform the actual data movement.

One connector can create multiple tasks for parallelism.


12. Kafka Security

84. How do you secure Kafka?

Answer:
Kafka security commonly includes:


85. What is SASL?

Answer:
SASL provides an authentication framework for Kafka clients and brokers.

Common mechanisms include:

SASL/SCRAM
SASL/GSSAPI
SASL/OAUTHBEARER

86. What are Kafka ACLs?

Answer:
ACLs control which principals can perform operations on Kafka resources.

For example:

Service A → READ orders
Service B → WRITE orders
Service C → READ payments

This follows the principle of least privilege.


87. How does TLS work with Kafka?

Answer:
TLS can encrypt communication between:

It protects data in transit and can also support certificate-based authentication.


13. Kafka Schema & Serialization

88. Why is schema management important in Kafka?

Answer:
Kafka messages are often consumed by many independent services.

Schema management helps prevent one producer change from unexpectedly breaking consumers.

Typical concerns include:


89. What is Avro?

Answer:
Avro is a compact binary serialization format commonly used with Kafka.

It provides:


90. What is a Schema Registry?

Answer:
A Schema Registry stores and manages schemas used by Kafka applications.

It can enforce compatibility rules when schemas evolve.

Typical flow:

Producer
   ↓
Schema Registry
   ↓
Kafka
   ↓
Consumer

91. JSON vs Avro vs Protobuf in Kafka?

Answer:

FormatAdvantagesDisadvantages
JSONSimple, human-readableLarger payload
AvroCompact, schema evolutionRequires schema tooling
ProtobufEfficient, strongly typedSchema tooling required

For large enterprise event platforms, Avro or Protobuf is often preferable to unrestricted JSON.


14. Error Handling & Reliability

92. How do you handle failed Kafka messages?

Answer:
Common patterns include:

Main Topic
    ↓
Consumer
    ↓
Processing Failure
    ↓
Retry Topic
    ↓
Retry
    ↓
DLQ

A Dead Letter Topic/Queue can hold messages that repeatedly fail.


93. What is a Dead Letter Topic?

Answer:
A Dead Letter Topic (DLT/DLQ) is a Kafka topic where messages that cannot be successfully processed are redirected.

It helps:


94. How would you implement retries in Kafka?

Answer:
Possible approaches include:

  1. Application-level retry
  2. Retry topics
  3. Delayed retry topics
  4. Exponential backoff
  5. Dead Letter Topics
  6. Kafka Streams error handlers

For long delays, dedicated retry topics are often preferable to blocking a consumer thread.


95. What is a poison message?

Answer:
A poison message is an event that repeatedly causes processing failures.

Examples:

A retry loop can repeatedly process the same message unless failure handling is designed carefully.


15. Advanced System Design Questions

96. How would you design Kafka for millions of events per second?

Answer:
I would consider:

                    ┌── Broker 1
Producers → Kafka ──┼── Broker 2
                    ├── Broker 3
                    ├── ...
                    └── Broker N
                         ↓
                    Consumer Groups

Key design considerations:

I would benchmark the workload rather than selecting partition counts purely from theoretical calculations.


97. How would you design an order-processing system using Kafka?

Answer:

A possible architecture:

                    ┌──────────────┐
                    │ Order API    │
                    └──────┬───────┘
                           │
                           ▼
                    ┌──────────────┐
                    │ orders topic │
                    └──────┬───────┘
                           │
             ┌─────────────┼─────────────┐
             ▼             ▼             ▼
        Payment Group  Inventory Group  Notification
             │             │             │
             ▼             ▼             ▼
         Payment DB     Inventory DB    Email/SMS

Use orderId as the Kafka key when ordering per order is important.


98. How would you guarantee ordering for customer transactions?

Answer:
Use the customer ID as the message key:

key = customerId

All events for the same customer should then be routed to the same partition.

However, this creates a trade-off:

Ordering is achieved per key, but a hot key can create partition imbalance.


99. How would you migrate a monolith to Kafka-based microservices?

Answer:
I would use an incremental migration rather than rewriting everything.

For example:

              ┌───────────────┐
              │   Monolith    │
              └───────┬───────┘
                      │
                 Domain Events
                      │
                      ▼
                 ┌─────────┐
                 │  Kafka  │
                 └────┬────┘
                      │
          ┌───────────┼───────────┐
          ▼           ▼           ▼
      Service A   Service B   Service C

Key considerations:


100. What is the transactional outbox pattern, and why is it useful with Kafka?

Answer:
The transactional outbox pattern solves the dual-write problem.

Without it:

Database Transaction
       +
Kafka Publish

The database operation could succeed while Kafka publishing fails, or vice versa.

With an outbox:

Application
    ↓
Database Transaction
 ┌───────────────┐
 │ Business Data │
 │ Outbox Event  │
 └───────────────┘
    ↓
Outbox Publisher / CDC
    ↓
Kafka

The business data and outbox event are committed atomically in the same database transaction.

A separate process then publishes the outbox event to Kafka.

This pattern is particularly valuable in microservices where database state and Kafka events must remain consistent.


Bonus: 10 Kafka Scenario-Based Interview Questions

These are especially useful for Senior Java / Spring Boot / Microservices interviews.

101. A consumer is constantly rebalancing. What would you investigate?

Check:


102. Kafka consumer lag suddenly increases. What could cause it?

Possible causes:


103. How do you prevent duplicate processing?

Use:

The key principle is:

Design for at-least-once delivery and make business operations idempotent.


104. How do you handle out-of-order events?

Possible approaches:


105. What happens if one Kafka partition becomes a hot partition?

The affected partition can become a throughput bottleneck.

Investigate:

key distribution
partitioner
producer traffic
consumer processing

Solutions may include:

But simply increasing partitions does not fix a poor key that always maps related traffic to one partition.


106. How would you handle a database failure while consuming Kafka?

Avoid committing the Kafka offset until the database operation succeeds.

Conceptually:

Kafka
 ↓
Consumer
 ↓
Database
 ↓ success
Commit Offset

If the database fails:

Kafka
 ↓
Consumer
 ↓
Database ❌
 ↓
Do NOT commit

The record can then be retried.

The database operation should ideally be idempotent.


107. How would you safely process financial transactions with Kafka?

I would focus on:

I would avoid claiming that Kafka alone guarantees end-to-end exactly-once behavior across external systems.


108. How do you choose the number of Kafka partitions?

Consider:

Required throughput
        ↓
Consumer parallelism
        ↓
Expected growth
        ↓
Ordering requirements
        ↓
Key distribution
        ↓
Broker capacity

A useful starting calculation is:

Required partitions ≈
max(
  producer throughput requirement / throughput per partition,
  consumer throughput requirement / throughput per partition
)

Then validate the result through load testing.


109. When would you choose Kafka over RabbitMQ?

Kafka is generally preferable when you need:

RabbitMQ may be preferable when you need:

The choice should be based on workload rather than simply asking which technology is “faster.”


110. What are the most important Kafka production metrics?

Monitor at least:

Broker metrics

Producer metrics

Consumer metrics

A particularly important production metric is:

Consumer lag by consumer group and partition.


Kafka Interview Cheat Sheet

ConceptKey Point
TopicLogical event stream
PartitionOrdered unit of parallelism
OffsetPosition within a partition
ProducerWrites events
ConsumerReads events
Consumer GroupParallel consumption
BrokerStores/serves partitions
Replication FactorNumber of partition copies
ISRIn-sync replicas
acks=allStrong producer acknowledgement
min.insync.replicasMinimum ISR required for writes
KRaftKafka metadata quorum architecture
RetentionControls how long records remain
CompactionKeeps latest value per key
Consumer LagProducer position minus consumer position
KStreamEvent stream
KTableCurrent state by key
Kafka ConnectExternal-system integration
Schema RegistrySchema management
TransactionAtomic Kafka writes
IdempotencySafe repeated processing
DLTFailed-message destination
OutboxSolves database/Kafka dual-write problem

Edit page
Share this post:

Previous Post
WhatsApp System Design Interview Question and Answer
Next Post
100 Google Cloud Platform interview Questions and Answers