100 Most Frequently Asked Kubernetes Interview Questions & Answers
Here’s a practical Kubernetes interview guide, progressing from fundamentals to advanced architecture, troubleshooting, security, networking, and production scenarios.
1. Kubernetes Fundamentals
1. What is Kubernetes?
Answer:
Kubernetes (K8s) is an open-source container orchestration platform originally developed by Google and now maintained by the Cloud Native Computing Foundation (CNCF).
It automates:
- Container deployment
- Scaling
- Service discovery
- Load balancing
- Rolling updates
- Self-healing
- Configuration management
- Secret management
2. Why do we need Kubernetes?
Answer:
Running containers manually becomes difficult when an application has many containers across multiple servers.
Kubernetes provides:
- Automatic scheduling
- Self-healing
- Horizontal scaling
- Service discovery
- Load balancing
- Rolling deployments
- Resource management
- Declarative configuration
3. What is a Kubernetes cluster?
Answer:
A Kubernetes cluster is a collection of machines running Kubernetes.
It consists primarily of:
Kubernetes Cluster
│
├── Control Plane
│ ├── API Server
│ ├── Scheduler
│ ├── Controller Manager
│ └── etcd
│
└── Worker Nodes
├── kubelet
├── kube-proxy
└── Container Runtime
4. What is a Pod?
Answer:
A Pod is the smallest deployable unit in Kubernetes.
A Pod can contain one or more containers that share:
- Network namespace
- IP address
- Ports
- Volumes
Most applications use one main container per Pod.
5. Why doesn’t Kubernetes deploy containers directly?
Answer:
Kubernetes manages Pods, not individual containers.
A Pod provides a shared execution environment for one or more tightly coupled containers.
6. What is a Node?
Answer:
A Node is a worker machine that runs Kubernetes workloads.
A Node typically contains:
- kubelet
- Container runtime
- kube-proxy
- CPU/memory resources
7. What is the Control Plane?
Answer:
The Control Plane manages the Kubernetes cluster.
Major components include:
- kube-apiserver
- etcd
- kube-scheduler
- kube-controller-manager
8. What is kube-apiserver?
Answer:
The Kubernetes API Server is the central entry point to the Kubernetes control plane.
Clients such as:
Bash
kubectl
communicate with the API Server.
It handles:
- Authentication
- Authorization
- Admission
- API requests
- Cluster state operations
9. What is etcd?
Answer:
etcd is a distributed key-value store used by Kubernetes to store cluster state.
It contains information such as:
- Pods
- Deployments
- Services
- ConfigMaps
- Secrets
- Cluster configuration
Important interview point:
Losing etcd without a backup can mean losing the cluster’s persistent control-plane state.
10. What is kube-scheduler?
Answer:
The scheduler determines which Node should run a newly created Pod.
It considers factors such as:
- CPU/memory requirements
- Node availability
- Affinity/anti-affinity
- Taints/tolerations
- Topology constraints
2. Kubernetes Architecture
11. What is kube-controller-manager?
Answer:
It runs Kubernetes controllers that continuously compare the desired state with the actual state.
For example:
Desired:
3 Pods
Actual:
2 Pods
Controller:
Create another Pod
12. What is kubelet?
Answer:
kubelet runs on every worker node.
Its responsibilities include:
- Starting containers
- Monitoring Pods
- Reporting node/Pod status
- Executing Pod lifecycle operations
13. What is kube-proxy?
Answer:
kube-proxy implements Kubernetes Service networking behavior on nodes.
It helps route traffic to the appropriate backend Pods.
Modern Kubernetes installations may use alternatives such as eBPF-based networking implementations.
14. What is a Container Runtime?
Answer:
The container runtime is responsible for running containers.
Kubernetes commonly uses runtimes implementing the Container Runtime Interface (CRI), such as:
- containerd
- CRI-O
15. What is the Kubernetes API?
Answer:
The Kubernetes API is the interface through which users and components interact with cluster resources.
Example:
Bash
kubectl get pods
ultimately communicates with the Kubernetes API Server.
16. What is a Namespace?
Answer:
A Namespace provides logical isolation within a Kubernetes cluster.
For example:
dev
test
staging
production
Namespaces can be used with:
- RBAC
- ResourceQuota
- NetworkPolicy
- ConfigMaps
- Secrets
17. Are all Kubernetes resources namespace-scoped?
Answer:
No.
Some are namespace-scoped:
Pod
Service
Deployment
ConfigMap
Secret
Others are cluster-scoped:
Node
Namespace
PersistentVolume
ClusterRole
ClusterRoleBinding
18. What is declarative configuration?
Answer:
Declarative configuration specifies what the desired state should be, rather than describing every step required to achieve it.
Example:
spec:
replicas: 3
Kubernetes continuously works toward maintaining that state.
19. What is imperative vs declarative management?
Answer:
Imperative:
Bash
kubectl create deployment myapp --image=nginx
Declarative:
Bash
kubectl apply -f deployment.yaml
Declarative management is generally preferred for production because configuration can be version-controlled.
20. What is kubectl?
Answer:
kubectl is the command-line client for interacting with Kubernetes clusters.
Examples:
kubectl get pods
kubectl describe pod mypod
kubectl logs mypod
kubectl apply -f deployment.yaml
kubectl delete pod mypod
3. Pods and Workloads
21. What happens when you create a Pod?
Answer:
A simplified flow is:
kubectl
↓
API Server
↓
etcd
↓
Scheduler
↓
Selected Node
↓
kubelet
↓
Container Runtime
↓
Container
22. Can a Pod contain multiple containers?
Answer:
Yes.
Multiple containers are useful when containers need to share:
- Network
- Storage
- Lifecycle
Example:
Pod
├── Application Container
└── Sidecar Container
23. What is a sidecar container?
Answer:
A sidecar is a secondary container running alongside the main application container.
Typical uses include:
- Log collection
- Proxying
- Metrics
- Security agents
- Configuration synchronization
24. What is an init container?
Answer:
An init container runs before application containers start.
It is commonly used for:
- Initialization
- Database checks
- Configuration preparation
- Dependency preparation
Example:
initContainers:
- name: init-db
image: busybox
25. What happens if a container inside a Pod crashes?
Answer:
The kubelet/container runtime may restart the container according to its restart policy.
For a typical Deployment Pod, Kubernetes attempts to keep the desired workload running.
26. What is a Deployment?
Answer:
A Deployment manages a replicated application workload.
It provides:
- Replica management
- Rolling updates
- Rollbacks
- Self-healing
- Version management
27. Deployment vs Pod?
Answer:
| Pod | Deployment |
|---|---|
| Runs containers | Manages Pods |
| Usually ephemeral | Maintains desired replicas |
| No rollout strategy | Supports rolling updates |
| Doesn’t provide workload history | Supports rollout/rollback |
28. What is a ReplicaSet?
Answer:
A ReplicaSet ensures that a specified number of Pod replicas are running.
Example:
YAML
replicas: 3
A Deployment normally manages ReplicaSets rather than Pods directly.
29. Deployment vs ReplicaSet?
Answer:
A ReplicaSet maintains Pod replicas.
A Deployment provides higher-level functionality:
Deployment
↓
ReplicaSet
↓
Pods
Deployments handle rollout and rollback of ReplicaSets.
30. What is a StatefulSet?
Answer:
A StatefulSet manages stateful applications requiring stable identity or storage.
It provides:
- Stable Pod names
- Stable network identities
- Persistent storage association
- Ordered operations
Example:
mysql-0
mysql-1
mysql-2
31. Deployment vs StatefulSet?
Answer:
Deployment:
Best for stateless applications.
web-xxxxx
web-yyyyy
StatefulSet:
Best for applications requiring stable identities.
db-0
db-1
db-2
32. What is a DaemonSet?
Answer:
A DaemonSet ensures that a Pod runs on every eligible Node.
Typical examples:
- Log collectors
- Monitoring agents
- Node security agents
- Network agents
33. What is a Job?
Answer:
A Job runs a workload until it successfully completes.
Example:
Run database migration
↓
Complete successfully
↓
Stop
34. What is a CronJob?
Answer:
A CronJob creates Jobs according to a schedule.
Example:
YAML
schedule: "0 2 * * *"
This could execute a backup every day at 2 AM.
35. Job vs Deployment?
Answer:
Deployment:
Long-running application
Job:
Run → Complete → Stop
4. Kubernetes Services and Networking
36. Why does Kubernetes need Services?
Answer:
Pods are ephemeral.
Their IP addresses can change when Pods are recreated.
A Service provides a stable network endpoint for a group of Pods.
Client
↓
Service
↓
Pod
Pod
Pod
37. What is a ClusterIP Service?
Answer:
ClusterIP exposes a Service internally within the cluster.
Example:
YAML
type: ClusterIP
It is the default Service type.
38. What is NodePort?
Answer:
NodePort exposes a Service through a port on each Node.
Example:
Client
↓
NodeIP:30080
↓
Service
↓
Pod
39. What is LoadBalancer Service?
Answer:
LoadBalancer exposes a Service through an external load balancer, typically provided by a cloud platform.
Examples:
- AWS
- Azure
- GCP
40. What is a Headless Service?
Answer:
A Headless Service has:
YAML
clusterIP: None
It doesn’t provide a normal virtual ClusterIP.
Instead, DNS can return the IP addresses of individual Pods.
This is especially useful with StatefulSets.
41. What is Kubernetes DNS?
Answer:
Kubernetes normally provides DNS-based service discovery through CoreDNS.
A Service can typically be accessed using a DNS name such as:
my-service.my-namespace.svc.cluster.local
42. What is Ingress?
Answer:
Ingress defines HTTP/HTTPS routing rules for external traffic.
Example:
Internet
↓
Ingress
↓
Service
↓
Pods
It can route based on:
- Host
- Path
43. What is an Ingress Controller?
Answer:
An Ingress resource defines routing rules, while an Ingress Controller implements those rules.
Examples include controllers based on:
- NGINX
- HAProxy
- Traefik
- Cloud provider load balancers
44. Ingress vs Service?
Answer:
Service:
Provides stable networking to Pods.
Ingress:
Provides HTTP/HTTPS routing into Services.
Internet
↓
Ingress
↓
Service
↓
Pods
45. What is NetworkPolicy?
Answer:
NetworkPolicy controls network traffic between Pods and/or other endpoints.
It can define:
- Ingress rules
- Egress rules
Example:
Frontend → Backend
Backend → Database
Frontend ✕ Database
The exact enforcement depends on the cluster’s network plugin.
46. What is a CNI?
Answer:
CNI stands for Container Network Interface.
It defines how networking is configured for containers.
Examples of Kubernetes networking implementations include:
- Cilium
- Calico
- Flannel
47. How do Pods communicate with each other?
Answer:
Kubernetes networking is designed so Pods can generally communicate directly across Nodes without requiring NAT between Pods.
The exact implementation is provided by the cluster’s networking layer/CNI.
48. What is a Service selector?
Answer:
A Service selector determines which Pods receive traffic.
Example:
selector:
app: payment
Pods with:
labels:
app: payment
can become Service endpoints.
49. What happens if a Service has no endpoints?
Answer:
Usually it means no eligible backend Pods are currently selected.
Common causes:
- Incorrect selector
- Pod labels don’t match
- Pods aren’t Ready
- Pods don’t exist
Useful commands:
kubectl get endpoints
kubectl get endpointslices
50. What is an EndpointSlice?
Answer:
EndpointSlice represents network endpoints associated with a Service.
It scales better than the older single Endpoints resource for Services with many endpoints.
5. Configuration and Storage
51. What is a ConfigMap?
Answer:
A ConfigMap stores non-sensitive configuration.
Examples:
Application properties
Environment variables
Configuration files
52. What is a Secret?
Answer:
A Secret is designed to store sensitive configuration such as:
- Passwords
- API keys
- Certificates
- Tokens
However, Kubernetes Secrets are not automatically equivalent to encrypted secret management. Their security depends on cluster configuration, access controls, and encryption-at-rest settings.
53. ConfigMap vs Secret?
Answer:
| ConfigMap | Secret |
|---|---|
| Non-sensitive configuration | Sensitive data |
| Application settings | Passwords/tokens |
| Usually plain configuration | Encoded by default |
Important:
Base64 encoding is not encryption.
54. What is a Volume?
Answer:
A Volume provides storage that can be mounted into a Pod.
Volumes can provide data persistence or shared storage depending on the volume type.
55. What is PersistentVolume?
Answer:
A PersistentVolume (PV) represents storage available to the cluster.
It abstracts the underlying storage implementation.
56. What is PersistentVolumeClaim?
Answer:
A PersistentVolumeClaim (PVC) is a request for storage by a workload.
Pod
↓
PVC
↓
PV
↓
Storage
57. PV vs PVC?
Answer:
PV:
Represents available storage.
PVC:
Represents a workload’s request for storage.
58. What is StorageClass?
Answer:
StorageClass defines how storage can be dynamically provisioned.
Example:
PVC
↓
StorageClass
↓
CSI Driver
↓
Cloud Disk
59. What is CSI?
Answer:
CSI stands for Container Storage Interface.
It provides a standard interface for integrating storage systems with Kubernetes.
60. What is a StatefulSet’s relationship with PVCs?
Answer:
StatefulSets can use volumeClaimTemplates to create persistent storage claims associated with individual Pods.
For example:
db-0 → pvc-db-0
db-1 → pvc-db-1
db-2 → pvc-db-2
6. Scheduling and Resources
61. What are resource requests?
Answer:
A resource request specifies the minimum amount of CPU or memory a container asks Kubernetes to reserve for scheduling purposes.
Example:
resources:
requests:
cpu: "500m"
memory: "512Mi"
62. What are resource limits?
Answer:
Limits specify the maximum resource usage permitted for a container.
Example:
resources:
limits:
cpu: "1"
memory: "1Gi"
63. Requests vs Limits?
Answer:
Request → Used for scheduling
Limit → Maximum allowed resource usage
For CPU, exceeding the limit generally results in throttling.
For memory, excessive usage can result in an OOM kill.
64. What is QoS in Kubernetes?
Answer:
Kubernetes assigns Pods a Quality of Service class.
The major classes are:
- Guaranteed
- Burstable
- BestEffort
QoS affects behavior under resource pressure.
65. What is a Taint?
Answer:
A Taint prevents Pods from being scheduled onto a Node unless they tolerate the taint.
Example:
Node:
dedicated=database:NoSchedule
66. What is a Toleration?
Answer:
A Toleration allows a Pod to be scheduled onto a Node with a matching taint.
Important:
Toleration does not force scheduling onto that Node.
It only permits it.
67. What is Node Selector?
Answer:
nodeSelector restricts a Pod to Nodes with specified labels.
Example:
nodeSelector:
disktype: ssd
68. What is Node Affinity?
Answer:
Node affinity provides more expressive scheduling rules than nodeSelector.
It supports:
- Required rules
- Preferred rules
69. What is Pod Affinity?
Answer:
Pod affinity allows Kubernetes to schedule Pods near other Pods based on labels and topology.
Example:
Application Pod
↓
Database Pod
You may want related workloads to run in the same zone.
70. What is Pod Anti-Affinity?
Answer:
Pod anti-affinity tries to prevent related Pods from being scheduled together.
This improves availability.
Example:
Pod A → Zone 1
Pod B → Zone 2
Pod C → Zone 3
71. What are topology spread constraints?
Answer:
Topology spread constraints help distribute Pods across failure domains such as:
- Nodes
- Availability Zones
- Regions
They are useful for improving workload resilience.
7. Scaling
72. What is Horizontal Pod Autoscaler?
Answer:
HPA automatically adjusts the number of Pod replicas based on metrics.
Example:
CPU > 70%
↓
Increase replicas
73. What is Vertical Pod Autoscaler?
Answer:
VPA adjusts CPU and memory resource requests/limits for workloads based on observed usage.
It complements HPA but should be designed carefully when both operate on the same resource dimensions.
74. What is Cluster Autoscaler?
Answer:
Cluster Autoscaler adjusts the number of Nodes in a cluster based on scheduling demand and cloud-provider capabilities.
Example:
Pods pending
↓
Insufficient capacity
↓
Add Node
75. HPA vs VPA vs Cluster Autoscaler?
Answer:
| Component | Scales |
|---|---|
| HPA | Pod replicas |
| VPA | Pod resource sizing |
| Cluster Autoscaler | Nodes |
76. What is Metrics Server?
Answer:
Metrics Server collects resource usage metrics from Nodes and Pods.
It commonly provides metrics used by HPA and commands such as:
kubectl top pods
kubectl top nodes
8. Deployment Strategies
77. What is a Rolling Update?
Answer:
A Rolling Update gradually replaces old Pods with new Pods.
Example:
v1 v1 v1
↓
v2 v1 v1
↓
v2 v2 v1
↓
v2 v2 v2
This can reduce downtime.
78. What is a Rollback?
Answer:
A rollback restores a previous Deployment revision.
Example:
Bash
kubectl rollout undo deployment myapp
79. What is Blue-Green Deployment?
Answer:
Blue → Current production
Green → New version
After validation, traffic switches from Blue to Green.
Advantages:
- Fast rollback
- Simple traffic switching
Disadvantage:
- Requires additional infrastructure capacity.
80. What is Canary Deployment?
Answer:
A Canary deployment sends a small percentage of traffic to the new version.
Example:
95% → v1
5% → v2
If v2 performs well:
80% → v1
20% → v2
Eventually:
100% → v2
81. Rolling vs Blue-Green vs Canary?
Answer:
| Strategy | Main Benefit |
|---|---|
| Rolling | Simple, resource efficient |
| Blue-Green | Fast rollback |
| Canary | Controlled risk |
9. Probes and Reliability
82. What is a liveness probe?
Answer:
A liveness probe determines whether a container is still functioning.
If it repeatedly fails, Kubernetes can restart the container.
83. What is a readiness probe?
Answer:
Readiness determines whether a Pod is ready to receive traffic.
If readiness fails:
Pod removed from Service endpoints
The container doesn’t necessarily restart.
84. What is a startup probe?
Answer:
A startup probe allows slow-starting applications to initialize before liveness/readiness checks become active.
This prevents Kubernetes from killing an application during a long startup period.
85. Liveness vs Readiness?
Answer:
Liveness:
"Is this container alive?"
Readiness:
"Can this Pod receive traffic?"
This distinction is extremely common in interviews.
86. What is self-healing in Kubernetes?
Answer:
Kubernetes continuously reconciles actual state with desired state.
For example:
Desired Pods = 3
Actual Pods = 2
Controller
↓
Create Pod
10. Security
87. What is RBAC?
Answer:
RBAC stands for Role-Based Access Control.
It controls who can perform which actions on which resources.
Core objects include:
- Role
- ClusterRole
- RoleBinding
- ClusterRoleBinding
88. Role vs ClusterRole?
Answer:
Role:
Namespace-scoped permissions.
ClusterRole:
Cluster-scoped permissions and can also define permissions applicable to namespaced resources when bound appropriately.
89. RoleBinding vs ClusterRoleBinding?
Answer:
RoleBinding:
Grants permissions within a namespace.
ClusterRoleBinding:
Grants permissions across the cluster according to the referenced ClusterRole.
90. What is a ServiceAccount?
Answer:
A ServiceAccount provides an identity for workloads running inside Kubernetes.
Applications can use that identity when communicating with the Kubernetes API or integrated cloud services.
91. What is Pod Security?
Answer:
Kubernetes provides Pod Security Standards with profiles such as:
- Privileged
- Baseline
- Restricted
These help enforce security requirements for workloads.
92. What is a SecurityContext?
Answer:
SecurityContext controls security-related settings for Pods and containers.
Examples:
runAsNonRoot: true
allowPrivilegeEscalation: false
It can also control Linux capabilities, user/group IDs, and filesystem behavior.
11. Troubleshooting
93. A Pod is stuck in Pending. What do you check?
Answer:
Start with:
Bash
kubectl describe pod <pod>
Look for scheduling events.
Common causes:
- Insufficient CPU/memory
- Node affinity
- Taints
- Missing toleration
- PVC issues
- Topology constraints
- Resource quotas
94. A Pod is in CrashLoopBackOff. What do you do?
Answer:
First inspect:
Bash
kubectl logs <pod>
Then:
Bash
kubectl logs <pod> --previous
And:
Bash
kubectl describe pod <pod>
Common causes:
- Application crash
- Invalid configuration
- Missing environment variables
- Failed dependency
- Incorrect command
- Liveness probe failure
95. A Pod is Running but the application is unreachable. What do you check?
Answer:
Check:
kubectl get pods
kubectl get svc
kubectl get endpoints
kubectl get endpointslices
Then investigate:
- Service selector
- Pod labels
- Readiness probe
- Container port
- Service
targetPort - NetworkPolicy
- Ingress
- DNS
96. A Deployment isn’t updating. What do you check?
Answer:
Use:
kubectl rollout status deployment/myapp
kubectl rollout history deployment/myapp
kubectl describe deployment myapp
kubectl get rs
kubectl get pods
Look for:
- Image pull errors
- Failed probes
- Resource constraints
- Scheduling problems
- Deployment strategy limits
- ReplicaSet failures
97. What is ImagePullBackOff?
Answer:
It means Kubernetes cannot successfully pull the container image and is backing off before retrying.
Common causes:
- Incorrect image name
- Invalid tag
- Private registry authentication failure
- Registry unavailable
- Network problems
Useful command:
Bash
kubectl describe pod <pod>
98. How would you troubleshoot high CPU usage?
Answer:
Start with:
kubectl top pods
kubectl top nodes
Then investigate:
Bash
kubectl describe pod <pod>
Check:
- CPU requests/limits
- HPA configuration
- Application behavior
- Traffic volume
- Thread usage
- JVM behavior for Java applications
For a Java/Spring Boot workload, I’d also investigate JVM metrics, GC behavior, thread pools, database calls, and downstream latency.
99. How would you troubleshoot a Kubernetes production outage?
Answer:
I would follow a structured approach:
1. Determine blast radius
2. Check recent deployments
3. Check Pods
4. Check Nodes
5. Check Services
6. Check Ingress/load balancer
7. Check application logs
8. Check metrics/traces
9. Check dependencies
10. Roll back if appropriate
11. Identify root cause
12. Document corrective actions
Useful commands:
kubectl get pods -A
kubectl get nodes
kubectl get events -A
kubectl get deployments -A
kubectl get svc -A
kubectl get ingress -A
The key interview point is to stabilize the system first, then perform root-cause analysis.
12. Advanced Kubernetes / System Design
100. How would you design a production-grade Kubernetes platform?
Answer:
A strong production architecture could look like:
Internet
│
▼
Cloud Load Balancer
│
▼
Ingress / Gateway
│
┌────────────┴────────────┐
▼ ▼
Frontend Service API Service
│ │
▼ ▼
Frontend Pods Backend Pods
│
┌───────────────┼───────────────┐
▼ ▼ ▼
Kafka Database Cache
The platform should also include:
Compute
- Multi-node worker pools
- Multiple availability zones
- Cluster autoscaling
- Appropriate resource requests/limits
Networking
- CNI
- NetworkPolicy
- Ingress/Gateway
- Internal Service discovery
- DNS
Security
- RBAC
- Workload identities
- Secrets management
- Pod Security Standards
- Network segmentation
- Image scanning
- Least privilege
Reliability
- Multiple replicas
- PodDisruptionBudgets
- Readiness/liveness/startup probes
- Topology spread constraints
- Multi-zone deployment
- Backup and disaster recovery
Observability
Metrics → Prometheus/Grafana
Logs → Centralized logging
Traces → OpenTelemetry
Deployment
Use GitOps/CI/CD:
Git
↓
CI
↓
Image Build
↓
Security Scan
↓
Container Registry
↓
Deployment/GitOps
↓
Kubernetes
Data protection
The control plane’s critical state, particularly etcd, should be backed up appropriately, while application data should use the storage/database platform’s own backup and disaster-recovery mechanisms.