Kubernetes Interview Questions for Hiring Managers (With What Good Answers Look Like)
Kubernetes is a technology where knowing the vocabulary and being able to operate the system are very different things. A candidate may understand what a Deployment is and still struggle to explain why a release is unavailable. If the job includes production responsibility, that second question deserves substantial attention.
I would start by defining the responsibility. Is the engineer deploying applications into a platform someone else runs, or are they responsible for the cluster, its security, upgrades, and recovery? Those roles overlap, but they should not receive an identical interview simply because both job descriptions mention Kubernetes.
The twenty Kubernetes interview questions below establish the technical model, followed by incident scenarios that give the candidate a chance to use it. The goal is to understand how they turn an ambiguous symptom into evidence and a decision. Someone who has operated your particular cloud service may move faster, but the more durable skill is knowing what to inspect and why.
What to Test in a Kubernetes Interview
Cover the workload model, connectivity, storage, access, and ongoing operation, then ask the candidate to use those concepts in an incident. A role that owns clusters needs deeper evidence about upgrades and recovery than a role deploying applications into an existing platform. Make that distinction explicit before choosing the questions.
Architecture and Workload Questions
- What do the control plane components do?
The API server exposes the Kubernetes API, etcd stores cluster state, the scheduler assigns suitable nodes to unscheduled Pods, and controllers reconcile observed state with desired state. Cloud integrations may also involve a cloud controller manager.
Ask what an API server outage changes. Existing workloads may continue running, while control plane operations and reconciliation are impaired. Do not accept either extreme: that nothing is affected or that every application necessarily stops. The candidate should distinguish application execution from the machinery coordinating it.
- How do Pods, Deployments, and Services differ?
A Pod groups containers that share a networking context and can share volumes. A Deployment manages replicated application instances through ReplicaSets and supports rollout behavior. A Service provides a stable way to reach a set of backends, often selected by labels.
Ask what recreates a workload after a node fails. A bare Pod lacks the controller behavior a Deployment supplies. Then ask what happens when a Service selector matches no Pods. Definitions become useful when the candidate can connect them to the failure you are investigating.
- How does scheduling choose a node?
Scheduling filters unsuitable nodes and scores the eligible ones. Resource requests, affinity, selectors, taints and tolerations, and topology constraints can all affect placement.
Give a scenario with several replicas and a requirement to survive an availability-zone failure. Ask how the candidate spreads them, what happens when one zone lacks capacity, and whether the configuration permits scheduling somewhere else. The answer should expose the availability trade-off rather than simply name anti-affinity.
- When would you use a Deployment, StatefulSet, or DaemonSet?
Deployments commonly suit interchangeable replicas. StatefulSets support stable identity and storage relationships. DaemonSets place a workload on eligible nodes, often for infrastructure agents.
Ask whether a database belongs in the cluster at all. A managed service may be a better fit for the team’s operating capacity. A StatefulSet provides useful mechanics; it does not automatically provide replication correctness, backups, or a capable database operator. The candidate should recognize the work that remains after choosing the controller.
- How do requests limits and autoscaling interact?
Requests inform scheduling and resource allocation. CPU limits can cause throttling, while memory limits may lead to a container being killed when it cannot stay within them. Resource-based autoscaling needs the metrics and request configuration its calculations depend on.
Ask what happens when requests are substantially below normal usage or when CPU limits are too restrictive. Then distinguish scaling application replicas from provisioning nodes. Autoscaling is only useful if it responds to a meaningful signal and can obtain the capacity the workload actually needs.
- How do readiness, liveness, and startup probes differ?
Readiness determines whether the workload should receive ordinary Service traffic. Liveness can trigger a restart when a container appears unhealthy. A startup probe gives a slow-starting application time before the other probes take effect.
Ask about a liveness check that depends on an unavailable database. Restarting every application container may amplify the incident without repairing the database. A good answer connects each probe to a recovery action and examines whether that action can improve the reported condition.
Networking and Storage Questions
- How does a Service reach its backends?
For a typical selector-based Service, matching backends are represented in EndpointSlices, and the cluster’s networking implementation routes traffic to them. kube-proxy is common, while some implementations provide alternatives.
Ask about ClusterIP, NodePort, LoadBalancer, and headless Services in relation to the actual environment. Then ask where the candidate would inspect an empty backend set. They should be able to follow the relationship from labels and selectors through endpoint readiness to the route traffic takes.
- How do Ingress and Gateway API expose applications?
Ingress describes HTTP routing implemented by a controller. Gateway API offers a broader model with separate infrastructure and route resources, including GatewayClass, Gateway, and HTTPRoute. The implementation and supported features matter in both cases.
Ask how a platform team and application teams would share ownership safely. Familiarity with newer APIs is useful, but an engineer maintaining an existing Ingress deployment should be assessed on its operation and a credible migration decision, rather than penalized simply for working in the installed environment.
- How would you investigate a connection failure?
Start by locating the failed layer. Confirm the workload state, examine name resolution, inspect the Service’s selectors and endpoints, and compare direct backend connectivity with Service connectivity. NetworkPolicies, ports, and the networking implementation then become evidence rather than guesses.
There is no universally correct command order for every symptom. Ask the candidate what a result would rule out and what they would inspect next. Methodical reasoning matters more than memorizing a troubleshooting sequence that does not fit the failure.
- How do volumes, claims, and StorageClasses work?
A PersistentVolumeClaim requests storage, a PersistentVolume represents it, and a StorageClass describes provisioning behavior. Ask about access modes, topology, binding, and what happens when a claim is deleted.
The candidate should connect those configuration choices to recovery. A Retain policy is not a backup, and a snapshot has not proven its usefulness until the restore path is understood. Ask what was last restored, which data was recovered, and how the application was returned to service.
Security and Access Questions
- How would you apply least-privilege RBAC?
The candidate should define the actions a user or workload needs and grant them at the appropriate scope. Roles, ClusterRoles, and bindings should be explained in that context. A ClusterRole can be bound within a namespace; it is not always a cluster-wide grant.
Ask about permission to read Secrets, create workloads, or execute inside a Pod. Those capabilities can expose credentials or provide broader access than the verb initially suggests. Convenient cluster-admin access deserves a concrete justification and a plan to reduce it.
- How should workloads authenticate to external services?
Ask about identities with a limited scope and lifetime. Projected service account tokens and cloud workload identity can avoid distributing long-lived credentials, depending on the platform and the service being accessed.
The candidate should explain audience, rotation, and which external permissions the identity receives. Mutual TLS can address a different service-to-service requirement. Naming several identity technologies is not enough; they need to explain which boundary each one secures and what happens when credentials expire.
- How would you protect Secrets?
Base64 encoding is not encryption. Upstream Kubernetes stores Secrets unencrypted in etcd by default, although a managed offering may configure additional protections. Ask the candidate to establish the actual cluster settings before describing the remaining exposure.
Controls include encryption at rest, limited API access, rotation, protection of manifests, and preventing values from entering logs or debugging output. An external secret manager may support central policy and rotation, but the path by which a secret reaches the workload still matters. Source: Kubernetes Secrets
- Which workload security controls would you apply?
Discuss Pod Security Admission and the privileged, baseline, and restricted standards. Running without root, limiting capabilities, preventing privilege escalation, and applying seccomp can reduce exposure. A read-only filesystem is another useful control where the application supports it.
Ask how exceptions are reviewed and how the cluster verifies image provenance. NetworkPolicy requires a supporting networking implementation, and namespaces do not automatically isolate traffic. Security should be a collection of explicit, checked controls rather than an assumption attached to a namespace label.
Operations Questions
- How would you plan an upgrade?
Ask for an inventory of versions, deprecated APIs, controllers, networking and storage components, and application dependencies. The candidate should check Kubernetes and provider-specific support policies rather than assume a generic sequence covers every cluster.
Testing, capacity for rolling changes, disruption budgets, and a recovery plan all belong in the discussion. Self-managed etcd recovery and managed-cluster replacement involve different responsibilities. Ask what happens if the control plane change succeeds but a critical controller no longer works. Reinstalling manifests alone may not recover application data.
- What does a PodDisruptionBudget protect?
A PDB limits certain voluntary disruptions through the eviction process. It does not prevent a node failure or node-pressure eviction, and application rollout behavior needs its own controls. That boundary is important when an interviewer uses PDBs as shorthand for availability.
Ask about a budget that permits no disruption with a single replica. It can block maintenance without making the workload resilient. The candidate should explain the relationship between replica count, actual readiness, drain behavior, and the application’s ability to tolerate a disruption. Source: Kubernetes disruptions
- How would you monitor cluster and application health?
Separate the application symptom from the infrastructure signals that help explain it. User-facing error rate and latency matter alongside saturation, node conditions, events, logs, and traces. metrics-server supports resource metrics use cases; it is not a complete historical monitoring system.
Ask how an alert reaches someone who can respond and what action it supports. A Pod restart may warrant investigation without warranting a page. The candidate should explain how alerting remains useful rather than simply enumerate every available metric.
- How do configuration updates reach an application?
Environment variables supplied from a ConfigMap or Secret are not refreshed inside an existing process merely because the object changes. Mounted data may update eventually, subject to configuration details, and the application still has to read it. A subPath mount does not receive those automatic updates.
Ask how the candidate coordinates a configuration change with a rollout and verifies that the running application uses it. Versioned objects or checksum annotations can make the relationship clearer. A successful API update is only one step in changing application behavior. Source: Kubernetes ConfigMaps
- How would you manage manifests across environments?
Helm, Kustomize, and other approaches can manage shared configuration and environment differences. GitOps can reconcile the cluster with declared state. Ask how review, secret handling, release promotion, and drift detection work in the chosen approach.
A Git revert can initiate a desired-state rollback, but it does not undo every database migration or restore overwritten data. The candidate should explain the scope of recovery, including what a reconciliation controller might do after an emergency manual change.
- How do namespaces, quotas, and limits support multiple teams?
Namespaces provide a useful boundary for resource organization and access policies. ResourceQuotas constrain specified aggregate resource use, while LimitRanges can establish defaults and constraints for individual resources.
Ask what they do not isolate. Shared nodes, networking, privileged workloads, and common infrastructure may still create cross-team risk. A quota is a useful control, but it is not proof that one tenant can never affect another. The design should match the degree of separation required.
Kubernetes Scenario Questions
A container repeatedly restarts
CrashLoopBackOff describes repeated container failure with a restart delay; it is a symptom rather than a diagnosis. Ask for events, previous logs, termination reasons, configuration changes, and probe behavior. Exit code 137 can indicate SIGKILL, but it does not establish an out-of-memory cause on its own. The engineer should check the recorded reason and supporting evidence before changing limits.
A rollout stalls
Ask whether users are affected, which versions remain ready, and what changed. A rollback may restore service, but its safety depends on compatibility and the release process. The candidate should understand rollout limits and distinguish marking a rollout as stalled from automatically reverting it. In a GitOps environment, recovery must also account for the declared state the controller will reapply.
A Pod remains Pending
Look at Pod conditions and events before adding capacity. Requests, placement constraints, volumes, and admission or quota failures can produce related symptoms at different stages. A quota may prevent Pod creation rather than leave an existing Pod Pending. The important skill is identifying the actual failing step instead of treating every unscheduled workload as a shortage of nodes. Source: Kubernetes resource quotas
A Service times out after a release
Correlate the symptoms with the change, compare old and new versions, and examine readiness, endpoints, policies, ports, downstream calls, and resource behavior. Ask when the engineer would restore a known good version and which evidence they would preserve. Recovery should be informed by the failure, especially when a rollback could conflict with changed data.
A node is under pressure
Determine whether memory, disk, or another resource is involved and identify the consumer. Cordoning can prevent new scheduling while the team responds; it does not remove existing pressure. Cleanup, capacity changes, resource settings, and workload correction depend on the cause. A PDB does not prevent node-pressure eviction, so it cannot serve as the remedy for the underlying resource problem.
Practical Assessment Options
Run a Live Troubleshooting Exercise
Use a disposable cluster with a bounded fault, such as a mismatched Service selector or a broken configuration reference. If that is impractical, supply command output and ask what the candidate would inspect next. Evaluate diagnosis, recovery, security, and communication, recording how each conclusion followed from the evidence.
Score Answers by Seniority
For a mid-level role, look for correct use of the existing platform and a structured diagnosis. A senior role may require independent recovery decisions, prevention, and an account of the risk to applications. A lead or architect role can add shared standards and decisions across teams. Score the level the job needs and record the evidence behind each judgment.
Red Flags That Signal Delivery Risk
I would probe random restarts, unexplained broad privileges, recovery plans without tested restores, and blame that arrives before diagnosis. A single error should invite a follow-up; persistent inability to reason through a responsibility the role requires is more consequential.
How Gun.io Vets Kubernetes Engineers
At Gun.io, a Kubernetes engineer’s fit depends on the operational responsibility and the environment, not just the technology name. The interview should leave you with a clearer account of what the engineer can own and where additional support is necessary. That is more useful than knowing they scored well on a vocabulary test.
Frequently asked questions
What Kubernetes Interview Questions Should I Ask?
Cover workloads, networking, storage, security, and operations, then use incidents to test how the candidate applies that knowledge. Weight the material against the actual job.
Do certifications prove production experience?
Certifications can provide evidence of knowledge and practical skills. They do not establish how someone has handled your environment’s ambiguity, recovery decisions, or operating responsibilities.
How do Kubernetes administrator and DevOps roles differ?
An administrator may own cluster operation, while a DevOps role may emphasize the path from code to production. Many organizations combine those duties, so define responsibilities rather than rely on the title.
What should a senior engineer add beyond cluster basics?
Independent diagnosis, safe recovery decisions, security judgment, and plans for upgrades and ongoing operation. They should explain the effects of their choices on applications and the people maintaining them.
How Do You Assess Kubernetes Skills?
Use comparable questions and a bounded troubleshooting exercise. Record the commands or evidence the candidate selects, how they interpret the results, and whether the recovery plan preserves security and application availability.
Choose Engineers Who Can Own Production Outcomes
The interview should establish how the engineer moves from a symptom to evidence and a safe decision. If you need help hiring through Gun.io, describe the cluster environment and who will own security, upgrades, and recovery. Those details make the discussion about fit much more useful.