Kubernetes Probes Types: Liveness, Readiness, Startup
Learn Kubernetes probes types, timing, YAML patterns, rollout guidance, and troubleshooting for reliable liveness, readiness, and startup checks.
A Kubernetes workload can be running and still be unable to serve traffic. It can also answer health checks while its main process is deadlocked. Treating those conditions as interchangeable leads to failed rollouts, unnecessary restarts, and noisy day-2 operations across a fleet.
The kubernetes probes types each answer a different operational question: startup determines when a slow-initializing container is ready to be evaluated. Liveness detects whether it should be restarted, and readiness determines whether it should receive traffic. Configure them around the service's actual startup and failure behavior.
That distinction is the foundation for safe probe configuration. The next step is to examine what each probe tells the kubelet, and what happens when its check succeeds or fails.
Build more reliable Kubernetes operations with Plural's fleet management platform.
What Are the Kubernetes Probe Types?
Kubernetes probes are periodic diagnostics performed by the kubelet against a container. The kubelet can execute code inside the container or make a network request, then use the result to decide whether the workload should receive traffic or be restarted. The three Kubernetes probe types are startup, liveness, and readiness. They inspect different stages of a container's operation, so treating them as interchangeable creates avoidable reliability problems. Kubernetes documents the probe mechanisms and their lifecycle behavior in detail.
The distinction matters because a container can be running without being ready, or responsive without making useful progress. Probe configuration is therefore part of pod lifecycle management, not merely a health-check implementation detail.
Startup probes protect initialization
A startup probe verifies that the application inside a container has finished initializing. While the startup probe is still failing, Kubernetes does not execute the liveness or readiness probes. This gives applications with slow or complex startup phases time to initialize without being mistaken for dead or unavailable. Legacy applications and services that load substantial state during startup are common examples.
Startup has a clear operational boundary: once the probe succeeds, the regular probes take over. If startup continues to fail, the kubelet kills the container and applies the Pod's restart policy. That makes the startup check a gate for initialization, not a replacement for the checks needed during steady-state operation.
Liveness probes detect failure that warrants a restart
A liveness probe answers a recovery question: is the process still making progress, or is it stuck in a condition that requires a restart? A deadlocked application may still have a running process, but it cannot serve its intended function. When a liveness probe fails beyond its configured tolerance, the kubelet restarts the container according to the Pod's restart policy. In some failure states, restarting can restore availability without manual intervention.
The check must represent a genuine unrecoverable application condition. A liveness endpoint that fails whenever the process is busy or a dependency is temporarily slow can cause unnecessary restarts. The result can be cascading failure under load, with fewer healthy containers handling more work.
Readiness probes control traffic
A readiness probe answers whether a container can accept traffic now. It can remain unsuccessful while an application completes initialization tasks, such as loading files, establishing connections, or warming a cache. If readiness fails, the EndpointSlice controller removes the Pod's IP address from the EndpointSlices for matching Services, so traffic is not routed to that container.
Readiness remains useful throughout the container's lifecycle. It can temporarily withdraw a Pod during overload or a recoverable fault without restarting the process. Liveness and readiness operate independently, so liveness does not wait for readiness to succeed. Use startup to gate both during initialization, then give each ongoing probe a narrowly defined responsibility.
How Do Liveness, Readiness, and Startup Probes Behave Differently?
The three Kubernetes probe types produce different operational outcomes, so treating them as interchangeable creates avoidable failure modes. The kubelet evaluates probe results and uses them to decide whether a container should continue running, receive traffic, or finish initialization. The distinctions are summarized in the table below.
| Probe type | Primary question | When it runs | What failure changes |
|---|---|---|---|
| Startup | Has the application finished initializing? | During startup, until it succeeds | The kubelet kills the container after repeated failure, then applies the Pod's restart policy. |
| Liveness | Is the running application still making progress? | Periodically after startup, unless gated by a startup probe | The kubelet restarts the container after failures exceed the configured tolerance. |
| Readiness | Can this container accept traffic now? | Periodically throughout the container's lifecycle | The Pod is removed from matching Service EndpointSlices, so traffic stops routing to it while it is unready. |
A startup probe is a gate, not a recurring availability signal. When it is configured, Kubernetes does not execute liveness or readiness probes until startup succeeds. This protects a slow-initializing service from being restarted before it has finished loading files, establishing connections, or completing other required work. If startup never succeeds, the container is killed and restarted according to the Pod's policy. The Kubernetes probe documentation describes this lifecycle behavior in detail.
Liveness and readiness then diverge during normal operation. A failed liveness probe says the process is no longer functioning well enough to recover in place, such as when it is stuck in a deadlock. Restarting can restore availability, but an overly broad liveness check can restart containers during high load and create a cascading failure. Liveness does not wait for readiness, so the two checks must represent separate conditions.
Readiness is less destructive by design. A failed check does not restart the container. Instead, the EndpointSlice controller removes its address from Services that match the Pod. Because readiness continues throughout the lifecycle, it can temporarily drain a Pod during overload or a recoverable dependency fault. Then return it to service when the application is ready again. This behavior is central to pod lifecycle management and to safe rollouts across a Kubernetes fleet.
How Should You Tune Probe Timing and Failure Thresholds?
Probe settings define how quickly Kubernetes reacts, and how much transient instability it tolerates. Tune them from observed application behavior rather than copying a default from another service. A database-backed API, a batch worker, and a stateless HTTP service can have very different startup and recovery profiles.
Build the detection budget first
periodSeconds controls how often the kubelet performs a periodic probe. timeoutSeconds limits how long an individual check can take before it is considered unsuccessful. failureThreshold sets how many consecutive failures are tolerated before the probe reports failure.
Together, these values create a detection budget. As a rough planning model, estimate the upper bound as the initial delay, plus the time allowed for the required failed attempts, including probe intervals and request timeouts. The exact observed timing can vary because scheduling and probe execution are not perfectly instantaneous. Validate the result under realistic load.
For liveness, make that budget long enough to distinguish an unrecoverable condition from a short CPU, network, or dependency delay. Kubernetes documentation warns that liveness checks must be configured carefully. An overly aggressive threshold can restart a container during a minor transient issue, potentially worsening an existing overload instead of recovering from it.
Separate startup protection from steady-state health
initialDelaySeconds can defer the first liveness check when an application needs predictable time before it is ready to be evaluated. For services with variable or complex initialization, a startup probe is usually the clearer boundary. While the startup probe is active, Kubernetes does not execute liveness or readiness probes. This prevents the steady-state checks from interpreting initialization as failure.
Use the startup probe's periodSeconds, timeoutSeconds, and failureThreshold to express the maximum initialization window your service can safely consume. If the startup probe fails beyond its configured tolerance, the kubelet kills the container and applies the Pod's restart policy. Set the budget from measured cold starts, including image initialization, migrations, cache warming, and dependency setup, then leave room for slower but valid launches.
Use success thresholds deliberately
successThreshold controls how many consecutive successes are required before a failed probe returns to a healthy state. It is especially useful for readiness when a service needs confirmation that recovery is stable before receiving traffic. Keep liveness recovery logic conservative: restarting a container should represent a genuine inability to make progress, not merely temporary unavailability. Readiness can continue evaluating throughout the lifecycle, so it is the right mechanism for removing a recovering or overloaded Pod from service traffic.
Record these assumptions in GitOps manifests and test them during slow starts, overload, dependency loss, and recovery. Consistent timing policies across clusters make probe behavior easier to reason about during fleet-wide rollouts.
See how Plural helps teams manage reliable Kubernetes operations across their fleet.
Sources: Kubernetes probe documentation.
What YAML Configuration Works for a Slow-Starting Service?
A slow initialization phase should not be confused with an unhealthy running process. A startup probe gives Kubernetes a separate gate for that phase. While it is failing, the kubelet does not execute the container's liveness or readiness probes. Once startup succeeds, the other checks take over their normal roles. If startup never succeeds, the kubelet kills the container and applies the Pod's restart policy.
See the Kubernetes probe documentation for the complete field reference.
apiVersion: apps/v1
kind: Deployment
metadata:
name: reports-api
spec:
replicas: 2
selector:
matchLabels:
app: reports-api
template:
metadata:
labels:
app: reports-api
spec:
containers:
- name: reports-api
image: registry.internal/reports-api:2.4.0
ports:
- name: http
containerPort: 8080
startupProbe:
httpGet:
path: /health/startup
port: http
periodSeconds: 10
failureThreshold: 30
readinessProbe:
httpGet:
path: /health/ready
port: http
periodSeconds: 10
failureThreshold: 3
livenessProbe:
httpGet:
path: /health/live
port: http
periodSeconds: 10
failureThreshold: 3- Make startup represent initialization, not general availability. The
/health/startuphandler should return success only after the service has loaded required configuration and completed its initialization work. The budget in this example is controlled by the probe interval and its failure threshold, so tune both against observed startup behavior rather than copying defaults. This is especially useful for legacy applications or services with expensive, multi-component initialization. - Keep readiness focused on traffic acceptance. The readiness endpoint can account for whether the application is able to serve requests now. If it fails later because of a temporary fault or overload, Kubernetes removes the Pod from matching Service EndpointSlices instead of routing new traffic to it. Readiness continues throughout the container lifecycle, so it can recover when the service becomes usable again.
- Keep liveness narrower than readiness. The liveness endpoint should identify a process that is stuck or deadlocked and can recover through a restart. It should not fail merely because a dependency is temporarily unavailable or the service is busy. Liveness and readiness are independent unless startup gates them, which is why the startup probe prevents premature checks during initialization.
- Use named ports and choose the probe mechanism deliberately. The named
httpport keeps the probes tied to the container port without repeating the numeric value. An HTTP request is appropriate when the service exposes lightweight health endpoints. For a process without HTTP, use a TCP socket or anexeccommand, but keep the check fast and deterministic. Do not use a heavyweight diagnostic operation as a health endpoint.
Store this manifest in your GitOps-based deployment workflow and tune it from rollout evidence. Consistent probe definitions make automated restarts and service rollouts more predictable across clusters.
See how Plural helps teams manage reliable Kubernetes fleets.
Which Probe Anti-Patterns Cause False Positives?
A probe is only useful when its failure means what the platform assumes it means. The most damaging configurations turn a temporary dependency failure, a busy process, or a small deployment mismatch into an apparent application failure. Kubernetes documentation warns that an incorrect liveness probe can restart containers under high load, reduce scalability, and contribute to cascading failures. These risks make probe design an operational decision, not a copy-and-paste exercise. See the Kubernetes probe documentation for the underlying behavior.
Checking dependencies from a liveness endpoint
A liveness check should answer whether the container can make progress, not whether every dependency is healthy. If the endpoint fails whenever a database, queue, or third-party API is slow, Kubernetes may restart an otherwise functioning process. Several replicas can then restart together while the dependency is already degraded. Use readiness for conditions that should temporarily remove a Pod from traffic, and reserve liveness for a condition such as an unrecoverable deadlock. A liveness failure can trigger a restart under the Pod's restart policy. This can improve availability when the process is genuinely stuck. It can amplify an incident when the test is too broad.
Using aggressive timing or heavyweight checks
Short delays, tight timeouts, and low failure thresholds leave little room for CPU contention, garbage collection, cold caches, or network jitter. The result is a false positive during normal load. The check itself can also become part of the problem. A probe that runs expensive queries, performs a full dependency graph check, or allocates substantial memory consumes resources on every interval. Under pressure, slower probes fail more often, prompting restarts that add still more initialization work. Keep probe handlers cheap, bounded, and local. Tune initialDelaySeconds, periodSeconds, timeoutSeconds, and failureThreshold against observed startup and runtime behavior rather than arbitrary defaults.
Testing the wrong port or path
A syntactically valid HTTP probe can still target the wrong named port, URL path, scheme, or listener. A service may be healthy on its application port while the probe reaches an administrative listener that is not exposed in the container. Validate the request from inside the Pod, confirm the process binds to the expected interface, and test the exact path used in the manifest. Treat a consistent failure immediately after a rollout as a configuration discrepancy until events and application logs show otherwise.
Omitting startup protection
Slow initialization is not the same as a deadlock. Without a startup probe, liveness may begin before migrations, cache warming, or connection setup has completed. Kubernetes provides startup protection so liveness and readiness are not executed until the startup probe succeeds. Omitting it can create a restart loop during every rollout, especially for legacy or multi-component services. Conversely, a startup probe that never succeeds still subjects the container to the restart policy, so its check must represent a real initialization milestone.
Track probe failures alongside latency, saturation, restarts, and dependency health. That broader view helps distinguish a bad signal from a real fault. Plural's guide to monitoring pod health provides the cluster-level context needed for consistent day-2 operations.
Manage probe reliability consistently across your Kubernetes fleet with Plural.
How Do You Roll Out and Troubleshoot Probes Across a Kubernetes Fleet?
Probe changes are application changes, even when they look like a small YAML diff. A timeout that works in one cluster can be too strict in another because of workload size, node pressure, network policy, or startup behavior. Treat probe configuration as versioned GitOps-based deployment code, and promote it through the fleet in controlled stages.
Use a staged rollout
Begin with a pull request that changes the probe definition, its endpoint, and its timing parameters together. Review the handler implementation with the manifest. A liveness endpoint should identify an unrecoverable process condition, while readiness should represent whether the instance can safely receive traffic. For a slow-starting service, verify that the startup probe gates the other checks rather than compensating for initialization with an excessively long liveness delay. Kubernetes does not run liveness or readiness probes until startup succeeds when a startup probe is configured. See the Kubernetes probe documentation for the execution model.
Apply the change to a representative, low-risk cluster first. Watch restart counts, probe failure events, request errors, and readiness transitions before expanding to additional environments. Keep the same manifest source and promotion rules across clusters, but allow documented environment-specific values where latency or capacity genuinely differs. Plural's Plural fleet management features support consistent day-2 operations through a self-hosted, agent-based pull architecture. That model is useful when rollout control and health automation must also work in zero-trust or air-gapped environments.
Diagnose the symptom before changing thresholds
Start with the Pod and its recent events. The troubleshoot failing probes guide is a practical reference for reading conditions, events, and container state. Then inspect application logs around the exact failure time. A timeout may indicate a saturated process, a wrong port, a dependency issue, or a handler that performs too much work. Increasing timeoutSeconds without identifying the cause can hide a real capacity problem.
| Observed symptom | What to check | Likely action |
|---|---|---|
| Pod never becomes Ready | Events, path, port, HTTP response, and container logs | Correct the handler or configuration; confirm the service is listening on the expected interface. |
| Restarts during startup | Startup duration and startup probe failures | Use a startup probe with a budget that covers measured initialization time. |
| Ready state drops during load | Handler latency, CPU or memory pressure, and dependency saturation | Keep readiness sensitive to traffic capacity, then address overload rather than weakening liveness. |
| Requests reach unhealthy replicas | Service selectors and EndpointSlices | Confirm readiness failures remove the Pod address from matching EndpointSlices. |
A failed readiness probe does not normally restart the container. It causes the EndpointSlice controller to remove the Pod IP from matching Service EndpointSlices, stopping new traffic while the check continues throughout the container lifecycle. This makes readiness valuable during temporary faults and overload. By contrast, repeated liveness failures can cause the kubelet to restart a container. Under high load, an incorrect liveness check can create cascading failures by restarting instances that would otherwise recover.
Record the diagnosis and the tested adjustment in the same change review. This keeps probe behavior auditable, reduces configuration drift, and gives operators a repeatable response across clusters. Correct health diagnostics are a baseline for reliable rollouts and automated restarts, not a substitute for application and capacity monitoring.
See how Plural helps teams operate reliable Kubernetes fleets.
Frequently Asked Questions
Do I need liveness, readiness, and startup probes on every container?
No. Choose probes based on the failure modes and startup behavior of the workload. Readiness is useful for services that should receive traffic only after they can serve requests. Add liveness when the process can become stuck or deadlocked and a restart is an appropriate recovery action. Add startup when initialization can take long enough that liveness would otherwise restart the container prematurely.
What happens when a Kubernetes probe fails?
A failed readiness probe removes the Pod from the matching Service EndpointSlices, so it stops receiving Service traffic while the condition persists. A failed liveness or startup probe can cause the kubelet to restart the container according to the Pod's restart policy. These outcomes make probe semantics operationally significant, not merely diagnostic. Source: Kubernetes documentation.
How should I configure probes for a slow-starting application?
Use a startup probe to cover the application's initialization window, then define liveness and readiness probes for steady-state behavior. Kubernetes waits for the startup probe to succeed before running the liveness and readiness probes. Set the startup period, timeout, and failure threshold from observed startup behavior, including image pulls, migrations, cache warming, and dependency initialization.
Does a readiness probe restart a Kubernetes container?
No. Readiness controls whether a Pod receives traffic; it does not restart the container. A Pod can remain running but unready during overload, a temporary fault, or controlled warm-up. Use a liveness probe when the intended response to a confirmed stuck process is a restart, and keep the two checks focused on their separate responsibilities.
Get started with more consistent fleet operations
Well-tuned probes are easier to operate when their manifests, rollout practices, and troubleshooting signals remain consistent across clusters. Plural helps platform teams manage that operational consistency with GitOps-based deployment and an agent-based pull architecture, while keeping the control plane self-hosted.
Get started with Plural for reliable Kubernetes fleet operations.
Newsletter
Join the newsletter to receive the latest updates in your inbox.