Platformkubernetescascading-failureload-sheddinghard

Traffic spikes. Pods start dying. Healthy pods, killed by their own cluster.

01Symptom

A Kubernetes service scales fine most of the time. But during traffic spikes, pod restart counts climb sharply, available replica count actually DROPS during the spike, and the outage gets worse before it recovers — the opposite of what autoscaling should do.

02Constraints

  • Liveness probe on /healthz, timeoutSeconds=1, periodSeconds=5, failureThreshold=1
  • The /healthz handler runs on the same request-handling thread pool/event loop as regular traffic
  • HPA scales on CPU
  • No OOM kills or crash logs during these events — kubectl describe shows 'Liveness probe failed' as the restart reason

03Evidence

  • Restart spikes correlate tightly with traffic spikes, not with any crash or OOM event in logs
  • Available healthy replica count drops during the highest-traffic windows, exactly when more capacity is needed
  • In a canary test, raising timeoutSeconds to 5 and failureThreshold to 3 under identical traffic produced zero restarts

→The question

Why would a healthy, merely busy pod get killed, and why does that make the situation worse rather than better?

04Your prediction

01What's actually happening?
02Why does this make the outage worse, not better?
0 / 600 chars