All blogs

kubernetes

Waybill API Down: Four Broken Layers

The waybill pods would not even schedule, and behind that waited three more faults: a missing Secret key, a purposeless sidecar probe, and a Service selector matching nothing.

By Ashutosh Patole9 min read
Waybill API Down: Four Broken Layers
kubernetes notes: waybill api down: four broken layers.

Scenario

I asked my local LLM for the hardest live scenario yet. It handed me an API whose pods would not even schedule, with three more faults stacked behind that one.

Just when you think you fixed the problem, another layer reveals itself.

The Incident

Waybill API down in the logistics namespace: deploy/waybill (2 replicas, containers web + sidecar), svc/waybill (ClusterIP :80), pod/tester.

  • waybill pods sit at 0/2 Pending, endpoints are <none>, and curl http://waybill/ from tester fails with exit 7.
  • Fixed = waybill-* pods at 2/2 Running, populated endpoints, and curl http://waybill/ returning waybill-api: OK.
  • Rules: Troubleshoot within the logistics namespace only. No rebuilding images, no touching other namespaces, no deleting and recreating the namespace.

Troubleshooting

This issue felt like peeling an onion. Every time one gate opened, the next one was locked. Let's walk through all four layers.

Layer 1: Pending on a nodeSelector Nobody Matches

I logged into my Lima machine running a local k3s cluster and checked the pods in the logistics namespace:

bash
ashutosh@lima-k8s-lab:~$ k get namespaceNAME              STATUS   AGEdefault           Active   7dkube-node-lease   Active   7dkube-public       Active   7dkube-system       Active   7dlogistics         Active   45mpayments          Active   6d23hprod              Active   7dshipping          Active   6d19hashutosh@lima-k8s-lab:~$ alias k="kubectl -n logistics"ashutosh@lima-k8s-lab:~$ k get podsNAME                       READY   STATUS    RESTARTS   AGEtester                     1/1     Running   0          45mwaybill-755df487d6-6jp2m   0/2     Pending   0          45mwaybill-755df487d6-cbt8k   0/2     Pending   0          45m

Both waybill pods were stuck in Pending. Let's inspect the deployment manifest to see why the scheduler refuses to touch them:

yaml
ashutosh@lima-k8s-lab:~$ k get deploy waybill -o yamlapiVersion: apps/v1kind: Deploymentmetadata:  annotations:    deployment.kubernetes.io/revision: "1"    kubectl.kubernetes.io/last-applied-configuration: |      {"apiVersion":"apps/v1","kind":"Deployment","metadata":{"annotations":{},"labels":{"app":"waybill"},"name":"waybill","namespace":"logistics"},"spec":{"replicas":2,"selector":{"matchLabels":{"app":"waybill"}},"strategy":{"type":"Recreate"},"template":{"metadata":{"labels":{"app":"waybill"}},"spec":{"containers":[{"command":["python3","/app/app.py"],"env":[{"name":"REDIS_PASSWORD","valueFrom":{"secretKeyRef":{"key":"redis-password","name":"waybill-creds"}}}],"image":"python:3.11-slim","livenessProbe":{"httpGet":{"path":"/healthz","port":8080},"initialDelaySeconds":15,"periodSeconds":10,"timeoutSeconds":2},"name":"web","ports":[{"containerPort":8080}],"readinessProbe":{"httpGet":{"path":"/healthz","port":8080},"initialDelaySeconds":5,"periodSeconds":5,"timeoutSeconds":2},"volumeMounts":[{"mountPath":"/app","name":"code"}]},{"command":["sh","-c","while true; do echo sidecar alive; sleep 30; done"],"image":"busybox:1.36","name":"sidecar","readinessProbe":{"exec":{"command":["cat","/sidecar/ready"]},"initialDelaySeconds":5,"periodSeconds":5}}],"nodeSelector":{"accelerator":"nvidia-t4"},"terminationGracePeriodSeconds":5,"volumes":[{"configMap":{"name":"waybill-code"},"name":"code"}]}}}}  creationTimestamp: "2024-09-25T05:49:51Z"  generation: 1  labels:    app: waybill  name: waybill  namespace: logistics  resourceVersion: "8504"  uid: 2db8e362-858c-4822-a213-1b6ce507712cspec:  progressDeadlineSeconds: 600  replicas: 2  revisionHistoryLimit: 10  selector:    matchLabels:      app: waybill  strategy:    type: Recreate  template:    metadata:      labels:        app: waybill    spec:      containers:      - command:        - python3        - /app/app.py        env:        - name: REDIS_PASSWORD          valueFrom:            secretKeyRef:              key: redis-password              name: waybill-creds        image: python:3.11-slim        imagePullPolicy: IfNotPresent        livenessProbe:          failureThreshold: 3          httpGet:            path: /healthz            port: 8080            scheme: HTTP          initialDelaySeconds: 15          periodSeconds: 10          successThreshold: 1          timeoutSeconds: 2        name: web        ports:        - containerPort: 8080          protocol: TCP        readinessProbe:          failureThreshold: 3          httpGet:            path: /healthz            port: 8080            scheme: HTTP          initialDelaySeconds: 5          periodSeconds: 5          successThreshold: 1          timeoutSeconds: 2        resources: {}        terminationMessagePath: /dev/termination-log        terminationMessagePolicy: File        volumeMounts:        - mountPath: /app          name: code      - command:        - sh        - -c        - while true; do echo sidecar alive; sleep 30; done        image: busybox:1.36        imagePullPolicy: IfNotPresent        name: sidecar        readinessProbe:          exec:            command:            - cat            - /sidecar/ready          failureThreshold: 3          initialDelaySeconds: 5          periodSeconds: 5          successThreshold: 1          timeoutSeconds: 1        resources: {}        terminationMessagePath: /dev/termination-log        terminationMessagePolicy: File      dnsPolicy: ClusterFirst      nodeSelector:        accelerator: nvidia-t4      restartPolicy: Always      schedulerName: default-scheduler      securityContext: {}      terminationGracePeriodSeconds: 5      volumes:      - configMap:          defaultMode: 420          name: waybill-code        name: codestatus:  conditions:  - lastTransitionTime: "2024-09-25T05:49:51Z"    lastUpdateTime: "2024-09-25T05:49:51Z"    message: Deployment does not have minimum availability.    reason: MinimumReplicasUnavailable    status: "False"    type: Available  - lastTransitionTime: "2024-09-25T06:36:48Z"    lastUpdateTime: "2024-09-25T06:36:48Z"    message: ReplicaSet "waybill-755df487d6" has timed out progressing.    reason: ProgressDeadlineExceeded    status: "False"    type: Progressing  observedGeneration: 1  replicas: 2  terminatingReplicas: 0  unavailableReplicas: 2  updatedReplicas: 2

In the deployment spec, there is a nodeSelector: accelerator: nvidia-t4.

Let's see if any node has that label:

bash
ashutosh@lima-k8s-lab:~$ k get nodes --show-labelsNAME           STATUS   ROLES           AGE    VERSION        LABELSlima-k8s-lab   Ready    control-plane   7d1h   v1.36.4+k3s1   beta.kubernetes.io/arch=arm64,beta.kubernetes.io/instance-type=k3s,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=lima-k8s-lab,kubernetes.io/os=linux,node-role.kubernetes.io/control-plane=true,node.kubernetes.io/instance-type=k3s

The single node carries no accelerator label, so nothing matches the nodeSelector.

Unlike a bare pod where the pod spec is immutable once created (as we saw in The Pod That Refused to Be Edited), this is a Deployment. We can safely edit the pod template and let the controller roll out new pods.

I edited the deployment to remove the nodeSelector:

bash
ashutosh@lima-k8s-lab:~$ k edit deploy waybilldeployment.apps/waybill editedashutosh@lima-k8s-lab:~$ k get podsNAME                      READY   STATUS                       RESTARTS   AGEtester                    1/1     Running                      0          50mwaybill-bbfbfff84-54rtb   0/2     CreateContainerConfigError   0          2swaybill-bbfbfff84-f6njq   0/2     CreateContainerConfigError   0          2s

The pods were scheduled immediately, but they did not run. They jumped straight into CreateContainerConfigError.

Layer 2: CreateContainerConfigError on a Missing Secret Key

Pods escaped Pending only to get blocked before the container process could even start. Describing one of the pods showed the cause:

bash
ashutosh@lima-k8s-lab:~$ k describe pod waybill-bbfbfff84-f6njqName:             waybill-bbfbfff84-f6njqNamespace:        logisticsPriority:         0Service Account:  defaultNode:             lima-k8s-lab/192.168.5.15Start Time:       Wed, 25 Sep 2024 12:10:05 +0530Labels:           app=waybill                  pod-template-hash=bbfbfff84Annotations:      <none>Status:           PendingIP:               10.42.0.67IPs:  IP:           10.42.0.67Controlled By:  ReplicaSet/waybill-bbfbfff84Containers:  web:    Container ID:    Image:         python:3.11-slim    Image ID:    Port:          8080/TCP    Host Port:     0/TCP    Command:      python3      /app/app.py    State:          Waiting      Reason:       CreateContainerConfigError    Ready:          False    Restart Count:  0    Liveness:       http-get http://:8080/healthz delay=15s timeout=2s period=10s #success=1 #failure=3    Readiness:      http-get http://:8080/healthz delay=5s timeout=2s period=5s #success=1 #failure=3    Environment:      REDIS_PASSWORD:  <set to the key 'redis-password' in secret 'waybill-creds'>  Optional: false    Mounts:      /app from code (rw)      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-6klgt (ro)  sidecar:    Container ID:  containerd://48a2fd8220ba5dc0bb3b4dd76852ecffedda8a76a45559f84d9e9f9259a950e5    Image:         busybox:1.36    Image ID:      docker.io/library/busybox@sha256:73aaf090f3d85aa34ee199857f03fa3a95c8ede2ffd4cc2cdb5b94e566b11662    Port:          <none>    Host Port:     <none>    Command:      sh      -c      while true; do echo sidecar alive; sleep 30; done    State:          Running      Started:      Wed, 25 Sep 2024 12:10:06 +0530    Ready:          False    Restart Count:  0    Readiness:      exec [cat /sidecar/ready] delay=5s timeout=1s period=5s #success=1 #failure=3    Environment:    <none>    Mounts:      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-6klgt (ro)Conditions:  Type                        Status  PodReadyToStartContainers   True  Initialized                 True  Ready                       False  ContainersReady             False  PodScheduled                TrueVolumes:  code:    Type:      ConfigMap (a volume populated by a ConfigMap)    Name:      waybill-code    Optional:  false  kube-api-access-6klgt:    Type:                    Projected (a volume that contains injected data from multiple sources)    TokenExpirationSeconds:  3607    ConfigMapName:           kube-root-ca.crt    Optional:                false    DownwardAPI:             trueQoS Class:                   BestEffortNode-Selectors:              <none>Tolerations:                 node.kubernetes.io/not-ready:NoExecute op=Exists for 300s                             node.kubernetes.io/unreachable:NoExecute op=Exists for 300sEvents:  Type     Reason     Age                  From               Message  ----     ------     ----                 ----               -------  Normal   Scheduled  110s                 default-scheduler  Successfully assigned logistics/waybill-bbfbfff84-f6njq to lima-k8s-lab  Normal   Pulled     109s                 kubelet            spec.containers{sidecar}: Container image "busybox:1.36" already present on machine and can be accessed by the pod  Normal   Created    109s                 kubelet            spec.containers{sidecar}: Container created  Normal   Started    109s                 kubelet            spec.containers{sidecar}: Container started  Warning  Failed     48s (x8 over 109s)   kubelet            spec.containers{web}: Error: couldn't find key redis-password in Secret logistics/waybill-creds  Warning  Unhealthy  38s (x17 over 102s)  kubelet            spec.containers{sidecar}: Readiness probe failed: cat: can't open '/sidecar/ready': No such file or directory  Normal   Pulled     10s (x11 over 109s)  kubelet            spec.containers{web}: Container image "python:3.11-slim" already present on machine and can be accessed by the pod

The events gave it away:

Error: couldn't find key redis-password in Secret logistics/waybill-creds

The deployment specifies an environment variable REDIS_PASSWORD sourced from the waybill-creds Secret, but the key redis-password was missing.

Also notice the warning right below that in the describe output:

Readiness probe failed: cat: can't open '/sidecar/ready': No such file or directory

We will keep that in mind for later. First, let's fix the Secret.

Remembering my earlier lesson from Payments API Down While Pods Are Ready, I used echo -n to avoid adding an unwanted trailing newline into the Secret:

bash
ashutosh@lima-k8s-lab:~$ echo -n "something" | base64c29tZXRoaW5nashutosh@lima-k8s-lab:~$ k edit secret waybill-credssecret/waybill-creds editedashutosh@lima-k8s-lab:~$ k get podsNAME                      READY   STATUS    RESTARTS   AGEtester                    1/1     Running   0          55mwaybill-bbfbfff84-54rtb   1/2     Running   0          5m33swaybill-bbfbfff84-f6njq   1/2     Running   0          5m33s

With the Secret key present, the kubelet successfully created the container config and started the web container.

Both pods transitioned to Running, but they were only 1/2 Ready.

Layer 3: 1/2 Ready on a Purposeless Sidecar Probe

The pods were running, but not ready for traffic.

We already saw the clue in the earlier pod describe:

Readiness probe failed: cat: can't open '/sidecar/ready': No such file or directory

Let's look at the sidecar container definition from the deployment spec:

yaml
- command:  - sh  - -c  - while true; do echo sidecar alive; sleep 30; done  image: busybox:1.36  imagePullPolicy: IfNotPresent  name: sidecar  readinessProbe:    exec:      command:      - cat      - /sidecar/ready    failureThreshold: 3    initialDelaySeconds: 5    periodSeconds: 5    successThreshold: 1    timeoutSeconds: 1

The sidecar is just a continuous loop printing sidecar alive every 30 seconds. No process inside the pod ever touches or writes to /sidecar/ready.

The probe was waiting for a file that was never going to exist.

Instead of hacking a dummy file into the container, the correct approach is to remove the bogus readiness probe from the deployment.

I edited the deployment and removed the probe:

bash
ashutosh@lima-k8s-lab:~$ k edit deploy waybilldeployment.apps/waybill editedashutosh@lima-k8s-lab:~$ k get podsNAME                      READY   STATUS        RESTARTS   AGEtester                    1/1     Running       0          100mwaybill-bbfbfff84-54rtb   1/2     Terminating   0          50mwaybill-bbfbfff84-f6njq   1/2     Terminating   0          50mashutosh@lima-k8s-lab:~$ k get podsNAME                       READY   STATUS    RESTARTS   AGEtester                     1/1     Running   0          100mwaybill-6bf4cfc7b5-mf9vg   2/2     Running   0          12swaybill-6bf4cfc7b5-n8qfh   2/2     Running   0          12s

Now both pods were 2/2 Ready and Running. Everything looked green on the surface.

Layer 4: A Selector Matching Nothing

With healthy pods running, I tried hitting the API from the tester pod:

bash
ashutosh@lima-k8s-lab:~$ k exec tester -- curl http://waybill  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current                                 Dload  Upload   Total   Spent    Left  Speed  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0curl: (7) Failed to connect to waybill port 80 after 0 ms: Could not connect to servercommand terminated with exit code 7 ashutosh@lima-k8s-lab:~$ k get svcNAME      TYPE        CLUSTER-IP     EXTERNAL-IP   PORT(S)   AGEwaybill   ClusterIP   10.43.235.88   <none>        80/TCP    102mashutosh@lima-k8s-lab:~$ k get endpointsliceNAME            ADDRESSTYPE   PORTS     ENDPOINTS   AGEwaybill-cpghd   IPv4          <unset>   <unset>     102m

Connection refused with exit code 7.

The pods were ready, the Service was active, but the EndpointSlice showed <unset> <unset>. No endpoints were registered.

When ready pods do not register into an EndpointSlice, the first place to look is the Service selector. Let's inspect the Service manifest:

yaml
ashutosh@lima-k8s-lab:~$ k get svc waybill -o yamlapiVersion: v1kind: Servicemetadata:  annotations:    kubectl.kubernetes.io/last-applied-configuration: |      {"apiVersion":"v1","kind":"Service","metadata":{"annotations":{},"name":"waybill","namespace":"logistics"},"spec":{"ports":[{"port":80,"protocol":"TCP","targetPort":8080}],"selector":{"app":"waybill-v2"},"type":"ClusterIP"}}  creationTimestamp: "2024-09-25T05:49:51Z"  name: waybill  namespace: logistics  resourceVersion: "8198"  uid: 929d4aec-61ce-4d66-b0a7-4724d8683d68spec:  clusterIP: 10.43.235.88  clusterIPs:  - 10.43.235.88  internalTrafficPolicy: Cluster  ipFamilies:  - IPv4  ipFamilyPolicy: SingleStack  ports:  - port: 80    protocol: TCP    targetPort: 8080  selector:    app: waybill-v2  sessionAffinity: None  type: ClusterIPstatus:  loadBalancer: {}

Aha! The Service was selecting app: waybill-v2, but our deployment pods were labelled app: waybill.

Because no pods matched app=waybill-v2, the endpoints remained empty.

I updated the Service selector to match the pod labels:

bash
ashutosh@lima-k8s-lab:~$ k edit svc waybillservice/waybill editedashutosh@lima-k8s-lab:~$ k get svcNAME      TYPE        CLUSTER-IP     EXTERNAL-IP   PORT(S)   AGEwaybill   ClusterIP   10.43.235.88   <none>        80/TCP    113mashutosh@lima-k8s-lab:~$ k get endpointsliceNAME            ADDRESSTYPE   PORTS   ENDPOINTS               AGEwaybill-cpghd   IPv4          8080    10.42.0.68,10.42.0.69   113mashutosh@lima-k8s-lab:~$ k exec tester -- curl -s http://waybillwaybill-api: OK

The EndpointSlice immediately populated with our pod IPs on port 8080, and the curl from the tester pod returned waybill-api: OK.

The Learnings

This scenario was a great exercise because each fault lived in a completely different part of the Kubernetes lifecycle:

  1. Scheduling (Kube-Scheduler): When pods linger in Pending, inspect nodeSelector, node taints, and resource limits against the actual node labels.
  2. Container Config (Kubelet): CreateContainerConfigError means the pod was scheduled, but the kubelet could not construct the container environment, usually due to a missing Secret or ConfigMap key.
  3. Readiness Probes (Kubelet): Never ignore a pod stuck at 1/2 Ready. Inspect the probe configuration. If a probe checks for an arbitrary file that nothing creates, it is dead weight.
  4. Service Routing (Endpoints / Kube-Proxy): Green pods mean nothing if your Service selector does not match the pod labels. If curl fails with connection refused, check your EndpointSlices first.

Four layers, but none of them required guesswork. Read the events, check the manifests, and verify the network path end-to-end.

Table of contents

kubernetes

Payments API Down While Pods Are Ready

The checkout pods were 1/1 Ready with populated endpoints, yet the payments API refused connections — then returned HTTP 500 after the first fix. Two layers: a wrong Service targetPort and a Secret with a trailing newline.

- 7 min read - kubernetes, troubleshooting, services

kubernetes

The Pod That Refused to Be Edited

A tale of a Pending Pod, an immutable nodeSelector, and the scheduler's hidden gate.

- 6 min read - kubernetes, troubleshooting, scheduling

kubernetes

Pod is running but not ready

In this article, we will troubleshoot a scenario where the pod is running but not ready. We will explore the possible reasons for this issue and how to resolve it.

- 6 min read - kubernetes, troubleshooting, readiness-probe