Scenario
For my next round of hands-on Kubernetes practice, I asked my local LLM for a harder live scenario.
It gave me a down payments API with perfectly green, ready pods. Looks easy, I thought probably just the service. It was not.
The Incident
Payments API down in the payments namespace — deploy/checkout (2 replicas), svc/checkout (ClusterIP :80), pod/tester.
- From
tester,curl http://checkout/fails — connection refused, HTTP 000, exit 7 — yetcheckoutpods show 1/1 Ready Running and endpoints are populated. Soget podslooks green. - Second layer: after the first fix,
curl http://checkout/returns HTTP 500payments-api: ERROR internalon/, while/healthzstill returns 200. - Fixed =
curl http://checkout/fromtesterreturns 200 with bodypayments-api: OK, stable across restarts. - Rules:
get/describe/logs/exec/edit/patch/set/delete/rolloutinpaymentsonly. No rebuilding images, no touchingkube-systemorprod, no deleting/recreating thepaymentsnamespace.
Troubleshooting
This issue is somewhat similar to the one we saw in Pod is running but not ready.
But here the checkout pods are running and ready, let's troubleshoot this one.
Layer 1: Connection Refused With Ready Pods
ashutosh@lima-k8s-lab:~$ k get namespaceNAME STATUS AGEdefault Active 84mkube-node-lease Active 84mkube-public Active 84mkube-system Active 84mpayments Active 13mprod Active 78mashutosh@lima-k8s-lab:~$ k get podsNAME READY STATUS RESTARTS AGEcheckout-778f56dbc5-9c58p 1/1 Running 0 12mcheckout-778f56dbc5-j4q4v 1/1 Running 0 12mtester 1/1 Running 0 13mashutosh@lima-k8s-lab:~$ k get svcNAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGEcheckout ClusterIP 10.43.60.94 <none> 80/TCP 13mashutosh@lima-k8s-lab:~$ k exec -it tester -- sh~ $ curl http://localhostcurl: (7) Failed to connect to localhost port 80 after 0 ms: Could not connect to server~ $The Service and pods were running, but the tester pod could not connect to the checkout service. Let's check the endpoints for the checkout service.
ashutosh@lima-k8s-lab:~$ k get endpointslicesNAME ADDRESSTYPE PORTS ENDPOINTS AGEcheckout-qts5d IPv4 9999 10.42.0.22,10.42.0.23 15mAha! The endpoints for the checkout service are not on port 80, but on port 9999. Let's check the service definition.
ashutosh@lima-k8s-lab:~$ k get svc checkout -o yamlapiVersion: v1kind: Servicemetadata: annotations: kubectl.kubernetes.io/last-applied-configuration: | {"apiVersion":"v1","kind":"Service","metadata":{"annotations":{},"name":"checkout","namespace":"payments"},"spec":{"ports":[{"port":80,"protocol":"TCP","targetPort":9999}],"selector":{"app":"checkout"}}} creationTimestamp: "2024-08-20T06:47:34Z" name: checkout namespace: payments resourceVersion: "3060" uid: 13b09816-2599-46fb-ba41-1893f17fba4bspec: clusterIP: 10.43.60.94 clusterIPs: - 10.43.60.94 internalTrafficPolicy: Cluster ipFamilies: - IPv4 ipFamilyPolicy: SingleStack ports: - port: 80 protocol: TCP targetPort: 9999 selector: app: checkout sessionAffinity: None type: ClusterIPstatus: loadBalancer: {}Let's check the targetPort of the checkout pods.
ashutosh@lima-k8s-lab:~$ k get deploy checkout -o yamlapiVersion: apps/v1kind: Deploymentmetadata: annotations: deployment.kubernetes.io/revision: "3" kubectl.kubernetes.io/last-applied-configuration: | {"apiVersion":"apps/v1","kind":"Deployment","metadata":{"annotations":{},"labels":{"app":"checkout"},"name":"checkout","namespace":"payments"},"spec":{"replicas":2,"selector":{"matchLabels":{"app":"checkout"}},"template":{"metadata":{"labels":{"app":"checkout"}},"spec":{"containers":[{"command":["python3","/app/app.py"],"env":[{"name":"API_KEY","valueFrom":{"secretKeyRef":{"key":"API_KEY","name":"app-secret"}}}],"image":"python:3.11-slim","name":"web","ports":[{"containerPort":8080}],"readinessProbe":{"failureThreshold":3,"httpGet":{"path":"/healthz","port":8080},"periodSeconds":5,"timeoutSeconds":2},"volumeMounts":[{"mountPath":"/app","name":"code"}]}],"volumes":[{"configMap":{"name":"app-code"},"name":"code"}]}}}} creationTimestamp: "2024-08-20T06:47:34Z" generation: 3 labels: app: checkout name: checkout namespace: payments resourceVersion: "3172" uid: 09fea314-37ca-4a1a-9e0f-1b144a8f63c9spec: progressDeadlineSeconds: 600 replicas: 2 revisionHistoryLimit: 10 selector: matchLabels: app: checkout strategy: rollingUpdate: maxSurge: 25% maxUnavailable: 25% type: RollingUpdate template: metadata: annotations: kubectl.kubernetes.io/restartedAt: "2024-08-20T12:18:22+05:30" labels: app: checkout spec: containers: - command: - python3 - /app/app.py env: - name: API_KEY valueFrom: secretKeyRef: key: API_KEY name: app-secret image: python:3.11-slim imagePullPolicy: IfNotPresent name: web ports: - containerPort: 8080 protocol: TCP readinessProbe: failureThreshold: 3 httpGet: path: /healthz port: 8080 scheme: HTTP periodSeconds: 5 successThreshold: 1 timeoutSeconds: 2 resources: {} terminationMessagePath: /dev/termination-log terminationMessagePolicy: File volumeMounts: - mountPath: /app name: code dnsPolicy: ClusterFirst restartPolicy: Always schedulerName: default-scheduler securityContext: {} terminationGracePeriodSeconds: 30 volumes: - configMap: defaultMode: 420 name: app-code name: codestatus: availableReplicas: 2 conditions: - lastTransitionTime: "2024-08-20T06:47:45Z" lastUpdateTime: "2024-08-20T06:47:45Z" message: Deployment has minimum availability. reason: MinimumReplicasAvailable status: "True" type: Available - lastTransitionTime: "2024-08-20T06:47:34Z" lastUpdateTime: "2024-08-20T06:48:29Z" message: ReplicaSet "checkout-778f56dbc5" has successfully progressed. reason: NewReplicaSetAvailable status: "True" type: Progressing observedGeneration: 3 readyReplicas: 2 replicas: 2 terminatingReplicas: 0 updatedReplicas: 2So the targetPort of the checkout pods is 8080, but the service is pointing to 9999. Let's fix the service to point to the correct targetPort.
ashutosh@lima-k8s-lab:~$ k edit svc checkoutservice/checkout editedashutosh@lima-k8s-lab:~$ k get svcNAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGEcheckout ClusterIP 10.43.60.94 <none> 80/TCP 19mashutosh@lima-k8s-lab:~$ k get endpointslicesNAME ADDRESSTYPE PORTS ENDPOINTS AGEcheckout-qts5d IPv4 8080 10.42.0.22,10.42.0.23 19mLayer 2: HTTP 500 From a Wrong Secret
Now the API was reachable from the tester pod, but we faced another problem, the API returned HTTP 500. Let's check the logs of the checkout pods.
ashutosh@lima-k8s-lab:~$ k exec -it tester -- sh~ $ curl http://checkoutpayments-api: ERROR internal~ $Upon checking the logs, we found this:
ashutosh@lima-k8s-lab:~$ k logs -f checkout-778f56dbc5-9c58ppayments-api listening on :8080ERROR GET / 500 invalid API_KEY (got len=13 prefix=pay-...)From the deployment definition, we can see that it uses a ConfigMap and a Secret. Let's check the Secret and the ConfigMap.
ashutosh@lima-k8s-lab:~$ k get configmapNAME DATA AGEapp-code 1 22mkube-root-ca.crt 1 22mashutosh@lima-k8s-lab:~$ k get secretsNAME TYPE DATA AGEapp-secret Opaque 1 22mashutosh@lima-k8s-lab:~$ k get configmap app-code -o yamlapiVersion: v1data: app.py: | import os from http.server import BaseHTTPRequestHandler, HTTPServer EXPECTED_API_KEY = "pay-live-7f3a9c2e" PORT = 8080 class Handler(BaseHTTPRequestHandler): def do_GET(self): if self.path == "/healthz": body = b"ok\n" self.send_response(200) self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) return if self.path == "/": actual = os.environ.get("API_KEY", "") if actual == EXPECTED_API_KEY: body = b"payments-api: OK\n" self.send_response(200) self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) print(f"INFO GET / 200 ok", flush=True) else: body = b"payments-api: ERROR internal\n" self.send_response(500) self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) prefix = (actual[:4] + "...") if actual else "(empty)" print(f"ERROR GET / 500 invalid API_KEY (got len={len(actual)} prefix={prefix})", flush=True) return body = b"not found\n" self.send_response(404) self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) def log_message(self, fmt, *args): pass if __name__ == "__main__": print("payments-api listening on :8080", flush=True) HTTPServer(("0.0.0.0", PORT), Handler).serve_forever()kind: ConfigMapmetadata: creationTimestamp: "2024-08-20T06:47:28Z" name: app-code namespace: payments resourceVersion: "2903" uid: a15a3cce-f1dc-4d6c-b3d7-ae073b83ffd1ashutosh@lima-k8s-lab:~$ k get secretNAME TYPE DATA AGEapp-secret Opaque 1 23mashutosh@lima-k8s-lab:~$ k get secret app-secret -o yamlapiVersion: v1data: API_KEY: cGF5LXRlc3QtMDAwMA==kind: Secretmetadata: annotations: kubectl.kubernetes.io/last-applied-configuration: | {"apiVersion":"v1","data":{"API_KEY":"cGF5LXRlc3QtMDAwMA=="},"kind":"Secret","metadata":{"annotations":{},"creationTimestamp":null,"name":"app-secret","namespace":"payments"}} creationTimestamp: "2024-08-20T06:47:28Z" name: app-secret namespace: payments resourceVersion: "3063" uid: 8a6a8d68-d855-4466-8a34-a19e9e9528a8type: OpaqueFrom the ConfigMap, we can see that the expected API_KEY is pay-live-7f3a9c2e, but the Secret has pay-test-0000. Let's fix the Secret to have the correct API_KEY.
ashutosh@lima-k8s-lab:~$ echo "pay-live-7f3a9c2e" | base64cGF5LWxpdmUtN2YzYTljMmUKashutosh@lima-k8s-lab:~$ k edit secret app-secretsecret/app-secret editedashutosh@lima-k8s-lab:~$ k get secret app-secret -o yamlapiVersion: v1data: API_KEY: cGF5LWxpdmUtN2YzYTljMmUKkind: Secretmetadata: annotations: kubectl.kubernetes.io/last-applied-configuration: | {"apiVersion":"v1","data":{"API_KEY":"cGF5LWxpdmUtN2YzYTljMmUK"},"kind":"Secret","metadata":{"annotations":{},"creationTimestamp":null,"name":"app-secret","namespace":"payments"}} creationTimestamp: "2024-08-20T06:47:28Z" name: app-secret namespace: payments resourceVersion: "3829" uid: 8a6a8d68-d855-4466-8a34-a19e9e9528a8type: OpaqueThe Stale Secret: Restart to Pick It Up
I went ahead and checked whether the API was reachable from the tester pod after fixing the Secret.
ashutosh@lima-k8s-lab:~$ k exec -it tester -- sh~ $ curl http://checkoutpayments-api: ERROR internal~ $But the API was still returning HTTP 500, which means the pods were still using the old Secret.
Upon checking the documentation, I found that when a secret is updated, the pods using that secret do not automatically get the new value. We need to restart the checkout pods to pick up the new secret.
Especially since the Secret is exposed as an environment variable, the pod needs to be restarted to get the new value.
As per the rules, we cannot scale the deployment down to 0 and back up, so we will do a rolling restart of the deployment instead.
ashutosh@lima-k8s-lab:~$ k rollout restart deploy checkoutdeployment.apps/checkout restarted ashutosh@lima-k8s-lab:~$ k get podsNAME READY STATUS RESTARTS AGEcheckout-58d7bd497b-4vpsr 1/1 Running 0 6scheckout-58d7bd497b-9kwrl 1/1 Running 0 4scheckout-778f56dbc5-9c58p 1/1 Terminating 0 32mcheckout-778f56dbc5-j4q4v 1/1 Terminating 0 32mtester 1/1 Running 0 33m ashutosh@lima-k8s-lab:~$ k get podsNAME READY STATUS RESTARTS AGEcheckout-58d7bd497b-4vpsr 1/1 Running 0 41scheckout-58d7bd497b-9kwrl 1/1 Running 0 39stester 1/1 Running 0 34mThe Trailing Newline
Still, I had the same issue, turns out there was an additional \n in the Secret value which was causing the issue. I fixed the Secret value and restarted the pods again.
ashutosh@lima-k8s-lab:~$ echo -n "pay-live-7f3a9c2e" | base64cGF5LWxpdmUtN2YzYTljMmU=ashutosh@lima-k8s-lab:~$ k edit secret app-secretsecret/app-secret editedashutosh@lima-k8s-lab:~$ k rollout restart deploy checkoutdeployment.apps/checkout restartedashutosh@lima-k8s-lab:~$ k get podsNAME READY STATUS RESTARTS AGEcheckout-58d7bd497b-7dbt8 1/1 Running 0 8scheckout-58d7bd497b-ds8kh 1/1 Running 0 8stester 1/1 Running 0 3h39mFinally, the API was reachable from the tester pod and returned HTTP 200.
ashutosh@lima-k8s-lab:~$ k exec -it tester -- sh~ $ curl http://checkoutpayments-api: OK~ $The Learnings
Whitespace matters! When you are using Secrets as environment variables, make sure to use echo -n to avoid adding a newline character at the end of the Secret value.
Or better, use printf instead of echo to avoid adding a newline character at the end of the Secret value.