Scenario
I wanted some real troubleshooting experience with Kubernetes, so I asked my local LLM to generate a scenario where I could troubleshoot a live environment.
It ran for a couple of minutes and then generated a scenario where I had a pod in a Running state.
But it was 0/1 status and the frontend was not accessible. I had to troubleshoot the issue and find out what was wrong with the pod.
Troubleshooting
I logged into my Lima machine hosting a k3s cluster and checked the pods in the prod namespace. I found that the storefront pod was in a Running state but not ready.
ashutosh@lima-k8s-lab:~$ k get namespaceNAME STATUS AGEdefault Active 14mkube-node-lease Active 14mkube-public Active 14mkube-system Active 14mprod Active 8m24sashutosh@lima-k8s-lab:~$ k get pods -n prodNAME READY STATUS RESTARTS AGEstorefront-646d9bdcdb-blwjw 0/1 Running 0 7m25sstorefront-646d9bdcdb-vjv87 0/1 Running 0 7m25stester 1/1 Running 0 8m31sashutosh@lima-k8s-lab:~$ k get deploy -n prodNAME READY UP-TO-DATE AVAILABLE AGEstorefront 0/2 2 0 8mashutosh@lima-k8s-lab:~$ k exec -it tester -n prod -- sh~ $ bashsh: bash: not found~ $command terminated with exit code 127ashutosh@lima-k8s-lab:~$ lsThe LLM also gave me a tester pod to test the frontend from outside the pod. I tried to access the frontend from the tester pod but it was not accessible.
ashutosh@lima-k8s-lab:~$ k get deploy storefront -o yaml -n prodapiVersion: apps/v1kind: Deploymentmetadata: annotations: deployment.kubernetes.io/revision: "1" kubectl.kubernetes.io/last-applied-configuration: | {"apiVersion":"apps/v1","kind":"Deployment","metadata":{"annotations":{},"labels":{"app":"storefront"},"name":"storefront","namespace":"prod"},"spec":{"replicas":2,"selector":{"matchLabels":{"app":"storefront"}},"template":{"metadata":{"labels":{"app":"storefront"}},"spec":{"containers":[{"image":"nginx:stable","name":"web","ports":[{"containerPort":80}],"readinessProbe":{"failureThreshold":3,"httpGet":{"path":"/healthz","port":80},"periodSeconds":5}}]}}}} creationTimestamp: "2024-08-16T05:43:23Z" generation: 1 labels: app: storefront name: storefront namespace: prod resourceVersion: "1378" uid: 4229e302-fab3-4ead-ad0e-9ce140972ba0spec: progressDeadlineSeconds: 600 replicas: 2 revisionHistoryLimit: 10 selector: matchLabels: app: storefront strategy: rollingUpdate: maxSurge: 25% maxUnavailable: 25% type: RollingUpdate template: metadata: labels: app: storefront spec: containers: - image: nginx:stable imagePullPolicy: IfNotPresent name: web ports: - containerPort: 80 protocol: TCP readinessProbe: failureThreshold: 3 httpGet: path: /healthz port: 80 scheme: HTTP periodSeconds: 5 successThreshold: 1 timeoutSeconds: 1 resources: {} terminationMessagePath: /dev/termination-log terminationMessagePolicy: File dnsPolicy: ClusterFirst restartPolicy: Always schedulerName: default-scheduler securityContext: {} terminationGracePeriodSeconds: 30status: conditions: - lastTransitionTime: "2024-08-16T05:43:23Z" lastUpdateTime: "2024-08-16T05:43:23Z" message: Deployment does not have minimum availability. reason: MinimumReplicasUnavailable status: "False" type: Available - lastTransitionTime: "2024-08-16T05:53:24Z" lastUpdateTime: "2024-08-16T05:53:24Z" message: ReplicaSet "storefront-646d9bdcdb" has timed out progressing. reason: ProgressDeadlineExceeded status: "False" type: Progressing observedGeneration: 1 replicas: 2 terminatingReplicas: 0 unavailableReplicas: 2 updatedReplicas: 2I went down the rabbit hole of checking the deployment, pods, services, and endpoints. Everything seemed to be in order.
- Labels matched between the Service and the Deployment.
- Service had an endpoint to connect to
ashutosh@lima-k8s-lab:~$ k get svc -ANAMESPACE NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGEdefault kubernetes ClusterIP 10.43.0.1 <none> 443/TCP 15mkube-system kube-dns ClusterIP 10.43.0.10 <none> 53/UDP,53/TCP,9153/TCP 15mkube-system metrics-server ClusterIP 10.43.23.8 <none> 443/TCP 15mkube-system traefik LoadBalancer 10.43.183.213 192.168.5.15 80:30497/TCP,443:31010/TCP 14mprod storefront ClusterIP 10.43.82.109 <none> 80/TCP 8m45s ashutosh@lima-k8s-lab:~$ k get endpointslice -n prodNAME ADDRESSTYPE PORTS ENDPOINTS AGEstorefront-qvwdd IPv4 80 10.42.0.13,10.42.0.14 9m29sashutosh@lima-k8s-lab:~$ k exec -it tester -n prod -- sh~ $ curl http://storefrontcurl: (7) Failed to connect to storefront port 80 after 0 ms: Could not connect to server~ $But I forgot to take a deep breath and understand the problem. The error was obvious but I was too focused on the deployment and service that I missed the point of a pod with running but not ready state.
I stopped for a few seconds, then described the pod and found that the readiness probe was failing. The pod was running but not ready for that reason.
ashutosh@lima-k8s-lab:~$ k describe pod storefront-646d9bdcdb-blwjw -n prodName: storefront-646d9bdcdb-blwjwNamespace: prodPriority: 0Service Account: defaultNode: lima-k8s-lab/192.168.5.15Start Time: Fri, 16 Aug 2024 11:13:23 +0530Labels: app=storefront pod-template-hash=646d9bdcdbAnnotations: <none>Status: RunningIP: 10.42.0.13IPs: IP: 10.42.0.13Controlled By: ReplicaSet/storefront-646d9bdcdbContainers: web: Container ID: containerd://49c047cacf6a52e09cf3bb211db0be8e6b2d2d30d346978515d38f19da0cd322 Image: nginx:stable Image ID: docker.io/library/nginx@sha256:f91bdb7aee4cba26f89b1c5c3aa12742ec3c91c6d70fd7007c6dc797e9676c45 Port: 80/TCP Host Port: 0/TCP State: Running Started: Fri, 16 Aug 2024 11:13:23 +0530 Ready: False Restart Count: 0 Readiness: http-get http://:80/healthz delay=0s timeout=1s period=5s #success=1 #failure=3 Environment: <none> Mounts: /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-qrwxb (ro)Conditions: Type Status PodReadyToStartContainers True Initialized True Ready False ContainersReady False PodScheduled TrueVolumes: kube-api-access-qrwxb: Type: Projected (a volume that contains injected data from multiple sources) TokenExpirationSeconds: 3607 ConfigMapName: kube-root-ca.crt Optional: false DownwardAPI: trueQoS Class: BestEffortNode-Selectors: <none>Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s node.kubernetes.io/unreachable:NoExecute op=Exists for 300sEvents: Type Reason Age From Message ---- ------ ---- ---- ------- Normal Scheduled 10m default-scheduler Successfully assigned prod/storefront-646d9bdcdb-blwjw to lima-k8s-lab Normal Pulled 10m kubelet spec.containers{web}: Container image "nginx:stable" already present on machine and can be accessed by the pod Normal Created 10m kubelet spec.containers{web}: Container created Normal Started 10m kubelet spec.containers{web}: Container started Warning Unhealthy 10m kubelet spec.containers{web}: Readiness probe failed: Get "http://10.42.0.13:80/healthz": dial tcp 10.42.0.13:80: connect: connection refused Warning Unhealthy 10s (x126 over 10m) kubelet spec.containers{web}: Readiness probe failed: HTTP probe failed with statuscode: 404I exec'd into the pod and curled the healthz endpoint, and found it returned 404. The readiness probe was failing because the healthz endpoint doesn't exist in the nginx image.
ashutosh@lima-k8s-lab:~$ k exec -it storefront-646d9bdcdb-blwjw -n prod -- sh# cd /usr/share/nginx/html# ls50x.html index.html# curl localhost<!DOCTYPE html><html><head><title>Welcome to nginx!</title><style>html { color-scheme: light dark; }body { width: 35em; margin: 0 auto;font-family: Tahoma, Verdana, Arial, sans-serif; }</style></head><body><h1>Welcome to nginx!</h1><p>If you see this page, nginx is successfully installed and working.Further configuration is required for the web server, reverse proxy,API gateway, load balancer, content cache, or other features.</p> <p>For online documentation and support please refer to<a href="https://nginx.org/">nginx.org</a>.<br/>To engage with the community please visit<a href="https://community.nginx.org/">community.nginx.org</a>.<br/>For enterprise grade support, professional services, additionalsecurity features and capabilities please refer to<a href="https://f5.com/nginx">f5.com/nginx</a>.</p> <p><em>Thank you for using nginx.</em></p></body></html># curl localhost/healthz<html><head><title>404 Not Found</title></head><body><center><h1>404 Not Found</h1></center><hr><center>nginx/1.30.5</center></body></html>#I edited the deployment and updated the readiness probe to use the correct endpoint / instead of /healthz. After updating the deployment, the pods became ready and I was able to access the frontend from the tester pod.
ashutosh@lima-k8s-lab:~$ k edit deploy storefront -n proddeployment.apps/storefront editedashutosh@lima-k8s-lab:~$ k get pods -n prodNAME READY STATUS RESTARTS AGEstorefront-6875bf8468-pbl66 1/1 Running 0 2sstorefront-6875bf8468-z5lfl 1/1 Running 0 3stester 1/1 Running 0 26mashutosh@lima-k8s-lab:~$ k exec -it tester -n prod -- sh~ $ curl localhostcurl: (7) Failed to connect to localhost port 80 after 0 ms: Could not connect to server~ $ curl http://storefront<!DOCTYPE html><html><head><title>Welcome to nginx!</title><style>html { color-scheme: light dark; }body { width: 35em; margin: 0 auto;font-family: Tahoma, Verdana, Arial, sans-serif; }</style></head><body><h1>Welcome to nginx!</h1><p>If you see this page, nginx is successfully installed and working.Further configuration is required for the web server, reverse proxy,API gateway, load balancer, content cache, or other features.</p> <p>For online documentation and support please refer to<a href="https://nginx.org/">nginx.org</a>.<br/>To engage with the community please visit<a href="https://community.nginx.org/">community.nginx.org</a>.<br/>For enterprise grade support, professional services, additionalsecurity features and capabilities please refer to<a href="https://f5.com/nginx">f5.com/nginx</a>.</p> <p><em>Thank you for using nginx.</em></p></body></html>~ $The Learning
I still get too excited when solving a problem, and I often forget to step back and understand it first. In this case, the pod was running but not ready because the readiness probe was failing. The readiness probe was failing because the healthz endpoint doesn't exist in the nginx image.
If I had spent a few seconds understanding the problem, I would have caught the issue earlier and saved myself a lot of time.