The Issue
I wanted to have real troubleshooting experience with Kubernetes, so I asked my local LLM to generate a scenario where I would troubleshoot a live environment.
It ran for a couple of mins and then it generated a scenario where I had a pod in a Pending state. As soon as I saw the scenario, my mind went, "Huh, this is a simple one; I will just check the events and see what is going on."
I had seen this before, so I was confident that I would be able to fix it in no time. I checked the events and saw that the pod was in a Pending state because of a nodeSelector mismatch. I thought, "Oh, this is easy; I will just edit the pod and fix the nodeSelector."
But I got humbled quickly.
The Learning
Checking the pod status
I logged into the Lima machine and checked the pod status.
[ashutosh@lima-default ~]$ alias k="kubectl"[ashutosh@lima-default ~]$ k get namespaceNAME STATUS AGEblog-prod Active 4h40mdefault Active 4h42mkube-node-lease Active 4h42mkube-public Active 4h42mkube-system Active 4h42m[ashutosh@lima-default ~]$ k get pods -n blog-prodNAME READY STATUS RESTARTS AGEblog-pending 0/1 Pending 0 3m40sReading the events
And there it was; I then checked its events and saw the following:
[ashutosh@lima-default ~]$ k describe pod blog-pending -n blog-prodName: blog-pendingNamespace: blog-prodPriority: 0Service Account: defaultNode: <none>Labels: app=blogAnnotations: <none>Status: PendingIP:IPs: <none>Containers: blog: Image: nginx:1.27 Port: <none> Host Port: <none> Limits: cpu: 8 memory: 16Gi Requests: cpu: 8 memory: 16Gi Environment: <none> Mounts: /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-9br4c (ro)Conditions: Type Status PodScheduled FalseVolumes: kube-api-access-9br4c: Type: Projected (a volume that contains injected data from multiple sources) TokenExpirationSeconds: 3607 ConfigMapName: kube-root-ca.crt Optional: false DownwardAPI: trueQoS Class: GuaranteedNode-Selectors: disktype=ssd-proTolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s node.kubernetes.io/unreachable:NoExecute op=Exists for 300sEvents: Type Reason Age From Message ---- ------ ---- ---- ------- Warning FailedScheduling 5m28s default-scheduler 0/1 nodes are available: 1 node(s) didn't match Pod's node affinity/selector. no new claims to deallocate, preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling. Warning FailedScheduling 26s default-scheduler 0/1 nodes are available: 1 node(s) didn't match Pod's node affinity/selector. no new claims to deallocate, preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling.Narrowing down the cause
I saw that 0/1 nodes were available, which meant one of a few things was happening:
- None of the resources requested by the pod were available on the node.
- The nodeSelector was not matching the node labels.
- Pod affinity/anti-affinity was not matching the node labels.
I continued troubleshooting and checked the pod manifest file, and I saw that the nodeSelector was set to disktype=ssd-pro. I then checked the node labels and saw that the no node had the label disktype=ssd-pro. So, the nodeSelector was not matching the node labels. I then checked the pod affinity/anti-affinity and saw that there was no affinity/anti-affinity set. So, the pod affinity/anti-affinity was not causing the issue.
[ashutosh@lima-default ~]$ k get nodes --show-labelsNAME STATUS ROLES AGE VERSION LABELSlima-default Ready control-plane 16h v1.36.4+k3s1 beta.kubernetes.io/arch=arm64,beta.kubernetes.io/instance-type=k3s,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=lima-default,kubernetes.io/os=linux,node-role.kubernetes.io/control-plane=true,node.kubernetes.io/instance-type=k3s[ashutosh@lima-default ~]$ k get pod -n blog-prodNAME READY STATUS RESTARTS AGEblog-pending 0/1 Pending 0 11h[ashutosh@lima-default ~]$ k get pod blog-pending -o yaml -n blog-prodapiVersion: v1kind: Podmetadata: annotations: kubectl.kubernetes.io/last-applied-configuration: | {"apiVersion":"v1","kind":"Pod","metadata":{"annotations":{},"labels":{"app":"blog"},"name":"blog-pending","namespace":"blog-prod"},"spec":{"containers":[{"image":"nginx:1.27","name":"blog","resources":{"limits":{"cpu":"8","memory":"16Gi"},"requests":{"cpu":"8","memory":"16Gi"}}}],"nodeSelector":{"disktype":"ssd-pro"}}} creationTimestamp: "2024-08-12T16:02:58Z" generation: 1 labels: app: blog name: blog-pending namespace: blog-prod resourceVersion: "3519" uid: 8645b28b-c492-48ca-aefa-83ec351b04e4spec: containers: - image: nginx:1.27 imagePullPolicy: IfNotPresent name: blog resources: limits: cpu: "8" memory: 16Gi requests: cpu: "8" memory: 16Gi terminationMessagePath: /dev/termination-log terminationMessagePolicy: File volumeMounts: - mountPath: /var/run/secrets/kubernetes.io/serviceaccount name: kube-api-access-9br4c readOnly: true dnsPolicy: ClusterFirst enableServiceLinks: true nodeSelector: disktype: ssd-pro preemptionPolicy: PreemptLowerPriority priority: 0 restartPolicy: Always schedulerName: default-scheduler securityContext: {} serviceAccount: default serviceAccountName: default terminationGracePeriodSeconds: 30 tolerations: - effect: NoExecute key: node.kubernetes.io/not-ready operator: Exists tolerationSeconds: 300 - effect: NoExecute key: node.kubernetes.io/unreachable operator: Exists tolerationSeconds: 300 volumes: - name: kube-api-access-9br4c projected: defaultMode: 420 sources: - serviceAccountToken: expirationSeconds: 3607 path: token - configMap: items: - key: ca.crt path: ca.crt name: kube-root-ca.crt - downwardAPI: items: - fieldRef: apiVersion: v1 fieldPath: metadata.namespace path: namespacestatus: conditions: - lastProbeTime: null lastTransitionTime: "2024-08-12T16:02:58Z" message: '0/1 nodes are available: 1 node(s) didn''t match Pod''s node affinity/selector. no new claims to deallocate, preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling.' observedGeneration: 1 reason: Unschedulable status: "False" type: PodScheduled phase: Pending qosClass: GuaranteedThe edit that failed
I then tried to edit the pod, either removing the nodeSelector or setting it to disktype=ssd, but I got the following error:
# oc edit pod blog-pending -n blog-prod...nodeselector: {}...Error:
pods "blog-pending" was not valid:* spec: Forbidden: pod updates may not change fields other than `spec.containers[*].image`,`spec.initContainers[*].image`,`spec.activeDeadlineSeconds`,`spec.tolerations` (only additions to existing tolerations),`spec.terminationGracePeriodSeconds` (allow it to be set to 1 if it was previously negative).Why nodeSelector is immutable
So, as per the error, only the specified fields can be updated.
But why can't nodeSelector be updated? I then checked the Kubernetes documentation and found that nodeSelector is immutable. So, you cannot update the nodeSelector of a pod once it is created.
The reason is that, once the pod is created, the scheduler has already made decisions and frozen the manifest in etcd.
You can edit the Deployment or StatefulSet, but not the pod directly. So, if you want to change the nodeSelector, you will have to delete the pod and create a new one with the updated nodeSelector.
The real blocker
Even if I had been able to edit the nodeSelector, the pod would still have been in a Pending state because the resources requested by the pod were not available on the node. The pod was requesting 8 CPUs and 16Gi of memory, but the node had only 4 CPUs and 8Gi of memory available.
Recreating the pod
I created a new manifest file with the updated nodeSelector and resource requests, and created a new pod. The new pod was scheduled successfully and was in a Running state.
[ashutosh@lima-default ~]$ k get pod blog-pending -o yaml -n blog-prod > fixed-blog-prod.yaml [ashutosh@lima-default ~]$ vi fixed-blog-prod.yaml [ashutosh@lima-default ~]$ cat fixed-blog-prod.yamlapiVersion: v1kind: Podmetadata: annotations: kubectl.kubernetes.io/last-applied-configuration: | {"apiVersion":"v1","kind":"Pod","metadata":{"annotations":{},"labels":{"app":"blog"},"name":"blog-pending","namespace":"blog-prod"},"spec":{"containers":[{"image":"nginx:1.27","name":"blog","resources":{"limits":{"cpu":"8","memory":"16Gi"},"requests":{"cpu":"8","memory":"16Gi"}}}],"nodeSelector":{"disktype":"ssd-pro"}}} creationTimestamp: "2024-08-12T16:02:58Z" generation: 1 labels: app: blog name: blog-pending namespace: blog-prod resourceVersion: "3519" uid: 8645b28b-c492-48ca-aefa-83ec351b04e4spec: containers: - image: nginx:1.27 imagePullPolicy: IfNotPresent name: blog resources: limits: cpu: "1" memory: 1Gi requests: cpu: "1" memory: 1Gi terminationMessagePath: /dev/termination-log terminationMessagePolicy: File volumeMounts: - mountPath: /var/run/secrets/kubernetes.io/serviceaccount name: kube-api-access-9br4c readOnly: true dnsPolicy: ClusterFirst enableServiceLinks: true preemptionPolicy: PreemptLowerPriority priority: 0 restartPolicy: Always schedulerName: default-scheduler securityContext: {} serviceAccount: default serviceAccountName: default terminationGracePeriodSeconds: 30 tolerations: - effect: NoExecute key: node.kubernetes.io/not-ready operator: Exists tolerationSeconds: 300 - effect: NoExecute key: node.kubernetes.io/unreachable operator: Exists tolerationSeconds: 300 volumes: - name: kube-api-access-9br4c projected: defaultMode: 420 sources: - serviceAccountToken: expirationSeconds: 3607 path: token - configMap: items: - key: ca.crt path: ca.crt name: kube-root-ca.crt - downwardAPI: items: - fieldRef: apiVersion: v1 fieldPath: metadata.namespace path: namespace [ashutosh@lima-default ~]$ k get pod -n blog-prodNAME READY STATUS RESTARTS AGEblog-pending 0/1 Pending 0 11h [ashutosh@lima-default ~]$ k delete pod blog-pending -n blog-prodpod "blog-pending" deleted from blog-prod namespace [ashutosh@lima-default ~]$ k get pod -n blog-prodNo resources found in blog-prod namespace. [ashutosh@lima-default ~]$ k create -f fixed-blog-prod.yamlpod/blog-pending created [ashutosh@lima-default ~]$ k get pod -n blog-prodNAME READY STATUS RESTARTS AGEblog-pending 1/1 Running 0 4sConclusion
The pod was in a Pending state because of two reasons:
- The nodeSelector was immutable, so I could not edit it to match the node labels
- The resources requested by the pod were not available on the node