All blogs

kubernetes

The Pod That Refused to Be Edited

A tale of a Pending Pod, an immutable nodeSelector, and the scheduler's hidden gate.

By Ashutosh Patole6 min read
The Pod That Refused to Be Edited
kubernetes notes: the pod that refused to be edited.

The Issue

I wanted to have real troubleshooting experience with Kubernetes, so I asked my local LLM to generate a scenario where I would troubleshoot a live environment.

It ran for a couple of mins and then it generated a scenario where I had a pod in a Pending state. As soon as I saw the scenario, my mind went, "Huh, this is a simple one; I will just check the events and see what is going on."

I had seen this before, so I was confident that I would be able to fix it in no time. I checked the events and saw that the pod was in a Pending state because of a nodeSelector mismatch. I thought, "Oh, this is easy; I will just edit the pod and fix the nodeSelector."

But I got humbled quickly.

The Learning

Checking the pod status

I logged into the Lima machine and checked the pod status.

bash
[ashutosh@lima-default ~]$ alias k="kubectl"[ashutosh@lima-default ~]$ k get namespaceNAME              STATUS   AGEblog-prod         Active   4h40mdefault           Active   4h42mkube-node-lease   Active   4h42mkube-public       Active   4h42mkube-system       Active   4h42m[ashutosh@lima-default ~]$ k get pods -n blog-prodNAME           READY   STATUS    RESTARTS   AGEblog-pending   0/1     Pending   0          3m40s

Reading the events

And there it was; I then checked its events and saw the following:

bash
[ashutosh@lima-default ~]$ k describe pod blog-pending -n blog-prodName:             blog-pendingNamespace:        blog-prodPriority:         0Service Account:  defaultNode:             <none>Labels:           app=blogAnnotations:      <none>Status:           PendingIP:IPs:              <none>Containers:  blog:    Image:      nginx:1.27    Port:       <none>    Host Port:  <none>    Limits:      cpu:     8      memory:  16Gi    Requests:      cpu:        8      memory:     16Gi    Environment:  <none>    Mounts:      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-9br4c (ro)Conditions:  Type           Status  PodScheduled   FalseVolumes:  kube-api-access-9br4c:    Type:                    Projected (a volume that contains injected data from multiple sources)    TokenExpirationSeconds:  3607    ConfigMapName:           kube-root-ca.crt    Optional:                false    DownwardAPI:             trueQoS Class:                   GuaranteedNode-Selectors:              disktype=ssd-proTolerations:                 node.kubernetes.io/not-ready:NoExecute op=Exists for 300s                             node.kubernetes.io/unreachable:NoExecute op=Exists for 300sEvents:  Type     Reason            Age    From               Message  ----     ------            ----   ----               -------  Warning  FailedScheduling  5m28s  default-scheduler  0/1 nodes are available: 1 node(s) didn't match Pod's node affinity/selector. no new claims to deallocate, preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling.  Warning  FailedScheduling  26s    default-scheduler  0/1 nodes are available: 1 node(s) didn't match Pod's node affinity/selector. no new claims to deallocate, preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling.

Narrowing down the cause

I saw that 0/1 nodes were available, which meant one of a few things was happening:

  • None of the resources requested by the pod were available on the node.
  • The nodeSelector was not matching the node labels.
  • Pod affinity/anti-affinity was not matching the node labels.

I continued troubleshooting and checked the pod manifest file, and I saw that the nodeSelector was set to disktype=ssd-pro. I then checked the node labels and saw that the no node had the label disktype=ssd-pro. So, the nodeSelector was not matching the node labels. I then checked the pod affinity/anti-affinity and saw that there was no affinity/anti-affinity set. So, the pod affinity/anti-affinity was not causing the issue.

bash
[ashutosh@lima-default ~]$ k get nodes --show-labelsNAME           STATUS   ROLES           AGE   VERSION        LABELSlima-default   Ready    control-plane   16h   v1.36.4+k3s1   beta.kubernetes.io/arch=arm64,beta.kubernetes.io/instance-type=k3s,beta.kubernetes.io/os=linux,kubernetes.io/arch=arm64,kubernetes.io/hostname=lima-default,kubernetes.io/os=linux,node-role.kubernetes.io/control-plane=true,node.kubernetes.io/instance-type=k3s[ashutosh@lima-default ~]$ k get pod -n blog-prodNAME           READY   STATUS    RESTARTS   AGEblog-pending   0/1     Pending   0          11h[ashutosh@lima-default ~]$ k get pod blog-pending -o yaml -n blog-prodapiVersion: v1kind: Podmetadata:  annotations:    kubectl.kubernetes.io/last-applied-configuration: |      {"apiVersion":"v1","kind":"Pod","metadata":{"annotations":{},"labels":{"app":"blog"},"name":"blog-pending","namespace":"blog-prod"},"spec":{"containers":[{"image":"nginx:1.27","name":"blog","resources":{"limits":{"cpu":"8","memory":"16Gi"},"requests":{"cpu":"8","memory":"16Gi"}}}],"nodeSelector":{"disktype":"ssd-pro"}}}  creationTimestamp: "2024-08-12T16:02:58Z"  generation: 1  labels:    app: blog  name: blog-pending  namespace: blog-prod  resourceVersion: "3519"  uid: 8645b28b-c492-48ca-aefa-83ec351b04e4spec:  containers:  - image: nginx:1.27    imagePullPolicy: IfNotPresent    name: blog    resources:      limits:        cpu: "8"        memory: 16Gi      requests:        cpu: "8"        memory: 16Gi    terminationMessagePath: /dev/termination-log    terminationMessagePolicy: File    volumeMounts:    - mountPath: /var/run/secrets/kubernetes.io/serviceaccount      name: kube-api-access-9br4c      readOnly: true  dnsPolicy: ClusterFirst  enableServiceLinks: true  nodeSelector:    disktype: ssd-pro  preemptionPolicy: PreemptLowerPriority  priority: 0  restartPolicy: Always  schedulerName: default-scheduler  securityContext: {}  serviceAccount: default  serviceAccountName: default  terminationGracePeriodSeconds: 30  tolerations:  - effect: NoExecute    key: node.kubernetes.io/not-ready    operator: Exists    tolerationSeconds: 300  - effect: NoExecute    key: node.kubernetes.io/unreachable    operator: Exists    tolerationSeconds: 300  volumes:  - name: kube-api-access-9br4c    projected:      defaultMode: 420      sources:      - serviceAccountToken:          expirationSeconds: 3607          path: token      - configMap:          items:          - key: ca.crt            path: ca.crt          name: kube-root-ca.crt      - downwardAPI:          items:          - fieldRef:              apiVersion: v1              fieldPath: metadata.namespace            path: namespacestatus:  conditions:  - lastProbeTime: null    lastTransitionTime: "2024-08-12T16:02:58Z"    message: '0/1 nodes are available: 1 node(s) didn''t match Pod''s node affinity/selector.      no new claims to deallocate, preemption: 0/1 nodes are available: 1 Preemption      is not helpful for scheduling.'    observedGeneration: 1    reason: Unschedulable    status: "False"    type: PodScheduled  phase: Pending  qosClass: Guaranteed

The edit that failed

I then tried to edit the pod, either removing the nodeSelector or setting it to disktype=ssd, but I got the following error:

bash
# oc edit pod blog-pending -n blog-prod...nodeselector: {}...

Error:

text
pods "blog-pending" was not valid:* spec: Forbidden: pod updates may not change fields other than `spec.containers[*].image`,`spec.initContainers[*].image`,`spec.activeDeadlineSeconds`,`spec.tolerations` (only additions to existing tolerations),`spec.terminationGracePeriodSeconds` (allow it to be set to 1 if it was previously negative).

Why nodeSelector is immutable

So, as per the error, only the specified fields can be updated.

But why can't nodeSelector be updated? I then checked the Kubernetes documentation and found that nodeSelector is immutable. So, you cannot update the nodeSelector of a pod once it is created.

The reason is that, once the pod is created, the scheduler has already made decisions and frozen the manifest in etcd.

You can edit the Deployment or StatefulSet, but not the pod directly. So, if you want to change the nodeSelector, you will have to delete the pod and create a new one with the updated nodeSelector.

The real blocker

Even if I had been able to edit the nodeSelector, the pod would still have been in a Pending state because the resources requested by the pod were not available on the node. The pod was requesting 8 CPUs and 16Gi of memory, but the node had only 4 CPUs and 8Gi of memory available.

Recreating the pod

I created a new manifest file with the updated nodeSelector and resource requests, and created a new pod. The new pod was scheduled successfully and was in a Running state.

bash
[ashutosh@lima-default ~]$ k get pod blog-pending -o yaml -n blog-prod > fixed-blog-prod.yaml [ashutosh@lima-default ~]$ vi fixed-blog-prod.yaml [ashutosh@lima-default ~]$ cat fixed-blog-prod.yamlapiVersion: v1kind: Podmetadata:  annotations:    kubectl.kubernetes.io/last-applied-configuration: |      {"apiVersion":"v1","kind":"Pod","metadata":{"annotations":{},"labels":{"app":"blog"},"name":"blog-pending","namespace":"blog-prod"},"spec":{"containers":[{"image":"nginx:1.27","name":"blog","resources":{"limits":{"cpu":"8","memory":"16Gi"},"requests":{"cpu":"8","memory":"16Gi"}}}],"nodeSelector":{"disktype":"ssd-pro"}}}  creationTimestamp: "2024-08-12T16:02:58Z"  generation: 1  labels:    app: blog  name: blog-pending  namespace: blog-prod  resourceVersion: "3519"  uid: 8645b28b-c492-48ca-aefa-83ec351b04e4spec:  containers:  - image: nginx:1.27    imagePullPolicy: IfNotPresent    name: blog    resources:      limits:        cpu: "1"        memory: 1Gi      requests:        cpu: "1"        memory: 1Gi    terminationMessagePath: /dev/termination-log    terminationMessagePolicy: File    volumeMounts:    - mountPath: /var/run/secrets/kubernetes.io/serviceaccount      name: kube-api-access-9br4c      readOnly: true  dnsPolicy: ClusterFirst  enableServiceLinks: true  preemptionPolicy: PreemptLowerPriority  priority: 0  restartPolicy: Always  schedulerName: default-scheduler  securityContext: {}  serviceAccount: default  serviceAccountName: default  terminationGracePeriodSeconds: 30  tolerations:  - effect: NoExecute    key: node.kubernetes.io/not-ready    operator: Exists    tolerationSeconds: 300  - effect: NoExecute    key: node.kubernetes.io/unreachable    operator: Exists    tolerationSeconds: 300  volumes:  - name: kube-api-access-9br4c    projected:      defaultMode: 420      sources:      - serviceAccountToken:          expirationSeconds: 3607          path: token      - configMap:          items:          - key: ca.crt            path: ca.crt          name: kube-root-ca.crt      - downwardAPI:          items:          - fieldRef:              apiVersion: v1              fieldPath: metadata.namespace            path: namespace [ashutosh@lima-default ~]$ k get pod -n blog-prodNAME           READY   STATUS    RESTARTS   AGEblog-pending   0/1     Pending   0          11h [ashutosh@lima-default ~]$ k delete pod blog-pending -n blog-prodpod "blog-pending" deleted from blog-prod namespace [ashutosh@lima-default ~]$ k get pod -n blog-prodNo resources found in blog-prod namespace. [ashutosh@lima-default ~]$ k create -f fixed-blog-prod.yamlpod/blog-pending created [ashutosh@lima-default ~]$ k get pod -n blog-prodNAME           READY   STATUS    RESTARTS   AGEblog-pending   1/1     Running   0          4s

Conclusion

The pod was in a Pending state because of two reasons:

  1. The nodeSelector was immutable, so I could not edit it to match the node labels
  2. The resources requested by the pod were not available on the node

References

Table of contents

kubernetes

Waybill API Down: Four Broken Layers

The waybill pods would not even schedule, and behind that waited three more faults: a missing Secret key, a purposeless sidecar probe, and a Service selector matching nothing.

- 9 min read - kubernetes, troubleshooting, scheduling

kubernetes

Payments API Down While Pods Are Ready

The checkout pods were 1/1 Ready with populated endpoints, yet the payments API refused connections — then returned HTTP 500 after the first fix. Two layers: a wrong Service targetPort and a Secret with a trailing newline.

- 7 min read - kubernetes, troubleshooting, services

kubernetes

Pod is running but not ready

In this article, we will troubleshoot a scenario where the pod is running but not ready. We will explore the possible reasons for this issue and how to resolve it.

- 6 min read - kubernetes, troubleshooting, readiness-probe