Software Engineering WikiSE Wiki

Kubernetes

Inspect workloads with kubectl, understand the object model behind them and find why a pod is not running, ready or reachable.

Reviewed MarkdownEdit

On this page

Cheatsheet#

Commands assume kubectl 1.34 or later against a supported cluster. kubectl top needs metrics-server installed.

TaskCommand
Which cluster am I onkubectl config current-context
Workloads in a namespacekubectl get all -n my-namespace
Why is this pod unhappykubectl describe pod my-pod -n my-namespace
Events for one objectkubectl events --for pod/my-pod -n my-namespace
Events, newest lastkubectl get events -n my-namespace --sort-by=.lastTimestamp
Logs of the crashed containerkubectl logs my-pod -c app --previous -n my-namespace
Shell in a running podkubectl exec -it my-pod -n my-namespace -- sh
Debug a distroless podkubectl debug -it my-pod --image=nicolaka/netshoot --target=app
Port-forward a Servicekubectl port-forward svc/my-app 8080:80 -n my-namespace
Restart a Deploymentkubectl rollout restart deploy/my-app -n my-namespace
Watch a rolloutkubectl rollout status deploy/my-app -n my-namespace
Roll backkubectl rollout undo deploy/my-app -n my-namespace
Scalekubectl scale deploy/my-app --replicas=5 -n my-namespace
Resource usagekubectl top pod -n my-namespace --sort-by=memory
Can I do thiskubectl auth can-i delete pods -n my-namespace
Validate against the live APIkubectl apply -f my-app.yaml --dry-run=server
Show what apply would changekubectl diff -f my-app.yaml
Object as storedkubectl get pod my-pod -o yaml

Start with a failing workload#

Confirm the context first, then read the pod’s state and its events. describe merges the object status with the events the scheduler and kubelet recorded, which is where the actual reason lives. Events expire after one hour by default, so check them early.

kubectl config current-context
kubectl get pods -n my-namespace -o wide
kubectl describe pod my-pod -n my-namespace | sed -n '/Events:/,$p'
kubectl logs my-pod -n my-namespace -c app --previous --tail 100
Pod stateWhat it meansNext command
PendingNo node fits: requests, taints, affinity, or an unbound PVCkubectl describe pod and read the FailedScheduling event
ContainerCreatingImage pull, volume mount or CNI setup still in progress or failingkubectl describe pod, then kubelet logs on the node
ImagePullBackOff / ErrImagePullWrong name or tag, private registry, missing imagePullSecretskubectl events --for pod/my-pod, kubectl get sa default -o yaml
CrashLoopBackOffThe process keeps exiting; restarts back off 10s, 20s, 40s up to 5 minuteskubectl logs --previous
CreateContainerConfigErrorA referenced ConfigMap, Secret or key does not existkubectl describe pod, then kubectl get cm,secret
Running, not ReadyReadiness probe failingkubectl describe pod, check probe path and port
OOMKilled (exit 137)Container memory exceeded limits.memorykubectl top pod, raise the limit or fix the leak
Terminating foreverFinalizer not cleared, or the node is unreachable so the kubelet never confirmskubectl get pod my-pod -o jsonpath='{.metadata.finalizers}', kubectl get node
EvictedNode pressure (memory, disk, PIDs) reclaimed itkubectl describe node, read the conditions

The crash backoff resets after the container runs for 10 minutes without failing. See pod lifecycle for the feature gates that shorten it.

Changes in a GitOps-managed cluster

A reconciler such as Argo CD reverts a manual edit. Use direct commands for diagnosis only and apply lasting fixes through the repository that owns the resource.

How the control plane reconciles state#

The API server is the only component that writes to etcd. Controllers, the scheduler and kubelets watch the API for desired state and act until reality matches. Every fix is therefore “change the desired state and wait”, never “make the change on the node”.

ComponentJob
kube-apiserverAuthenticates, runs admission, validates and stores objects; the single write path
etcdConsistent key-value store holding cluster state
kube-schedulerBinds a pending pod to a node that satisfies its requests and constraints
kube-controller-managerReconciliation loops for Deployments, ReplicaSets, Jobs, nodes and EndpointSlices
kubeletRuns the containers for pods bound to its node and reports status
kube-proxy or a CNI replacementPrograms Service load balancing on each node (Cilium can replace it)

A Deployment owns a ReplicaSet, which owns Pods. Changing the pod template creates a new ReplicaSet and scales the old one down. That indirection is why kubectl rollout undo works, and why deleting a pod with a bad image only brings back another pod with the same bad image.

kubectl contexts, queries and applies#

kubectl config get-contexts
kubectl config use-context prod
kubectl config set-context --current --namespace=my-namespace   # default namespace for this context

kubectl get deploy,sts,ds,job -A                          # workloads in every namespace
kubectl get pods -A -o wide --field-selector status.phase!=Running
kubectl get pod my-pod -o jsonpath='{.spec.containers[*].image}{"\n"}'
kubectl explain deployment.spec.strategy --recursive       # schema served by this cluster
kubectl api-resources --namespaced=true                    # kinds this cluster serves
kubectl diff -f manifest.yaml                              # what applying would change, exit 1 if different
kubectl apply -f manifest.yaml --server-side               # server tracks field ownership per manager

kubectl get -o yaml returns the object after defaulting and admission, not what you submitted. Client-side apply records what you sent in the kubectl.kubernetes.io/last-applied-configuration annotation. Server-side apply records ownership in metadata.managedFields instead, and reports a conflict when another manager owns a field you try to set. Add --force-conflicts only when you intend to take that field over.

Workloads and rollouts#

apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
spec:
  replicas: 3
  revisionHistoryLimit: 5
  strategy:
    type: RollingUpdate
    rollingUpdate: { maxSurge: 1, maxUnavailable: 0 }
  selector:
    matchLabels: { app: my-app }                 # immutable after creation
  template:
    metadata:
      labels: { app: my-app }
    spec:
      terminationGracePeriodSeconds: 45
      securityContext:
        runAsNonRoot: true
        seccompProfile: { type: RuntimeDefault }
      containers:
        - name: app
          image: registry.example.com/my-app@sha256:<digest>   # digest, not a moving tag
          ports: [{ containerPort: 8080 }]
          resources:
            requests: { cpu: 100m, memory: 256Mi }   # what the scheduler reserves
            limits: { memory: 512Mi }                # what the kernel enforces
          readinessProbe:
            httpGet: { path: /readyz, port: 8080 }
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /healthz, port: 8080 }
            initialDelaySeconds: 20
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities: { drop: ["ALL"] }

Requests drive scheduling and CPU weight. Limits drive throttling and OOM kills. A CPU limit throttles rather than kills, which shows up as latency, not restarts, so many teams set memory limits and leave CPU limits off.

Readiness removes a pod from Service endpoints. Liveness restarts the container. Pointing liveness at an endpoint that checks a database turns a database blip into restarts across every replica. Use a startupProbe for slow starters instead of a long initialDelaySeconds.

kubectl rollout status deploy/my-app -n my-namespace --timeout=5m
kubectl rollout history deploy/my-app -n my-namespace
kubectl rollout undo deploy/my-app --to-revision=3 -n my-namespace
kubectl rollout restart deploy/my-app -n my-namespace    # new pods, same spec: picks up rotated Secrets in env vars
kubectl scale deploy/my-app --replicas=0 -n my-namespace  # stops every pod; the Deployment remains
KindUse it for
DeploymentStateless replicas, rolling updates, rollback
StatefulSetStable network identity and per-replica storage; ordered, slower updates
DaemonSetOne pod per node: node agents, CNI, log shippers
Job / CronJobRun to completion, bounded by backoffLimit and activeDeadlineSeconds

Native sidecars are init containers with restartPolicy: Always. They start before the app containers and stop after them, which fixes Jobs that never complete because a proxy sidecar keeps running. They are stable since v1.33.

A StatefulSet’s PVCs survive deletion of the StatefulSet by default. Removing them is a separate kubectl delete pvc, which destroys the data when the reclaim policy is Delete.

Probes and graceful shutdown#

Probe defaults are initialDelaySeconds: 0, periodSeconds: 10, timeoutSeconds: 1, failureThreshold: 3 and successThreshold: 1. A startup probe suspends the other two until it succeeds once, so failureThreshold * periodSeconds is the longest start you permit, and liveness and readiness then run with short thresholds of their own. The one-second timeout is the usual surprise: a handler that takes 1.5 s under load counts as a failure. This extends the probe rules in Workloads and rollouts.

    startupProbe:
      httpGet: { path: /healthz, port: 8080 }
      periodSeconds: 5
      failureThreshold: 60             # up to 5 minutes to start; liveness and readiness wait
    readinessProbe:
      httpGet: { path: /readyz, port: 8080 }
      periodSeconds: 5
      timeoutSeconds: 2
      failureThreshold: 2              # out of the endpoints after 10 s of failures
    livenessProbe:
      grpc: { port: 9090 }             # gRPC health protocol; kubelet speaks it natively
      periodSeconds: 20
      timeoutSeconds: 5
    lifecycle:
      preStop:
        sleep: { seconds: 5 }          # stable in v1.34; keeps serving while endpoint removal propagates

Deleting a pod starts two things in parallel: the kubelet begins the local shutdown, and the EndpointSlice controller marks the endpoint terminating with ready: false, which kube-proxy and external load balancers act on after their own delay. The kubelet runs the preStop hook to completion, then has the runtime send SIGTERM to PID 1 of each container, and when terminationGracePeriodSeconds (default 30, counted from the start of the hook) expires it sends SIGKILL. Hook and shutdown share that one budget: a 25 s hook plus a 10 s shutdown under a 30 s grace period ends in SIGKILL. A short preStop sleep covers the propagation gap so requests are not routed to a container that has already stopped listening. Containers receive SIGTERM in arbitrary order unless the helper is a native sidecar, which stops after the main containers.

kubectl exec my-pod -n my-namespace -- cat /proc/1/cmdline | tr '\0' ' '   # what PID 1 is; sh -c 'app' does not forward SIGTERM
kubectl get pod my-pod -n my-namespace -o jsonpath='{.metadata.deletionTimestamp} {.metadata.deletionGracePeriodSeconds}{"\n"}'

Resource requests, limits and QoS classes#

The API server derives a QoS class from the containers’ requests and limits, and the kubelet evicts in reverse order of that class under node pressure. Guaranteed requires every container to set CPU and memory limits equal to its requests, Burstable is any pod with at least one request or limit, and BestEffort has none. Guaranteed pods are the only ones the static CPU manager policy can pin to exclusive cores. The Deployment in Workloads and rollouts is Burstable, which is the normal choice for services.

kubectl get pods -n my-namespace -o custom-columns='POD:.metadata.name,QOS:.status.qosClass,NODE:.spec.nodeName'
kubectl set resources deploy/my-app -n my-namespace -c app --requests=cpu=100m,memory=256Mi --limits=memory=512Mi   # changes the template: rollout
kubectl describe quota -n my-namespace                    # ResourceQuota usage; a full quota rejects new pods at admission
kubectl get limitrange -n my-namespace -o yaml            # defaults injected into containers that set nothing

A LimitRange with default and defaultRequest stops BestEffort pods appearing in a namespace, and a ResourceQuota on requests.cpu makes any pod without that request fail admission with must specify requests.cpu. Memory requests set the container’s OOM score adjustment, so a container far above its request is killed first when the node runs out, before it reaches its own limit.

PodDisruptionBudgets#

A PDB limits voluntary disruptions: kubectl drain, the eviction API, cluster autoscaler scale-down and managed node upgrades. It does nothing for crashes, OOM kills or hardware failure. minAvailable and maxUnavailable take a count or a percentage, only one may be set, and maxUnavailable requires that every selected pod share one controller. minAvailable equal to the replica count blocks every drain forever, which is the most common way to stall a node upgrade.

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: my-app, namespace: my-namespace }
spec:
  maxUnavailable: 1                          # or "25%"; percentages round up
  unhealthyPodEvictionPolicy: AlwaysAllow    # stable in v1.31: pods that are not Ready may always be evicted
  selector:
    matchLabels: { app: my-app }

The default IfHealthyBudget refuses to evict a crash-looping pod while the budget is unmet, so a broken deployment blocks a drain; AlwaysAllow is the right setting for almost every service. Check what a drain will run into before starting it:

kubectl get pdb -A                                                                   # ALLOWED DISRUPTIONS 0 means drain waits
kubectl get pdb my-app -n my-namespace -o jsonpath='{.status.disruptionsAllowed}{"\n"}'
kubectl drain node-1 --ignore-daemonsets --delete-emptydir-data --dry-run=server     # lists the pods it would evict

Horizontal Pod Autoscaler#

An autoscaling/v2 HPA sets replicas through the target’s scale subresource from ceil(currentReplicas * currentMetric / targetMetric), evaluated every 15 s and ignored when the ratio is within 10% of 1. Utilization targets are percentages of the container requests, so a pod whose container has no request for that metric is skipped and the HPA reports FailedGetResourceMetric. Remove replicas from the Deployment manifest once an HPA owns it: every apply resets the count and the HPA scales it back, which shows up as a rollout on each GitOps sync.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: my-app, namespace: my-namespace }
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: my-app }
  minReplicas: 2
  maxReplicas: 20
  metrics:
    - type: ContainerResource              # one container, not the pod total: ignores the sidecar
      containerResource:
        name: cpu
        container: app
        target: { type: Utilization, averageUtilization: 70 }
    - type: Resource
      resource:
        name: memory
        target: { type: AverageValue, averageValue: 400Mi }
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0        # act on the latest recommendation immediately
      policies:
        - { type: Percent, value: 100, periodSeconds: 15 }   # the default: double, or add 4 pods, per 15 s
    scaleDown:
      stabilizationWindowSeconds: 300      # the default: use the highest recommendation of the last 5 minutes
      policies:
        - { type: Pods, value: 2, periodSeconds: 60 }
      selectPolicy: Min                    # Max (default), Min, or Disabled to never scale down
kubectl autoscale deploy/my-app -n my-namespace --min=2 --max=10 --cpu=70%   # --cpu and --memory take a percentage or a quantity such as 500m
kubectl get hpa -n my-namespace                                                # TARGETS shows <unknown> until metrics arrive
kubectl describe hpa my-app -n my-namespace | sed -n '/Conditions:/,$p'        # ScalingActive False says why nothing happens

With several metrics the HPA takes the largest desired replica count. Memory rarely drops when load does because most runtimes keep freed heap, so a memory target mostly scales up.

Node affinity, taints and topology spread#

The scheduler filters out nodes that fail hard constraints, then scores the rest. nodeSelector and requiredDuringSchedulingIgnoredDuringExecution filter; preferredDuringSchedulingIgnoredDuringExecution adds a weighted score. IgnoredDuringExecution means a running pod stays put when node labels change. Taints are the reverse: a node repels pods that do not tolerate the taint. NoSchedule affects new pods, PreferNoSchedule is the soft form, and NoExecute also evicts running pods after tolerationSeconds. The control plane taints unhealthy nodes with node.kubernetes.io/not-ready and node.kubernetes.io/unreachable as NoExecute, and every pod gets a default 300 s toleration for both, which is the five-minute wait before pods on a dead node are replaced.

kubectl label nodes node-1 workload=batch                        # for selectors and affinity
kubectl taint nodes node-1 dedicated=batch:NoSchedule            # repel pods without the toleration
kubectl taint nodes node-1 dedicated=batch:NoSchedule-           # trailing dash removes it
kubectl cordon node-1                                            # adds node.kubernetes.io/unschedulable:NoSchedule; running pods stay
kubectl get nodes -o custom-columns='NODE:.metadata.name,TAINTS:.spec.taints[*].key,ZONE:.metadata.labels.topology\.kubernetes\.io/zone'
spec:
  tolerations:
    - { key: dedicated, operator: Equal, value: batch, effect: NoSchedule }
    - key: node.kubernetes.io/unreachable
      operator: Exists
      effect: NoExecute
      tolerationSeconds: 60                 # replace this pod after 1 minute instead of 5
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:                  # terms are ORed; expressions within a term are ANDed
          - matchExpressions:
              - { key: kubernetes.io/arch, operator: In, values: [amd64, arm64] }
              - { key: node.kubernetes.io/instance-type, operator: NotIn, values: [t3.micro] }
      preferredDuringSchedulingIgnoredDuringExecution:
        - weight: 80
          preference:
            matchExpressions: [{ key: workload, operator: In, values: [batch] }]
  topologySpreadConstraints:
    - maxSkew: 1
      topologyKey: topology.kubernetes.io/zone
      whenUnsatisfiable: DoNotSchedule      # hard; ScheduleAnyway only affects scoring
      labelSelector: { matchLabels: { app: my-app } }
      minDomains: 3                         # zones with no matching pod count as domains (stable in v1.30)
    - maxSkew: 1
      topologyKey: kubernetes.io/hostname
      whenUnsatisfiable: ScheduleAnyway
      labelSelector: { matchLabels: { app: my-app } }
      matchLabelKeys: [pod-template-hash]   # count only this ReplicaSet, so a rollout ignores old pods (beta since v1.27)

Topology spread replaces podAntiAffinity for “one replica per zone or node” and scales better, because anti-affinity is evaluated against every pod in the cluster. A hard spread with maxSkew: 1 and a zone that has no schedulable capacity leaves pods Pending with didn't match pod topology spread constraints; nodeTaintsPolicy: Honor and nodeAffinityPolicy: Honor (both beta since v1.26) make the scheduler exclude such nodes from the skew calculation.

Jobs and CronJobs#

A Job runs pods until completions succeed (default 1), parallelism at a time, and fails after backoffLimit failed pods (default 6) or activeDeadlineSeconds, whichever comes first. The pod’s restartPolicy must be Never or OnFailure; with OnFailure the retries hide inside one pod’s restartCount, so Never gives a clearer failure history. ttlSecondsAfterFinished deletes the Job and its pods after it finishes; without it, finished Jobs and their logs accumulate.

apiVersion: batch/v1
kind: Job
metadata: { name: migrate, namespace: my-namespace }
spec:
  backoffLimit: 3
  activeDeadlineSeconds: 900               # kill everything 15 minutes after the Job starts
  ttlSecondsAfterFinished: 86400           # delete the Job and its pods a day after it finishes
  podFailurePolicy:                        # stable in v1.31
    rules:
      - action: FailJob                    # do not retry a configuration error
        onExitCodes: { containerName: migrate, operator: In, values: [2] }
      - action: Ignore                     # preemption or a drain does not count against backoffLimit
        onPodConditions: [{ type: DisruptionTarget }]
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: migrate
          image: registry.example.com/my-app@sha256:<digest>
          args: [migrate, --to, latest]
apiVersion: batch/v1
kind: CronJob
metadata: { name: report, namespace: my-namespace }
spec:
  schedule: "15 2 * * *"
  timeZone: Australia/Melbourne            # stable in v1.27; unset means the controller-manager's zone
  concurrencyPolicy: Forbid                # skip the run if the previous Job is still active; Replace kills it
  startingDeadlineSeconds: 600             # give up on a run that could not start within 10 minutes
  successfulJobsHistoryLimit: 3
  failedJobsHistoryLimit: 3                # default 1 hides the failure you are looking for
  jobTemplate:
    spec:
      backoffLimit: 2
      template:
        spec:
          restartPolicy: OnFailure
          containers:
            - name: report
              image: registry.example.com/report@sha256:<digest>

backoffLimitPerIndex and successPolicy (both stable in v1.33) need completionMode: Indexed. The controller checks schedules every 10 s. A CronJob that misses more than 100 start times, typically after a long suspend or controller outage, stops scheduling and logs Too many missed start time (> 100); setting startingDeadlineSeconds bounds the window it counts over and is the fix.

kubectl create job report-now --from=cronjob/report -n my-namespace         # run the template immediately
kubectl get jobs -n my-namespace -o wide                                   # COMPLETIONS and DURATION per run
kubectl logs job/report-now -n my-namespace --all-containers                # logs of the Job's pod
kubectl patch cronjob report -n my-namespace -p '{"spec":{"suspend":true}}' # pause; a running Job continues
kubectl delete jobs -n my-namespace --field-selector status.successful=1    # remove finished Jobs and their pods

Services and networking#

A Service is a stable virtual IP plus a label selector. The EndpointSlice controller keeps a list of the selected pods’ IPs and readiness, and kube-proxy (or a replacement such as Cilium) programs the dataplane from it. No ready pods means no endpoints, which usually presents as connection refused or a timeout rather than an error message. The older Endpoints API is deprecated since v1.33; read EndpointSlices.

kubectl get svc,endpointslice -n my-namespace
kubectl get endpointslice -l kubernetes.io/service-name=my-app -n my-namespace -o yaml | grep -A3 addresses
kubectl run tmp --rm -it --image=nicolaka/netshoot -n my-namespace -- sh   # curl, dig, tcpdump in-cluster; pod deleted on exit
kubectl port-forward svc/my-app 8080:80 -n my-namespace
TypeBehaviour
ClusterIPIn-cluster virtual IP; the default
NodePortClusterIP plus a port (30000-32767 by default) on every node
LoadBalancerNodePort plus an external load balancer provisioned by a controller
ExternalNameDNS CNAME only, no proxying
Headless (clusterIP: None)DNS returns pod IPs directly; how StatefulSet members are addressed

DNS names follow <service>.<namespace>.svc.cluster.local. Inside a pod, my-app resolves through the search domains in /etc/resolv.conf. Across namespaces, my-app.other-namespace is the shortest reliable form. The default ndots:5 makes external names try every search domain first; see DNS for the effect. For HTTP routing into the cluster, use Gateway API.

NetworkPolicies are allow-lists. A pod selected by no policy for a direction allows all traffic in that direction. Once any policy selects it for ingress or egress, everything in that direction not explicitly allowed is denied. When both the client’s egress and the server’s ingress are isolated, both sides must allow the flow. Enforcement depends on the CNI; a CNI without policy support accepts the objects and ignores them.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: my-app-allow }
spec:
  podSelector: { matchLabels: { app: my-app } }
  policyTypes: [Ingress, Egress]
  ingress:
    - from:
        - podSelector: { matchLabels: { app: web } }
      ports: [{ protocol: TCP, port: 8080 }]
  egress:
    - to: [{ namespaceSelector: { matchLabels: { kubernetes.io/metadata.name: kube-system } } }]
      ports:
        - { protocol: UDP, port: 53 }   # without DNS egress, every name lookup fails
        - { protocol: TCP, port: 53 }

Storage#

A PVC is a request, a PV is the volume that satisfies it, and a StorageClass provisions PVs on demand. volumeBindingMode: WaitForFirstConsumer delays provisioning until a pod is scheduled, so the volume is created in the same zone as the node.

kubectl get pvc,pv -n my-namespace
kubectl get sc
kubectl describe pvc data-my-app-0 -n my-namespace   # provisioning errors appear as events
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: data }
spec:
  accessModes: [ReadWriteOnce]              # one node, not one pod
  storageClassName: gp3
  resources: { requests: { storage: 20Gi } }

ReadWriteOnce allows many pods on the same node to mount the volume. Use ReadWriteOncePod when exactly one pod may mount it. Expansion works in place when the StorageClass sets allowVolumeExpansion: true; shrinking is not supported. A PVC stuck Terminating is still mounted by a pod; the kubernetes.io/pvc-protection finalizer holds it until that pod is gone.

ConfigMaps and Secrets#

ConfigMaps and Secrets use the same mechanism with different handling. Secret values are base64-encoded in the API (encoding, not encryption), encrypted in etcd only if the cluster configures encryption at rest, and mounted on tmpfs.

kubectl create configmap app-config --from-file=config.yaml --dry-run=client -o yaml > cm.yaml
kubectl create secret generic db --from-literal=password="$DB_PASSWORD" --dry-run=client -o yaml > secret.yaml   # file holds the value; do not commit it
kubectl get secret db -o jsonpath='{.data.password}' | base64 -d   # prints the secret to the terminal
    envFrom:
      - configMapRef: { name: app-config }
    env:
      - name: DB_PASSWORD
        valueFrom:
          secretKeyRef: { name: db, key: password }
    volumeMounts:
      - { name: config, mountPath: /etc/app, readOnly: true }
  volumes:
    - name: config
      configMap: { name: app-config }

Mounted ConfigMaps and Secrets update in place after the kubelet sync period plus cache delay, typically within a minute or two. Environment variables and subPath mounts never update. Applications that read configuration once need kubectl rollout restart after a change.

RBAC and service accounts#

Every pod runs as a ServiceAccount. Its short-lived token is projected into the pod and used for API calls. RBAC binds Roles (namespaced) or ClusterRoles to users, groups or ServiceAccounts. Permissions are additive; there is no deny rule.

kubectl auth can-i --list -n my-namespace                                          # my permissions here
kubectl auth can-i get secrets -n my-namespace --as system:serviceaccount:my-namespace:my-app
kubectl auth whoami                                                                 # identity the API server sees
kubectl get rolebinding,clusterrolebinding -A -o wide | grep my-namespace
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata: { name: pod-reader, namespace: my-namespace }
rules:
  - apiGroups: [""]
    resources: ["pods", "pods/log"]
    verbs: ["get", "list", "watch"]

Set automountServiceAccountToken: false on workloads that never call the API. Anything that can read Secrets in a namespace, or create pods there, can obtain every ServiceAccount token in it.

Troubleshooting#

kubectl get events -A --sort-by=.lastTimestamp | tail -30
kubectl describe node node-1 | sed -n '/Conditions:/,/Events:/p'
kubectl top node; kubectl top pod -A --sort-by=cpu
kubectl get pod my-pod -o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'
kubectl debug node/node-1 -it --image=busybox     # pod in host namespaces; host filesystem at /host; delete the pod afterwards
kubectl get --raw='/readyz?verbose'               # API server health checks, one line each
SymptomLikely causeCheck
Pods Pending, nodes look idleRequests exceed allocatable, or taints without matching tolerationskubectl describe node, Allocated resources section
Service fails intermittentlySome replicas failing readinessEndpointSlice membership over time
DNS slow or failingCoreDNS unhealthy, or a NetworkPolicy blocking port 53kubectl get pods -n kube-system -l k8s-app=kube-dns, DNS
exec works, curl from another pod does notProcess listening on 127.0.0.1 inside the podkubectl exec my-pod -- ss -ltn
Changes revert within minutesA GitOps controller owns the resourcekubectl get <kind> <name> -o yaml, look for Argo CD or Flux labels and annotations
Node NotReadykubelet stopped, disk or PID pressure, or CNI failureNode conditions, then journalctl -u kubelet on the node (systemd)
forbidden from the APIRBAC missing for that verb, resource or namespacekubectl auth can-i <verb> <resource> --as <subject>
Pod restarts with no error in logsLiveness probe killing a slow containerkubectl describe pod, look for Liveness probe failed events

Oneliners#

# Pods not Running or Succeeded, cluster-wide
kubectl get pods -A --field-selector 'status.phase!=Running,status.phase!=Succeeded'

# Top restart counts (first container of each pod)
kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}{"\t"}{.metadata.name}{"\t"}{.status.containerStatuses[0].restartCount}{"\n"}{end}' | sort -k3 -nr | head

# Every image running in the cluster, deduplicated
kubectl get pods -A -o jsonpath='{range .items[*]}{range .spec.containers[*]}{.image}{"\n"}{end}{end}' | sort -u

# Requests per pod
kubectl get pods -A -o custom-columns='NS:.metadata.namespace,POD:.metadata.name,CPU:.spec.containers[*].resources.requests.cpu,MEM:.spec.containers[*].resources.requests.memory'

# Requested versus allocatable on a node
kubectl describe node node-1 | awk '/Allocated resources/,/Events/'

# Pods on one node
kubectl get pods -A -o wide --field-selector spec.nodeName=node-1

# Which pods mount a given Secret as a volume
kubectl get pods -A -o json | jq -r '.items[] | select(.spec.volumes[]?.secret.secretName=="db") | "\(.metadata.namespace)/\(.metadata.name)"'

# Drain a node for maintenance (evicts pods, respects PodDisruptionBudgets), then allow scheduling again
kubectl drain node-1 --ignore-daemonsets --delete-emptydir-data --timeout=5m && kubectl uncordon node-1

# Watch rollouts of every Deployment in a namespace in parallel
kubectl get deploy -n my-namespace -o name | xargs -n1 -P0 kubectl rollout status -n my-namespace

# Object counts by resource, to find what is filling etcd (metric renamed in v1.34)
kubectl get --raw=/metrics | grep -E '^apiserver_(storage|resource)_objects' | sort -t' ' -k2 -nr | head

# Copy a file out of a pod without tar in the image
kubectl exec my-pod -n my-namespace -- cat /app/report.csv > report.csv

Two commands here need care:

# Print every key of a Secret in plain text. Output contains credentials.
kubectl get secret db -o go-template='{{range $k,$v := .data}}{{$k}}={{$v|base64decode}}{{"\n"}}{{end}}'

Force deletion

kubectl delete pod my-pod --grace-period=0 --force removes the pod object without waiting for the kubelet to confirm the containers stopped. If the node is only partitioned, the old container can keep running alongside its replacement, which corrupts data for StatefulSets. It does not bypass finalizers; a pod with finalizers stays Terminating until they are removed.