Kubernetes
Inspect workloads with kubectl, understand the object model behind them and find why a pod is not running, ready or reachable.
On this page
Cheatsheet#
Commands assume kubectl 1.34 or later against a supported cluster. kubectl top needs metrics-server installed.
| Task | Command |
|---|---|
| Which cluster am I on | kubectl config current-context |
| Workloads in a namespace | kubectl get all -n my-namespace |
| Why is this pod unhappy | kubectl describe pod my-pod -n my-namespace |
| Events for one object | kubectl events --for pod/my-pod -n my-namespace |
| Events, newest last | kubectl get events -n my-namespace --sort-by=.lastTimestamp |
| Logs of the crashed container | kubectl logs my-pod -c app --previous -n my-namespace |
| Shell in a running pod | kubectl exec -it my-pod -n my-namespace -- sh |
| Debug a distroless pod | kubectl debug -it my-pod --image=nicolaka/netshoot --target=app |
| Port-forward a Service | kubectl port-forward svc/my-app 8080:80 -n my-namespace |
| Restart a Deployment | kubectl rollout restart deploy/my-app -n my-namespace |
| Watch a rollout | kubectl rollout status deploy/my-app -n my-namespace |
| Roll back | kubectl rollout undo deploy/my-app -n my-namespace |
| Scale | kubectl scale deploy/my-app --replicas=5 -n my-namespace |
| Resource usage | kubectl top pod -n my-namespace --sort-by=memory |
| Can I do this | kubectl auth can-i delete pods -n my-namespace |
| Validate against the live API | kubectl apply -f my-app.yaml --dry-run=server |
| Show what apply would change | kubectl diff -f my-app.yaml |
| Object as stored | kubectl get pod my-pod -o yaml |
Start with a failing workload#
Confirm the context first, then read the pod’s state and its events. describe merges the object status with the events the scheduler and kubelet recorded, which is where the actual reason lives. Events expire after one hour by default, so check them early.
kubectl config current-context
kubectl get pods -n my-namespace -o wide
kubectl describe pod my-pod -n my-namespace | sed -n '/Events:/,$p'
kubectl logs my-pod -n my-namespace -c app --previous --tail 100| Pod state | What it means | Next command |
|---|---|---|
Pending | No node fits: requests, taints, affinity, or an unbound PVC | kubectl describe pod and read the FailedScheduling event |
ContainerCreating | Image pull, volume mount or CNI setup still in progress or failing | kubectl describe pod, then kubelet logs on the node |
ImagePullBackOff / ErrImagePull | Wrong name or tag, private registry, missing imagePullSecrets | kubectl events --for pod/my-pod, kubectl get sa default -o yaml |
CrashLoopBackOff | The process keeps exiting; restarts back off 10s, 20s, 40s up to 5 minutes | kubectl logs --previous |
CreateContainerConfigError | A referenced ConfigMap, Secret or key does not exist | kubectl describe pod, then kubectl get cm,secret |
Running, not Ready | Readiness probe failing | kubectl describe pod, check probe path and port |
OOMKilled (exit 137) | Container memory exceeded limits.memory | kubectl top pod, raise the limit or fix the leak |
Terminating forever | Finalizer not cleared, or the node is unreachable so the kubelet never confirms | kubectl get pod my-pod -o jsonpath='{.metadata.finalizers}', kubectl get node |
Evicted | Node pressure (memory, disk, PIDs) reclaimed it | kubectl describe node, read the conditions |
The crash backoff resets after the container runs for 10 minutes without failing. See pod lifecycle for the feature gates that shorten it.
Changes in a GitOps-managed cluster
A reconciler such as Argo CD reverts a manual edit. Use direct commands for diagnosis only and apply lasting fixes through the repository that owns the resource.
How the control plane reconciles state#
The API server is the only component that writes to etcd. Controllers, the scheduler and kubelets watch the API for desired state and act until reality matches. Every fix is therefore “change the desired state and wait”, never “make the change on the node”.
| Component | Job |
|---|---|
kube-apiserver | Authenticates, runs admission, validates and stores objects; the single write path |
etcd | Consistent key-value store holding cluster state |
kube-scheduler | Binds a pending pod to a node that satisfies its requests and constraints |
kube-controller-manager | Reconciliation loops for Deployments, ReplicaSets, Jobs, nodes and EndpointSlices |
kubelet | Runs the containers for pods bound to its node and reports status |
kube-proxy or a CNI replacement | Programs Service load balancing on each node (Cilium can replace it) |
A Deployment owns a ReplicaSet, which owns Pods. Changing the pod template creates a new ReplicaSet and scales the old one down. That indirection is why kubectl rollout undo works, and why deleting a pod with a bad image only brings back another pod with the same bad image.
kubectl contexts, queries and applies#
kubectl config get-contexts
kubectl config use-context prod
kubectl config set-context --current --namespace=my-namespace # default namespace for this context
kubectl get deploy,sts,ds,job -A # workloads in every namespace
kubectl get pods -A -o wide --field-selector status.phase!=Running
kubectl get pod my-pod -o jsonpath='{.spec.containers[*].image}{"\n"}'
kubectl explain deployment.spec.strategy --recursive # schema served by this cluster
kubectl api-resources --namespaced=true # kinds this cluster serves
kubectl diff -f manifest.yaml # what applying would change, exit 1 if different
kubectl apply -f manifest.yaml --server-side # server tracks field ownership per managerkubectl get -o yaml returns the object after defaulting and admission, not what you submitted. Client-side apply records what you sent in the kubectl.kubernetes.io/last-applied-configuration annotation. Server-side apply records ownership in metadata.managedFields instead, and reports a conflict when another manager owns a field you try to set. Add --force-conflicts only when you intend to take that field over.
Workloads and rollouts#
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-app
spec:
replicas: 3
revisionHistoryLimit: 5
strategy:
type: RollingUpdate
rollingUpdate: { maxSurge: 1, maxUnavailable: 0 }
selector:
matchLabels: { app: my-app } # immutable after creation
template:
metadata:
labels: { app: my-app }
spec:
terminationGracePeriodSeconds: 45
securityContext:
runAsNonRoot: true
seccompProfile: { type: RuntimeDefault }
containers:
- name: app
image: registry.example.com/my-app@sha256:<digest> # digest, not a moving tag
ports: [{ containerPort: 8080 }]
resources:
requests: { cpu: 100m, memory: 256Mi } # what the scheduler reserves
limits: { memory: 512Mi } # what the kernel enforces
readinessProbe:
httpGet: { path: /readyz, port: 8080 }
periodSeconds: 5
livenessProbe:
httpGet: { path: /healthz, port: 8080 }
initialDelaySeconds: 20
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities: { drop: ["ALL"] }Requests drive scheduling and CPU weight. Limits drive throttling and OOM kills. A CPU limit throttles rather than kills, which shows up as latency, not restarts, so many teams set memory limits and leave CPU limits off.
Readiness removes a pod from Service endpoints. Liveness restarts the container. Pointing liveness at an endpoint that checks a database turns a database blip into restarts across every replica. Use a startupProbe for slow starters instead of a long initialDelaySeconds.
kubectl rollout status deploy/my-app -n my-namespace --timeout=5m
kubectl rollout history deploy/my-app -n my-namespace
kubectl rollout undo deploy/my-app --to-revision=3 -n my-namespace
kubectl rollout restart deploy/my-app -n my-namespace # new pods, same spec: picks up rotated Secrets in env vars
kubectl scale deploy/my-app --replicas=0 -n my-namespace # stops every pod; the Deployment remains| Kind | Use it for |
|---|---|
Deployment | Stateless replicas, rolling updates, rollback |
StatefulSet | Stable network identity and per-replica storage; ordered, slower updates |
DaemonSet | One pod per node: node agents, CNI, log shippers |
Job / CronJob | Run to completion, bounded by backoffLimit and activeDeadlineSeconds |
Native sidecars are init containers with restartPolicy: Always. They start before the app containers and stop after them, which fixes Jobs that never complete because a proxy sidecar keeps running. They are stable since v1.33.
A StatefulSet’s PVCs survive deletion of the StatefulSet by default. Removing them is a separate kubectl delete pvc, which destroys the data when the reclaim policy is Delete.
Probes and graceful shutdown#
Probe defaults are initialDelaySeconds: 0, periodSeconds: 10, timeoutSeconds: 1, failureThreshold: 3 and successThreshold: 1. A startup probe suspends the other two until it succeeds once, so failureThreshold * periodSeconds is the longest start you permit, and liveness and readiness then run with short thresholds of their own. The one-second timeout is the usual surprise: a handler that takes 1.5 s under load counts as a failure. This extends the probe rules in Workloads and rollouts.
startupProbe:
httpGet: { path: /healthz, port: 8080 }
periodSeconds: 5
failureThreshold: 60 # up to 5 minutes to start; liveness and readiness wait
readinessProbe:
httpGet: { path: /readyz, port: 8080 }
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 2 # out of the endpoints after 10 s of failures
livenessProbe:
grpc: { port: 9090 } # gRPC health protocol; kubelet speaks it natively
periodSeconds: 20
timeoutSeconds: 5
lifecycle:
preStop:
sleep: { seconds: 5 } # stable in v1.34; keeps serving while endpoint removal propagatesDeleting a pod starts two things in parallel: the kubelet begins the local shutdown, and the EndpointSlice controller marks the endpoint terminating with ready: false, which kube-proxy and external load balancers act on after their own delay. The kubelet runs the preStop hook to completion, then has the runtime send SIGTERM to PID 1 of each container, and when terminationGracePeriodSeconds (default 30, counted from the start of the hook) expires it sends SIGKILL. Hook and shutdown share that one budget: a 25 s hook plus a 10 s shutdown under a 30 s grace period ends in SIGKILL. A short preStop sleep covers the propagation gap so requests are not routed to a container that has already stopped listening. Containers receive SIGTERM in arbitrary order unless the helper is a native sidecar, which stops after the main containers.
kubectl exec my-pod -n my-namespace -- cat /proc/1/cmdline | tr '\0' ' ' # what PID 1 is; sh -c 'app' does not forward SIGTERM
kubectl get pod my-pod -n my-namespace -o jsonpath='{.metadata.deletionTimestamp} {.metadata.deletionGracePeriodSeconds}{"\n"}'Resource requests, limits and QoS classes#
The API server derives a QoS class from the containers’ requests and limits, and the kubelet evicts in reverse order of that class under node pressure. Guaranteed requires every container to set CPU and memory limits equal to its requests, Burstable is any pod with at least one request or limit, and BestEffort has none. Guaranteed pods are the only ones the static CPU manager policy can pin to exclusive cores. The Deployment in Workloads and rollouts is Burstable, which is the normal choice for services.
kubectl get pods -n my-namespace -o custom-columns='POD:.metadata.name,QOS:.status.qosClass,NODE:.spec.nodeName'
kubectl set resources deploy/my-app -n my-namespace -c app --requests=cpu=100m,memory=256Mi --limits=memory=512Mi # changes the template: rollout
kubectl describe quota -n my-namespace # ResourceQuota usage; a full quota rejects new pods at admission
kubectl get limitrange -n my-namespace -o yaml # defaults injected into containers that set nothingA LimitRange with default and defaultRequest stops BestEffort pods appearing in a namespace, and a ResourceQuota on requests.cpu makes any pod without that request fail admission with must specify requests.cpu. Memory requests set the container’s OOM score adjustment, so a container far above its request is killed first when the node runs out, before it reaches its own limit.
PodDisruptionBudgets#
A PDB limits voluntary disruptions: kubectl drain, the eviction API, cluster autoscaler scale-down and managed node upgrades. It does nothing for crashes, OOM kills or hardware failure. minAvailable and maxUnavailable take a count or a percentage, only one may be set, and maxUnavailable requires that every selected pod share one controller. minAvailable equal to the replica count blocks every drain forever, which is the most common way to stall a node upgrade.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: my-app, namespace: my-namespace }
spec:
maxUnavailable: 1 # or "25%"; percentages round up
unhealthyPodEvictionPolicy: AlwaysAllow # stable in v1.31: pods that are not Ready may always be evicted
selector:
matchLabels: { app: my-app }The default IfHealthyBudget refuses to evict a crash-looping pod while the budget is unmet, so a broken deployment blocks a drain; AlwaysAllow is the right setting for almost every service. Check what a drain will run into before starting it:
kubectl get pdb -A # ALLOWED DISRUPTIONS 0 means drain waits
kubectl get pdb my-app -n my-namespace -o jsonpath='{.status.disruptionsAllowed}{"\n"}'
kubectl drain node-1 --ignore-daemonsets --delete-emptydir-data --dry-run=server # lists the pods it would evictHorizontal Pod Autoscaler#
An autoscaling/v2 HPA sets replicas through the target’s scale subresource from ceil(currentReplicas * currentMetric / targetMetric), evaluated every 15 s and ignored when the ratio is within 10% of 1. Utilization targets are percentages of the container requests, so a pod whose container has no request for that metric is skipped and the HPA reports FailedGetResourceMetric. Remove replicas from the Deployment manifest once an HPA owns it: every apply resets the count and the HPA scales it back, which shows up as a rollout on each GitOps sync.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: my-app, namespace: my-namespace }
spec:
scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: my-app }
minReplicas: 2
maxReplicas: 20
metrics:
- type: ContainerResource # one container, not the pod total: ignores the sidecar
containerResource:
name: cpu
container: app
target: { type: Utilization, averageUtilization: 70 }
- type: Resource
resource:
name: memory
target: { type: AverageValue, averageValue: 400Mi }
behavior:
scaleUp:
stabilizationWindowSeconds: 0 # act on the latest recommendation immediately
policies:
- { type: Percent, value: 100, periodSeconds: 15 } # the default: double, or add 4 pods, per 15 s
scaleDown:
stabilizationWindowSeconds: 300 # the default: use the highest recommendation of the last 5 minutes
policies:
- { type: Pods, value: 2, periodSeconds: 60 }
selectPolicy: Min # Max (default), Min, or Disabled to never scale downkubectl autoscale deploy/my-app -n my-namespace --min=2 --max=10 --cpu=70% # --cpu and --memory take a percentage or a quantity such as 500m
kubectl get hpa -n my-namespace # TARGETS shows <unknown> until metrics arrive
kubectl describe hpa my-app -n my-namespace | sed -n '/Conditions:/,$p' # ScalingActive False says why nothing happensWith several metrics the HPA takes the largest desired replica count. Memory rarely drops when load does because most runtimes keep freed heap, so a memory target mostly scales up.
Node affinity, taints and topology spread#
The scheduler filters out nodes that fail hard constraints, then scores the rest. nodeSelector and requiredDuringSchedulingIgnoredDuringExecution filter; preferredDuringSchedulingIgnoredDuringExecution adds a weighted score. IgnoredDuringExecution means a running pod stays put when node labels change. Taints are the reverse: a node repels pods that do not tolerate the taint. NoSchedule affects new pods, PreferNoSchedule is the soft form, and NoExecute also evicts running pods after tolerationSeconds. The control plane taints unhealthy nodes with node.kubernetes.io/not-ready and node.kubernetes.io/unreachable as NoExecute, and every pod gets a default 300 s toleration for both, which is the five-minute wait before pods on a dead node are replaced.
kubectl label nodes node-1 workload=batch # for selectors and affinity
kubectl taint nodes node-1 dedicated=batch:NoSchedule # repel pods without the toleration
kubectl taint nodes node-1 dedicated=batch:NoSchedule- # trailing dash removes it
kubectl cordon node-1 # adds node.kubernetes.io/unschedulable:NoSchedule; running pods stay
kubectl get nodes -o custom-columns='NODE:.metadata.name,TAINTS:.spec.taints[*].key,ZONE:.metadata.labels.topology\.kubernetes\.io/zone'spec:
tolerations:
- { key: dedicated, operator: Equal, value: batch, effect: NoSchedule }
- key: node.kubernetes.io/unreachable
operator: Exists
effect: NoExecute
tolerationSeconds: 60 # replace this pod after 1 minute instead of 5
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms: # terms are ORed; expressions within a term are ANDed
- matchExpressions:
- { key: kubernetes.io/arch, operator: In, values: [amd64, arm64] }
- { key: node.kubernetes.io/instance-type, operator: NotIn, values: [t3.micro] }
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 80
preference:
matchExpressions: [{ key: workload, operator: In, values: [batch] }]
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule # hard; ScheduleAnyway only affects scoring
labelSelector: { matchLabels: { app: my-app } }
minDomains: 3 # zones with no matching pod count as domains (stable in v1.30)
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector: { matchLabels: { app: my-app } }
matchLabelKeys: [pod-template-hash] # count only this ReplicaSet, so a rollout ignores old pods (beta since v1.27)Topology spread replaces podAntiAffinity for “one replica per zone or node” and scales better, because anti-affinity is evaluated against every pod in the cluster. A hard spread with maxSkew: 1 and a zone that has no schedulable capacity leaves pods Pending with didn't match pod topology spread constraints; nodeTaintsPolicy: Honor and nodeAffinityPolicy: Honor (both beta since v1.26) make the scheduler exclude such nodes from the skew calculation.
Jobs and CronJobs#
A Job runs pods until completions succeed (default 1), parallelism at a time, and fails after backoffLimit failed pods (default 6) or activeDeadlineSeconds, whichever comes first. The pod’s restartPolicy must be Never or OnFailure; with OnFailure the retries hide inside one pod’s restartCount, so Never gives a clearer failure history. ttlSecondsAfterFinished deletes the Job and its pods after it finishes; without it, finished Jobs and their logs accumulate.
apiVersion: batch/v1
kind: Job
metadata: { name: migrate, namespace: my-namespace }
spec:
backoffLimit: 3
activeDeadlineSeconds: 900 # kill everything 15 minutes after the Job starts
ttlSecondsAfterFinished: 86400 # delete the Job and its pods a day after it finishes
podFailurePolicy: # stable in v1.31
rules:
- action: FailJob # do not retry a configuration error
onExitCodes: { containerName: migrate, operator: In, values: [2] }
- action: Ignore # preemption or a drain does not count against backoffLimit
onPodConditions: [{ type: DisruptionTarget }]
template:
spec:
restartPolicy: Never
containers:
- name: migrate
image: registry.example.com/my-app@sha256:<digest>
args: [migrate, --to, latest]apiVersion: batch/v1
kind: CronJob
metadata: { name: report, namespace: my-namespace }
spec:
schedule: "15 2 * * *"
timeZone: Australia/Melbourne # stable in v1.27; unset means the controller-manager's zone
concurrencyPolicy: Forbid # skip the run if the previous Job is still active; Replace kills it
startingDeadlineSeconds: 600 # give up on a run that could not start within 10 minutes
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3 # default 1 hides the failure you are looking for
jobTemplate:
spec:
backoffLimit: 2
template:
spec:
restartPolicy: OnFailure
containers:
- name: report
image: registry.example.com/report@sha256:<digest>backoffLimitPerIndex and successPolicy (both stable in v1.33) need completionMode: Indexed. The controller checks schedules every 10 s. A CronJob that misses more than 100 start times, typically after a long suspend or controller outage, stops scheduling and logs Too many missed start time (> 100); setting startingDeadlineSeconds bounds the window it counts over and is the fix.
kubectl create job report-now --from=cronjob/report -n my-namespace # run the template immediately
kubectl get jobs -n my-namespace -o wide # COMPLETIONS and DURATION per run
kubectl logs job/report-now -n my-namespace --all-containers # logs of the Job's pod
kubectl patch cronjob report -n my-namespace -p '{"spec":{"suspend":true}}' # pause; a running Job continues
kubectl delete jobs -n my-namespace --field-selector status.successful=1 # remove finished Jobs and their podsServices and networking#
A Service is a stable virtual IP plus a label selector. The EndpointSlice controller keeps a list of the selected pods’ IPs and readiness, and kube-proxy (or a replacement such as Cilium) programs the dataplane from it. No ready pods means no endpoints, which usually presents as connection refused or a timeout rather than an error message. The older Endpoints API is deprecated since v1.33; read EndpointSlices.
kubectl get svc,endpointslice -n my-namespace
kubectl get endpointslice -l kubernetes.io/service-name=my-app -n my-namespace -o yaml | grep -A3 addresses
kubectl run tmp --rm -it --image=nicolaka/netshoot -n my-namespace -- sh # curl, dig, tcpdump in-cluster; pod deleted on exit
kubectl port-forward svc/my-app 8080:80 -n my-namespace| Type | Behaviour |
|---|---|
ClusterIP | In-cluster virtual IP; the default |
NodePort | ClusterIP plus a port (30000-32767 by default) on every node |
LoadBalancer | NodePort plus an external load balancer provisioned by a controller |
ExternalName | DNS CNAME only, no proxying |
Headless (clusterIP: None) | DNS returns pod IPs directly; how StatefulSet members are addressed |
DNS names follow <service>.<namespace>.svc.cluster.local. Inside a pod, my-app resolves through the search domains in /etc/resolv.conf. Across namespaces, my-app.other-namespace is the shortest reliable form. The default ndots:5 makes external names try every search domain first; see DNS for the effect. For HTTP routing into the cluster, use Gateway API.
NetworkPolicies are allow-lists. A pod selected by no policy for a direction allows all traffic in that direction. Once any policy selects it for ingress or egress, everything in that direction not explicitly allowed is denied. When both the client’s egress and the server’s ingress are isolated, both sides must allow the flow. Enforcement depends on the CNI; a CNI without policy support accepts the objects and ignores them.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: my-app-allow }
spec:
podSelector: { matchLabels: { app: my-app } }
policyTypes: [Ingress, Egress]
ingress:
- from:
- podSelector: { matchLabels: { app: web } }
ports: [{ protocol: TCP, port: 8080 }]
egress:
- to: [{ namespaceSelector: { matchLabels: { kubernetes.io/metadata.name: kube-system } } }]
ports:
- { protocol: UDP, port: 53 } # without DNS egress, every name lookup fails
- { protocol: TCP, port: 53 }Storage#
A PVC is a request, a PV is the volume that satisfies it, and a StorageClass provisions PVs on demand. volumeBindingMode: WaitForFirstConsumer delays provisioning until a pod is scheduled, so the volume is created in the same zone as the node.
kubectl get pvc,pv -n my-namespace
kubectl get sc
kubectl describe pvc data-my-app-0 -n my-namespace # provisioning errors appear as eventsapiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: data }
spec:
accessModes: [ReadWriteOnce] # one node, not one pod
storageClassName: gp3
resources: { requests: { storage: 20Gi } }ReadWriteOnce allows many pods on the same node to mount the volume. Use ReadWriteOncePod when exactly one pod may mount it. Expansion works in place when the StorageClass sets allowVolumeExpansion: true; shrinking is not supported. A PVC stuck Terminating is still mounted by a pod; the kubernetes.io/pvc-protection finalizer holds it until that pod is gone.
ConfigMaps and Secrets#
ConfigMaps and Secrets use the same mechanism with different handling. Secret values are base64-encoded in the API (encoding, not encryption), encrypted in etcd only if the cluster configures encryption at rest, and mounted on tmpfs.
kubectl create configmap app-config --from-file=config.yaml --dry-run=client -o yaml > cm.yaml
kubectl create secret generic db --from-literal=password="$DB_PASSWORD" --dry-run=client -o yaml > secret.yaml # file holds the value; do not commit it
kubectl get secret db -o jsonpath='{.data.password}' | base64 -d # prints the secret to the terminal envFrom:
- configMapRef: { name: app-config }
env:
- name: DB_PASSWORD
valueFrom:
secretKeyRef: { name: db, key: password }
volumeMounts:
- { name: config, mountPath: /etc/app, readOnly: true }
volumes:
- name: config
configMap: { name: app-config }Mounted ConfigMaps and Secrets update in place after the kubelet sync period plus cache delay, typically within a minute or two. Environment variables and subPath mounts never update. Applications that read configuration once need kubectl rollout restart after a change.
RBAC and service accounts#
Every pod runs as a ServiceAccount. Its short-lived token is projected into the pod and used for API calls. RBAC binds Roles (namespaced) or ClusterRoles to users, groups or ServiceAccounts. Permissions are additive; there is no deny rule.
kubectl auth can-i --list -n my-namespace # my permissions here
kubectl auth can-i get secrets -n my-namespace --as system:serviceaccount:my-namespace:my-app
kubectl auth whoami # identity the API server sees
kubectl get rolebinding,clusterrolebinding -A -o wide | grep my-namespaceapiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata: { name: pod-reader, namespace: my-namespace }
rules:
- apiGroups: [""]
resources: ["pods", "pods/log"]
verbs: ["get", "list", "watch"]Set automountServiceAccountToken: false on workloads that never call the API. Anything that can read Secrets in a namespace, or create pods there, can obtain every ServiceAccount token in it.
Troubleshooting#
kubectl get events -A --sort-by=.lastTimestamp | tail -30
kubectl describe node node-1 | sed -n '/Conditions:/,/Events:/p'
kubectl top node; kubectl top pod -A --sort-by=cpu
kubectl get pod my-pod -o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'
kubectl debug node/node-1 -it --image=busybox # pod in host namespaces; host filesystem at /host; delete the pod afterwards
kubectl get --raw='/readyz?verbose' # API server health checks, one line each| Symptom | Likely cause | Check |
|---|---|---|
Pods Pending, nodes look idle | Requests exceed allocatable, or taints without matching tolerations | kubectl describe node, Allocated resources section |
| Service fails intermittently | Some replicas failing readiness | EndpointSlice membership over time |
| DNS slow or failing | CoreDNS unhealthy, or a NetworkPolicy blocking port 53 | kubectl get pods -n kube-system -l k8s-app=kube-dns, DNS |
exec works, curl from another pod does not | Process listening on 127.0.0.1 inside the pod | kubectl exec my-pod -- ss -ltn |
| Changes revert within minutes | A GitOps controller owns the resource | kubectl get <kind> <name> -o yaml, look for Argo CD or Flux labels and annotations |
Node NotReady | kubelet stopped, disk or PID pressure, or CNI failure | Node conditions, then journalctl -u kubelet on the node (systemd) |
forbidden from the API | RBAC missing for that verb, resource or namespace | kubectl auth can-i <verb> <resource> --as <subject> |
| Pod restarts with no error in logs | Liveness probe killing a slow container | kubectl describe pod, look for Liveness probe failed events |
Oneliners#
# Pods not Running or Succeeded, cluster-wide
kubectl get pods -A --field-selector 'status.phase!=Running,status.phase!=Succeeded'
# Top restart counts (first container of each pod)
kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}{"\t"}{.metadata.name}{"\t"}{.status.containerStatuses[0].restartCount}{"\n"}{end}' | sort -k3 -nr | head
# Every image running in the cluster, deduplicated
kubectl get pods -A -o jsonpath='{range .items[*]}{range .spec.containers[*]}{.image}{"\n"}{end}{end}' | sort -u
# Requests per pod
kubectl get pods -A -o custom-columns='NS:.metadata.namespace,POD:.metadata.name,CPU:.spec.containers[*].resources.requests.cpu,MEM:.spec.containers[*].resources.requests.memory'
# Requested versus allocatable on a node
kubectl describe node node-1 | awk '/Allocated resources/,/Events/'
# Pods on one node
kubectl get pods -A -o wide --field-selector spec.nodeName=node-1
# Which pods mount a given Secret as a volume
kubectl get pods -A -o json | jq -r '.items[] | select(.spec.volumes[]?.secret.secretName=="db") | "\(.metadata.namespace)/\(.metadata.name)"'
# Drain a node for maintenance (evicts pods, respects PodDisruptionBudgets), then allow scheduling again
kubectl drain node-1 --ignore-daemonsets --delete-emptydir-data --timeout=5m && kubectl uncordon node-1
# Watch rollouts of every Deployment in a namespace in parallel
kubectl get deploy -n my-namespace -o name | xargs -n1 -P0 kubectl rollout status -n my-namespace
# Object counts by resource, to find what is filling etcd (metric renamed in v1.34)
kubectl get --raw=/metrics | grep -E '^apiserver_(storage|resource)_objects' | sort -t' ' -k2 -nr | head
# Copy a file out of a pod without tar in the image
kubectl exec my-pod -n my-namespace -- cat /app/report.csv > report.csvTwo commands here need care:
# Print every key of a Secret in plain text. Output contains credentials.
kubectl get secret db -o go-template='{{range $k,$v := .data}}{{$k}}={{$v|base64decode}}{{"\n"}}{{end}}'Force deletion
kubectl delete pod my-pod --grace-period=0 --force removes the pod object without waiting for the kubelet to confirm the containers stopped. If the node is only partitioned, the old container can keep running alongside its replacement, which corrupts data for StatefulSets. It does not bypass finalizers; a pod with finalizers stays Terminating until they are removed.