# Kubernetes

> Inspect workloads with kubectl, understand the object model behind them and find why a pod is not running, ready or reachable.

Canonical: https://www.wiki.jodisand.me/kubernetes/
Reviewed: 2026-09-24
Related: [Helm](https://www.wiki.jodisand.me/helm/index.md), [Argo CD](https://www.wiki.jodisand.me/argocd/index.md), [Cilium](https://www.wiki.jodisand.me/cilium/index.md), [Gateway API](https://www.wiki.jodisand.me/gateway-api/index.md), [Docker](https://www.wiki.jodisand.me/docker/index.md)


## Cheatsheet

Commands assume `kubectl` 1.34 or later against a supported cluster. `kubectl top` needs metrics-server installed.

| Task | Command |
| --- | --- |
| Which cluster am I on | `kubectl config current-context` |
| Workloads in a namespace | `kubectl get all -n my-namespace` |
| Why is this pod unhappy | `kubectl describe pod my-pod -n my-namespace` |
| Events for one object | `kubectl events --for pod/my-pod -n my-namespace` |
| Events, newest last | `kubectl get events -n my-namespace --sort-by=.lastTimestamp` |
| Logs of the crashed container | `kubectl logs my-pod -c app --previous -n my-namespace` |
| Shell in a running pod | `kubectl exec -it my-pod -n my-namespace -- sh` |
| Debug a distroless pod | `kubectl debug -it my-pod --image=nicolaka/netshoot --target=app` |
| Port-forward a Service | `kubectl port-forward svc/my-app 8080:80 -n my-namespace` |
| Restart a Deployment | `kubectl rollout restart deploy/my-app -n my-namespace` |
| Watch a rollout | `kubectl rollout status deploy/my-app -n my-namespace` |
| Roll back | `kubectl rollout undo deploy/my-app -n my-namespace` |
| Scale | `kubectl scale deploy/my-app --replicas=5 -n my-namespace` |
| Resource usage | `kubectl top pod -n my-namespace --sort-by=memory` |
| Can I do this | `kubectl auth can-i delete pods -n my-namespace` |
| Validate against the live API | `kubectl apply -f my-app.yaml --dry-run=server` |
| Show what apply would change | `kubectl diff -f my-app.yaml` |
| Object as stored | `kubectl get pod my-pod -o yaml` |

## Start with a failing workload

Confirm the context first, then read the pod's state and its events. `describe` merges the object status with the events the scheduler and kubelet recorded, which is where the actual reason lives. Events expire after one hour by default, so check them early.

```sh
kubectl config current-context
kubectl get pods -n my-namespace -o wide
kubectl describe pod my-pod -n my-namespace | sed -n '/Events:/,$p'
kubectl logs my-pod -n my-namespace -c app --previous --tail 100
```

| Pod state | What it means | Next command |
| --- | --- | --- |
| `Pending` | No node fits: requests, taints, affinity, or an unbound PVC | `kubectl describe pod` and read the `FailedScheduling` event |
| `ContainerCreating` | Image pull, volume mount or CNI setup still in progress or failing | `kubectl describe pod`, then kubelet logs on the node |
| `ImagePullBackOff` / `ErrImagePull` | Wrong name or tag, private registry, missing `imagePullSecrets` | `kubectl events --for pod/my-pod`, `kubectl get sa default -o yaml` |
| `CrashLoopBackOff` | The process keeps exiting; restarts back off 10s, 20s, 40s up to 5 minutes | `kubectl logs --previous` |
| `CreateContainerConfigError` | A referenced ConfigMap, Secret or key does not exist | `kubectl describe pod`, then `kubectl get cm,secret` |
| `Running`, not `Ready` | Readiness probe failing | `kubectl describe pod`, check probe path and port |
| `OOMKilled` (exit 137) | Container memory exceeded `limits.memory` | `kubectl top pod`, raise the limit or fix the leak |
| `Terminating` forever | Finalizer not cleared, or the node is unreachable so the kubelet never confirms | `kubectl get pod my-pod -o jsonpath='{.metadata.finalizers}'`, `kubectl get node` |
| `Evicted` | Node pressure (memory, disk, PIDs) reclaimed it | `kubectl describe node`, read the conditions |

The crash backoff resets after the container runs for 10 minutes without failing. See [pod lifecycle](https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/) for the feature gates that shorten it.

> [!WARNING] Changes in a GitOps-managed cluster
> A reconciler such as [Argo CD](https://www.wiki.jodisand.me/argocd/) reverts a manual edit. Use direct commands for diagnosis only and apply lasting fixes through the repository that owns the resource.

## How the control plane reconciles state

The API server is the only component that writes to etcd. Controllers, the scheduler and kubelets watch the API for desired state and act until reality matches. Every fix is therefore "change the desired state and wait", never "make the change on the node".

| Component | Job |
| --- | --- |
| `kube-apiserver` | Authenticates, runs admission, validates and stores objects; the single write path |
| `etcd` | Consistent key-value store holding cluster state |
| `kube-scheduler` | Binds a pending pod to a node that satisfies its requests and constraints |
| `kube-controller-manager` | Reconciliation loops for Deployments, ReplicaSets, Jobs, nodes and EndpointSlices |
| `kubelet` | Runs the containers for pods bound to its node and reports status |
| `kube-proxy` or a CNI replacement | Programs Service load balancing on each node ([Cilium](https://www.wiki.jodisand.me/cilium/#kube-proxy-replacement) can replace it) |

A Deployment owns a ReplicaSet, which owns Pods. Changing the pod template creates a new ReplicaSet and scales the old one down. That indirection is why `kubectl rollout undo` works, and why deleting a pod with a bad image only brings back another pod with the same bad image.

## kubectl contexts, queries and applies

```sh
kubectl config get-contexts
kubectl config use-context prod
kubectl config set-context --current --namespace=my-namespace   # default namespace for this context

kubectl get deploy,sts,ds,job -A                          # workloads in every namespace
kubectl get pods -A -o wide --field-selector status.phase!=Running
kubectl get pod my-pod -o jsonpath='{.spec.containers[*].image}{"\n"}'
kubectl explain deployment.spec.strategy --recursive       # schema served by this cluster
kubectl api-resources --namespaced=true                    # kinds this cluster serves
kubectl diff -f manifest.yaml                              # what applying would change, exit 1 if different
kubectl apply -f manifest.yaml --server-side               # server tracks field ownership per manager
```

`kubectl get -o yaml` returns the object after defaulting and admission, not what you submitted. Client-side apply records what you sent in the `kubectl.kubernetes.io/last-applied-configuration` annotation. Server-side apply records ownership in `metadata.managedFields` instead, and reports a conflict when another manager owns a field you try to set. Add `--force-conflicts` only when you intend to take that field over.

## Workloads and rollouts

```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
spec:
  replicas: 3
  revisionHistoryLimit: 5
  strategy:
    type: RollingUpdate
    rollingUpdate: { maxSurge: 1, maxUnavailable: 0 }
  selector:
    matchLabels: { app: my-app }                 # immutable after creation
  template:
    metadata:
      labels: { app: my-app }
    spec:
      terminationGracePeriodSeconds: 45
      securityContext:
        runAsNonRoot: true
        seccompProfile: { type: RuntimeDefault }
      containers:
        - name: app
          image: registry.example.com/my-app@sha256:<digest>   # digest, not a moving tag
          ports: [{ containerPort: 8080 }]
          resources:
            requests: { cpu: 100m, memory: 256Mi }   # what the scheduler reserves
            limits: { memory: 512Mi }                # what the kernel enforces
          readinessProbe:
            httpGet: { path: /readyz, port: 8080 }
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /healthz, port: 8080 }
            initialDelaySeconds: 20
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities: { drop: ["ALL"] }
```

Requests drive scheduling and CPU weight. Limits drive throttling and OOM kills. A CPU limit throttles rather than kills, which shows up as latency, not restarts, so many teams set memory limits and leave CPU limits off.

Readiness removes a pod from Service endpoints. Liveness restarts the container. Pointing liveness at an endpoint that checks a database turns a database blip into restarts across every replica. Use a `startupProbe` for slow starters instead of a long `initialDelaySeconds`.

```sh
kubectl rollout status deploy/my-app -n my-namespace --timeout=5m
kubectl rollout history deploy/my-app -n my-namespace
kubectl rollout undo deploy/my-app --to-revision=3 -n my-namespace
kubectl rollout restart deploy/my-app -n my-namespace    # new pods, same spec: picks up rotated Secrets in env vars
kubectl scale deploy/my-app --replicas=0 -n my-namespace  # stops every pod; the Deployment remains
```

| Kind | Use it for |
| --- | --- |
| `Deployment` | Stateless replicas, rolling updates, rollback |
| `StatefulSet` | Stable network identity and per-replica storage; ordered, slower updates |
| `DaemonSet` | One pod per node: node agents, CNI, log shippers |
| `Job` / `CronJob` | Run to completion, bounded by `backoffLimit` and `activeDeadlineSeconds` |

Native sidecars are init containers with `restartPolicy: Always`. They start before the app containers and stop after them, which fixes Jobs that never complete because a proxy sidecar keeps running. They are stable since v1.33.

A StatefulSet's PVCs survive deletion of the StatefulSet by default. Removing them is a separate `kubectl delete pvc`, which destroys the data when the reclaim policy is `Delete`.

## Probes and graceful shutdown

Probe defaults are `initialDelaySeconds: 0`, `periodSeconds: 10`, `timeoutSeconds: 1`, `failureThreshold: 3` and `successThreshold: 1`. A startup probe suspends the other two until it succeeds once, so `failureThreshold * periodSeconds` is the longest start you permit, and liveness and readiness then run with short thresholds of their own. The one-second timeout is the usual surprise: a handler that takes 1.5 s under load counts as a failure. This extends the probe rules in [Workloads and rollouts](#workloads-and-rollouts).

```yaml
    startupProbe:
      httpGet: { path: /healthz, port: 8080 }
      periodSeconds: 5
      failureThreshold: 60             # up to 5 minutes to start; liveness and readiness wait
    readinessProbe:
      httpGet: { path: /readyz, port: 8080 }
      periodSeconds: 5
      timeoutSeconds: 2
      failureThreshold: 2              # out of the endpoints after 10 s of failures
    livenessProbe:
      grpc: { port: 9090 }             # gRPC health protocol; kubelet speaks it natively
      periodSeconds: 20
      timeoutSeconds: 5
    lifecycle:
      preStop:
        sleep: { seconds: 5 }          # stable in v1.34; keeps serving while endpoint removal propagates
```

Deleting a pod starts two things in parallel: the kubelet begins the local shutdown, and the EndpointSlice controller marks the endpoint `terminating` with `ready: false`, which kube-proxy and external load balancers act on after their own delay. The kubelet runs the `preStop` hook to completion, then has the runtime send `SIGTERM` to PID 1 of each container, and when `terminationGracePeriodSeconds` (default 30, counted from the start of the hook) expires it sends `SIGKILL`. Hook and shutdown share that one budget: a 25 s hook plus a 10 s shutdown under a 30 s grace period ends in `SIGKILL`. A short `preStop` sleep covers the propagation gap so requests are not routed to a container that has already stopped listening. Containers receive `SIGTERM` in arbitrary order unless the helper is a native sidecar, which stops after the main containers.

```sh
kubectl exec my-pod -n my-namespace -- cat /proc/1/cmdline | tr '\0' ' '   # what PID 1 is; sh -c 'app' does not forward SIGTERM
kubectl get pod my-pod -n my-namespace -o jsonpath='{.metadata.deletionTimestamp} {.metadata.deletionGracePeriodSeconds}{"\n"}'
```

## Resource requests, limits and QoS classes

The API server derives a QoS class from the containers' requests and limits, and the kubelet evicts in reverse order of that class under node pressure. `Guaranteed` requires every container to set CPU and memory limits equal to its requests, `Burstable` is any pod with at least one request or limit, and `BestEffort` has none. `Guaranteed` pods are the only ones the static CPU manager policy can pin to exclusive cores. The Deployment in [Workloads and rollouts](#workloads-and-rollouts) is `Burstable`, which is the normal choice for services.

```sh
kubectl get pods -n my-namespace -o custom-columns='POD:.metadata.name,QOS:.status.qosClass,NODE:.spec.nodeName'
kubectl set resources deploy/my-app -n my-namespace -c app --requests=cpu=100m,memory=256Mi --limits=memory=512Mi   # changes the template: rollout
kubectl describe quota -n my-namespace                    # ResourceQuota usage; a full quota rejects new pods at admission
kubectl get limitrange -n my-namespace -o yaml            # defaults injected into containers that set nothing
```

A `LimitRange` with `default` and `defaultRequest` stops `BestEffort` pods appearing in a namespace, and a `ResourceQuota` on `requests.cpu` makes any pod without that request fail admission with `must specify requests.cpu`. Memory requests set the container's OOM score adjustment, so a container far above its request is killed first when the node runs out, before it reaches its own limit.

## PodDisruptionBudgets

A PDB limits voluntary disruptions: `kubectl drain`, the eviction API, cluster autoscaler scale-down and managed node upgrades. It does nothing for crashes, OOM kills or hardware failure. `minAvailable` and `maxUnavailable` take a count or a percentage, only one may be set, and `maxUnavailable` requires that every selected pod share one controller. `minAvailable` equal to the replica count blocks every drain forever, which is the most common way to stall a node upgrade.

```yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: my-app, namespace: my-namespace }
spec:
  maxUnavailable: 1                          # or "25%"; percentages round up
  unhealthyPodEvictionPolicy: AlwaysAllow    # stable in v1.31: pods that are not Ready may always be evicted
  selector:
    matchLabels: { app: my-app }
```

The default `IfHealthyBudget` refuses to evict a crash-looping pod while the budget is unmet, so a broken deployment blocks a drain; `AlwaysAllow` is the right setting for almost every service. Check what a drain will run into before starting it:

```sh
kubectl get pdb -A                                                                   # ALLOWED DISRUPTIONS 0 means drain waits
kubectl get pdb my-app -n my-namespace -o jsonpath='{.status.disruptionsAllowed}{"\n"}'
kubectl drain node-1 --ignore-daemonsets --delete-emptydir-data --dry-run=server     # lists the pods it would evict
```

## Horizontal Pod Autoscaler

An `autoscaling/v2` HPA sets `replicas` through the target's scale subresource from `ceil(currentReplicas * currentMetric / targetMetric)`, evaluated every 15 s and ignored when the ratio is within 10% of 1. `Utilization` targets are percentages of the container requests, so a pod whose container has no request for that metric is skipped and the HPA reports `FailedGetResourceMetric`. Remove `replicas` from the Deployment manifest once an HPA owns it: every apply resets the count and the HPA scales it back, which shows up as a rollout on each GitOps sync.

```yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: my-app, namespace: my-namespace }
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: my-app }
  minReplicas: 2
  maxReplicas: 20
  metrics:
    - type: ContainerResource              # one container, not the pod total: ignores the sidecar
      containerResource:
        name: cpu
        container: app
        target: { type: Utilization, averageUtilization: 70 }
    - type: Resource
      resource:
        name: memory
        target: { type: AverageValue, averageValue: 400Mi }
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0        # act on the latest recommendation immediately
      policies:
        - { type: Percent, value: 100, periodSeconds: 15 }   # the default: double, or add 4 pods, per 15 s
    scaleDown:
      stabilizationWindowSeconds: 300      # the default: use the highest recommendation of the last 5 minutes
      policies:
        - { type: Pods, value: 2, periodSeconds: 60 }
      selectPolicy: Min                    # Max (default), Min, or Disabled to never scale down
```

```sh
kubectl autoscale deploy/my-app -n my-namespace --min=2 --max=10 --cpu=70%   # --cpu and --memory take a percentage or a quantity such as 500m
kubectl get hpa -n my-namespace                                                # TARGETS shows <unknown> until metrics arrive
kubectl describe hpa my-app -n my-namespace | sed -n '/Conditions:/,$p'        # ScalingActive False says why nothing happens
```

With several metrics the HPA takes the largest desired replica count. Memory rarely drops when load does because most runtimes keep freed heap, so a memory target mostly scales up.

## Node affinity, taints and topology spread

The scheduler filters out nodes that fail hard constraints, then scores the rest. `nodeSelector` and `requiredDuringSchedulingIgnoredDuringExecution` filter; `preferredDuringSchedulingIgnoredDuringExecution` adds a weighted score. `IgnoredDuringExecution` means a running pod stays put when node labels change. Taints are the reverse: a node repels pods that do not tolerate the taint. `NoSchedule` affects new pods, `PreferNoSchedule` is the soft form, and `NoExecute` also evicts running pods after `tolerationSeconds`. The control plane taints unhealthy nodes with `node.kubernetes.io/not-ready` and `node.kubernetes.io/unreachable` as `NoExecute`, and every pod gets a default 300 s toleration for both, which is the five-minute wait before pods on a dead node are replaced.

```sh
kubectl label nodes node-1 workload=batch                        # for selectors and affinity
kubectl taint nodes node-1 dedicated=batch:NoSchedule            # repel pods without the toleration
kubectl taint nodes node-1 dedicated=batch:NoSchedule-           # trailing dash removes it
kubectl cordon node-1                                            # adds node.kubernetes.io/unschedulable:NoSchedule; running pods stay
kubectl get nodes -o custom-columns='NODE:.metadata.name,TAINTS:.spec.taints[*].key,ZONE:.metadata.labels.topology\.kubernetes\.io/zone'
```

```yaml
spec:
  tolerations:
    - { key: dedicated, operator: Equal, value: batch, effect: NoSchedule }
    - key: node.kubernetes.io/unreachable
      operator: Exists
      effect: NoExecute
      tolerationSeconds: 60                 # replace this pod after 1 minute instead of 5
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:                  # terms are ORed; expressions within a term are ANDed
          - matchExpressions:
              - { key: kubernetes.io/arch, operator: In, values: [amd64, arm64] }
              - { key: node.kubernetes.io/instance-type, operator: NotIn, values: [t3.micro] }
      preferredDuringSchedulingIgnoredDuringExecution:
        - weight: 80
          preference:
            matchExpressions: [{ key: workload, operator: In, values: [batch] }]
  topologySpreadConstraints:
    - maxSkew: 1
      topologyKey: topology.kubernetes.io/zone
      whenUnsatisfiable: DoNotSchedule      # hard; ScheduleAnyway only affects scoring
      labelSelector: { matchLabels: { app: my-app } }
      minDomains: 3                         # zones with no matching pod count as domains (stable in v1.30)
    - maxSkew: 1
      topologyKey: kubernetes.io/hostname
      whenUnsatisfiable: ScheduleAnyway
      labelSelector: { matchLabels: { app: my-app } }
      matchLabelKeys: [pod-template-hash]   # count only this ReplicaSet, so a rollout ignores old pods (beta since v1.27)
```

Topology spread replaces `podAntiAffinity` for "one replica per zone or node" and scales better, because anti-affinity is evaluated against every pod in the cluster. A hard spread with `maxSkew: 1` and a zone that has no schedulable capacity leaves pods `Pending` with `didn't match pod topology spread constraints`; `nodeTaintsPolicy: Honor` and `nodeAffinityPolicy: Honor` (both beta since v1.26) make the scheduler exclude such nodes from the skew calculation.

## Jobs and CronJobs

A Job runs pods until `completions` succeed (default 1), `parallelism` at a time, and fails after `backoffLimit` failed pods (default 6) or `activeDeadlineSeconds`, whichever comes first. The pod's `restartPolicy` must be `Never` or `OnFailure`; with `OnFailure` the retries hide inside one pod's `restartCount`, so `Never` gives a clearer failure history. `ttlSecondsAfterFinished` deletes the Job and its pods after it finishes; without it, finished Jobs and their logs accumulate.

```yaml
apiVersion: batch/v1
kind: Job
metadata: { name: migrate, namespace: my-namespace }
spec:
  backoffLimit: 3
  activeDeadlineSeconds: 900               # kill everything 15 minutes after the Job starts
  ttlSecondsAfterFinished: 86400           # delete the Job and its pods a day after it finishes
  podFailurePolicy:                        # stable in v1.31
    rules:
      - action: FailJob                    # do not retry a configuration error
        onExitCodes: { containerName: migrate, operator: In, values: [2] }
      - action: Ignore                     # preemption or a drain does not count against backoffLimit
        onPodConditions: [{ type: DisruptionTarget }]
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: migrate
          image: registry.example.com/my-app@sha256:<digest>
          args: [migrate, --to, latest]
```

```yaml
apiVersion: batch/v1
kind: CronJob
metadata: { name: report, namespace: my-namespace }
spec:
  schedule: "15 2 * * *"
  timeZone: Australia/Melbourne            # stable in v1.27; unset means the controller-manager's zone
  concurrencyPolicy: Forbid                # skip the run if the previous Job is still active; Replace kills it
  startingDeadlineSeconds: 600             # give up on a run that could not start within 10 minutes
  successfulJobsHistoryLimit: 3
  failedJobsHistoryLimit: 3                # default 1 hides the failure you are looking for
  jobTemplate:
    spec:
      backoffLimit: 2
      template:
        spec:
          restartPolicy: OnFailure
          containers:
            - name: report
              image: registry.example.com/report@sha256:<digest>
```

`backoffLimitPerIndex` and `successPolicy` (both stable in v1.33) need `completionMode: Indexed`. The controller checks schedules every 10 s. A CronJob that misses more than 100 start times, typically after a long `suspend` or controller outage, stops scheduling and logs `Too many missed start time (> 100)`; setting `startingDeadlineSeconds` bounds the window it counts over and is the fix.

```sh
kubectl create job report-now --from=cronjob/report -n my-namespace         # run the template immediately
kubectl get jobs -n my-namespace -o wide                                   # COMPLETIONS and DURATION per run
kubectl logs job/report-now -n my-namespace --all-containers                # logs of the Job's pod
kubectl patch cronjob report -n my-namespace -p '{"spec":{"suspend":true}}' # pause; a running Job continues
kubectl delete jobs -n my-namespace --field-selector status.successful=1    # remove finished Jobs and their pods
```

## Services and networking

A Service is a stable virtual IP plus a label selector. The EndpointSlice controller keeps a list of the selected pods' IPs and readiness, and kube-proxy (or a replacement such as Cilium) programs the dataplane from it. No ready pods means no endpoints, which usually presents as connection refused or a timeout rather than an error message. The older `Endpoints` API is deprecated since v1.33; read EndpointSlices.

```sh
kubectl get svc,endpointslice -n my-namespace
kubectl get endpointslice -l kubernetes.io/service-name=my-app -n my-namespace -o yaml | grep -A3 addresses
kubectl run tmp --rm -it --image=nicolaka/netshoot -n my-namespace -- sh   # curl, dig, tcpdump in-cluster; pod deleted on exit
kubectl port-forward svc/my-app 8080:80 -n my-namespace
```

| Type | Behaviour |
| --- | --- |
| `ClusterIP` | In-cluster virtual IP; the default |
| `NodePort` | ClusterIP plus a port (30000-32767 by default) on every node |
| `LoadBalancer` | NodePort plus an external load balancer provisioned by a controller |
| `ExternalName` | DNS CNAME only, no proxying |
| Headless (`clusterIP: None`) | DNS returns pod IPs directly; how StatefulSet members are addressed |

DNS names follow `<service>.<namespace>.svc.cluster.local`. Inside a pod, `my-app` resolves through the `search` domains in `/etc/resolv.conf`. Across namespaces, `my-app.other-namespace` is the shortest reliable form. The default `ndots:5` makes external names try every search domain first; see [DNS](https://www.wiki.jodisand.me/dns/) for the effect. For HTTP routing into the cluster, use [Gateway API](https://www.wiki.jodisand.me/gateway-api/).

NetworkPolicies are allow-lists. A pod selected by no policy for a direction allows all traffic in that direction. Once any policy selects it for ingress or egress, everything in that direction not explicitly allowed is denied. When both the client's egress and the server's ingress are isolated, both sides must allow the flow. Enforcement depends on the CNI; a CNI without policy support accepts the objects and ignores them.

```yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: my-app-allow }
spec:
  podSelector: { matchLabels: { app: my-app } }
  policyTypes: [Ingress, Egress]
  ingress:
    - from:
        - podSelector: { matchLabels: { app: web } }
      ports: [{ protocol: TCP, port: 8080 }]
  egress:
    - to: [{ namespaceSelector: { matchLabels: { kubernetes.io/metadata.name: kube-system } } }]
      ports:
        - { protocol: UDP, port: 53 }   # without DNS egress, every name lookup fails
        - { protocol: TCP, port: 53 }
```

## Storage

A PVC is a request, a PV is the volume that satisfies it, and a StorageClass provisions PVs on demand. `volumeBindingMode: WaitForFirstConsumer` delays provisioning until a pod is scheduled, so the volume is created in the same zone as the node.

```sh
kubectl get pvc,pv -n my-namespace
kubectl get sc
kubectl describe pvc data-my-app-0 -n my-namespace   # provisioning errors appear as events
```

```yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: data }
spec:
  accessModes: [ReadWriteOnce]              # one node, not one pod
  storageClassName: gp3
  resources: { requests: { storage: 20Gi } }
```

`ReadWriteOnce` allows many pods on the same node to mount the volume. Use `ReadWriteOncePod` when exactly one pod may mount it. Expansion works in place when the StorageClass sets `allowVolumeExpansion: true`; shrinking is not supported. A PVC stuck `Terminating` is still mounted by a pod; the `kubernetes.io/pvc-protection` finalizer holds it until that pod is gone.

## ConfigMaps and Secrets

ConfigMaps and Secrets use the same mechanism with different handling. Secret values are base64-encoded in the API (encoding, not encryption), encrypted in etcd only if the cluster configures encryption at rest, and mounted on tmpfs.

```sh
kubectl create configmap app-config --from-file=config.yaml --dry-run=client -o yaml > cm.yaml
kubectl create secret generic db --from-literal=password="$DB_PASSWORD" --dry-run=client -o yaml > secret.yaml   # file holds the value; do not commit it
kubectl get secret db -o jsonpath='{.data.password}' | base64 -d   # prints the secret to the terminal
```

```yaml
    envFrom:
      - configMapRef: { name: app-config }
    env:
      - name: DB_PASSWORD
        valueFrom:
          secretKeyRef: { name: db, key: password }
    volumeMounts:
      - { name: config, mountPath: /etc/app, readOnly: true }
  volumes:
    - name: config
      configMap: { name: app-config }
```

Mounted ConfigMaps and Secrets update in place after the kubelet sync period plus cache delay, typically within a minute or two. Environment variables and `subPath` mounts never update. Applications that read configuration once need `kubectl rollout restart` after a change.

## RBAC and service accounts

Every pod runs as a ServiceAccount. Its short-lived token is projected into the pod and used for API calls. RBAC binds Roles (namespaced) or ClusterRoles to users, groups or ServiceAccounts. Permissions are additive; there is no deny rule.

```sh
kubectl auth can-i --list -n my-namespace                                          # my permissions here
kubectl auth can-i get secrets -n my-namespace --as system:serviceaccount:my-namespace:my-app
kubectl auth whoami                                                                 # identity the API server sees
kubectl get rolebinding,clusterrolebinding -A -o wide | grep my-namespace
```

```yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata: { name: pod-reader, namespace: my-namespace }
rules:
  - apiGroups: [""]
    resources: ["pods", "pods/log"]
    verbs: ["get", "list", "watch"]
```

Set `automountServiceAccountToken: false` on workloads that never call the API. Anything that can read Secrets in a namespace, or create pods there, can obtain every ServiceAccount token in it.

## Troubleshooting

```sh
kubectl get events -A --sort-by=.lastTimestamp | tail -30
kubectl describe node node-1 | sed -n '/Conditions:/,/Events:/p'
kubectl top node; kubectl top pod -A --sort-by=cpu
kubectl get pod my-pod -o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'
kubectl debug node/node-1 -it --image=busybox     # pod in host namespaces; host filesystem at /host; delete the pod afterwards
kubectl get --raw='/readyz?verbose'               # API server health checks, one line each
```

| Symptom | Likely cause | Check |
| --- | --- | --- |
| Pods `Pending`, nodes look idle | Requests exceed allocatable, or taints without matching tolerations | `kubectl describe node`, `Allocated resources` section |
| Service fails intermittently | Some replicas failing readiness | EndpointSlice membership over time |
| DNS slow or failing | CoreDNS unhealthy, or a NetworkPolicy blocking port 53 | `kubectl get pods -n kube-system -l k8s-app=kube-dns`, [DNS](https://www.wiki.jodisand.me/dns/#a-name-that-will-not-resolve) |
| `exec` works, `curl` from another pod does not | Process listening on `127.0.0.1` inside the pod | `kubectl exec my-pod -- ss -ltn` |
| Changes revert within minutes | A GitOps controller owns the resource | `kubectl get <kind> <name> -o yaml`, look for Argo CD or Flux labels and annotations |
| Node `NotReady` | kubelet stopped, disk or PID pressure, or CNI failure | Node conditions, then `journalctl -u kubelet` on the node ([systemd](https://www.wiki.jodisand.me/systemd/#a-failing-service)) |
| `forbidden` from the API | RBAC missing for that verb, resource or namespace | `kubectl auth can-i <verb> <resource> --as <subject>` |
| Pod restarts with no error in logs | Liveness probe killing a slow container | `kubectl describe pod`, look for `Liveness probe failed` events |

## Oneliners

```sh
# Pods not Running or Succeeded, cluster-wide
kubectl get pods -A --field-selector 'status.phase!=Running,status.phase!=Succeeded'

# Top restart counts (first container of each pod)
kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}{"\t"}{.metadata.name}{"\t"}{.status.containerStatuses[0].restartCount}{"\n"}{end}' | sort -k3 -nr | head

# Every image running in the cluster, deduplicated
kubectl get pods -A -o jsonpath='{range .items[*]}{range .spec.containers[*]}{.image}{"\n"}{end}{end}' | sort -u

# Requests per pod
kubectl get pods -A -o custom-columns='NS:.metadata.namespace,POD:.metadata.name,CPU:.spec.containers[*].resources.requests.cpu,MEM:.spec.containers[*].resources.requests.memory'

# Requested versus allocatable on a node
kubectl describe node node-1 | awk '/Allocated resources/,/Events/'

# Pods on one node
kubectl get pods -A -o wide --field-selector spec.nodeName=node-1

# Which pods mount a given Secret as a volume
kubectl get pods -A -o json | jq -r '.items[] | select(.spec.volumes[]?.secret.secretName=="db") | "\(.metadata.namespace)/\(.metadata.name)"'

# Drain a node for maintenance (evicts pods, respects PodDisruptionBudgets), then allow scheduling again
kubectl drain node-1 --ignore-daemonsets --delete-emptydir-data --timeout=5m && kubectl uncordon node-1

# Watch rollouts of every Deployment in a namespace in parallel
kubectl get deploy -n my-namespace -o name | xargs -n1 -P0 kubectl rollout status -n my-namespace

# Object counts by resource, to find what is filling etcd (metric renamed in v1.34)
kubectl get --raw=/metrics | grep -E '^apiserver_(storage|resource)_objects' | sort -t' ' -k2 -nr | head

# Copy a file out of a pod without tar in the image
kubectl exec my-pod -n my-namespace -- cat /app/report.csv > report.csv
```

Two commands here need care:

```sh
# Print every key of a Secret in plain text. Output contains credentials.
kubectl get secret db -o go-template='{{range $k,$v := .data}}{{$k}}={{$v|base64decode}}{{"\n"}}{{end}}'
```

> [!WARNING] Force deletion
> `kubectl delete pod my-pod --grace-period=0 --force` removes the pod object without waiting for the kubelet to confirm the containers stopped. If the node is only partitioned, the old container can keep running alongside its replacement, which corrupts data for StatefulSets. It does not bypass finalizers; a pod with finalizers stays `Terminating` until they are removed.


