# Docker

> Run, build, inspect and clean up containers, images, volumes and networks on a single Docker daemon, and diagnose containers that exit or cannot connect.

Canonical: https://www.wiki.jodisand.me/docker/
Reviewed: 2026-09-24
Related: [Docker Compose](https://www.wiki.jodisand.me/docker-compose/index.md), [Kubernetes](https://www.wiki.jodisand.me/kubernetes/index.md), [Linux performance](https://www.wiki.jodisand.me/linux-performance/index.md), [iproute2](https://www.wiki.jodisand.me/iproute2/index.md)


## Cheatsheet

| Task | Command |
| --- | --- |
| Which daemon the CLI talks to | `docker context show` |
| All containers, including stopped | `docker ps -a` |
| Last 100 log lines, then follow | `docker logs -f --tail 100 my-app` |
| Shell inside a running container | `docker exec -it my-app sh` |
| Throwaway container, removed on exit | `docker run --rm -it alpine:3.24 sh` |
| Publish a port on localhost only | `docker run -p 127.0.0.1:8080:80 nginx:1.30-alpine` |
| Why it exited | `docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}} {{.State.Error}}' my-app` |
| Disk used by images, containers, volumes, cache | `docker system df` |
| Resolve a name from inside a container | `docker exec my-app getent hosts db` |
| Copy a file out of a container | `docker cp my-app:/app/config.yaml .` |
| Build and tag an image | `docker build -t my-app:1.0 .` |
| Layer sizes of an image | `docker history my-app:1.0` |

For multi-service stacks defined in a file, see [Docker Compose](https://www.wiki.jodisand.me/docker-compose/). For the same concepts under an orchestrator, see [Kubernetes](https://www.wiki.jodisand.me/kubernetes/).

## Start with a container problem

A container is an ordinary Linux process with namespaces (its own view of PIDs, network, mounts) and cgroups (resource limits) applied. It stops when its PID 1 exits, so a "crashed" container is almost always a process that returned or was killed, not Docker losing it.

```sh
docker context show                 # which daemon the CLI targets (local socket, remote host, Desktop VM)
docker ps -a                        # status, exit code, published ports
docker logs --tail 100 my-app       # stdout/stderr of PID 1 only
docker inspect -f '{{.State.Status}} {{.State.ExitCode}} {{.State.OOMKilled}}' my-app
```

```text
exited 137 true
```

| Symptom | Cause to check first |
| --- | --- |
| `Exited (0)` immediately | The command finished. `CMD` runs a one-shot command or a daemon that forks into the background |
| `Exited (137)` | SIGKILL. `OOMKilled: true` means the memory limit, otherwise `docker kill` or `docker stop` timing out |
| `Exited (143)` | SIGTERM handled and the process exited, usually a normal `docker stop` |
| `Exited (1)` with empty logs | The application logs to a file inside the container, not stdout |
| Restarting in a loop | Crash on start; `docker logs` shows each attempt. Restart policies back off, they do not stop |
| Port published but connection refused | Process bound to `127.0.0.1` inside the container instead of `0.0.0.0` |
| Container name does not resolve | Containers are on different networks, or on the default `bridge` network, which has no DNS |
| Data gone after `docker rm` | Writes went to the container layer, not a volume or bind mount |
| `permission denied` on a bind mount | Host UID/GID or SELinux label. On Fedora/RHEL add `:Z` (private) or `:z` (shared) to relabel |

## Containers

`docker run` is `create` plus `start`. Flags that shape the sandbox (network, mounts, user, capabilities) are fixed at create time, so changing them means removing and recreating the container. Resource limits and restart policy are the exception: `docker update` changes them in place.

```sh
docker run -d --name my-app \
  -p 127.0.0.1:8080:80 \
  -e DB_HOST=db \
  -v my-data:/app/data \
  --memory 512m --cpus 1.5 \
  --restart unless-stopped \
  nginx:1.30-alpine

docker stop my-app                  # SIGTERM to PID 1, SIGKILL after 10 s
docker stop -t 30 my-app            # allow 30 s to drain
docker start my-app                 # same filesystem and config, new process
docker rm -f my-app                 # kill and remove; the container layer is lost
docker update --memory 1g --restart on-failure:5 my-app
```

| Flag | Effect |
| --- | --- |
| `-p 8080:80` | Host port 8080 on every interface to container port 80 |
| `-p 127.0.0.1:8080:80` | Same, reachable only from the host |
| `--network my-net` | Join a user-defined network, which gives DNS by container name |
| `--restart unless-stopped` | Restart whenever it exits, and after a daemon restart, unless someone stopped it |
| `--restart on-failure:5` | Restart only on non-zero exit, at most 5 times; not after a daemon restart |
| `--user 10001:10001` | Run as that UID/GID, overriding the image's `USER` |
| `--read-only --tmpfs /tmp` | Read-only root filesystem with a writable scratch directory |
| `--cap-drop ALL --cap-add NET_BIND_SERVICE` | Drop every capability, add back only what the process needs |
| `--init` | Run a minimal init as PID 1 that forwards signals and reaps zombies |

A restart policy only takes effect after the container has run for at least 10 seconds, which stops a container that fails at start from spinning tightly. See [restart policies](https://docs.docker.com/engine/containers/start-containers-automatically/).

## Images and layers

An image is an ordered stack of read-only layers plus a config (entrypoint, env, user). A running container adds one writable copy-on-write layer on top. Each Dockerfile instruction that changes the filesystem adds a layer, and the build cache reuses a layer only while its instruction and inputs are unchanged, so instruction order decides rebuild time.

```sh
docker pull nginx:1.30-alpine
docker image ls
docker history my-app:1.0                                  # per-layer size and the instruction that made it
docker image inspect -f '{{.Config.Entrypoint}} {{.Config.Cmd}} {{.Config.User}}' my-app:1.0
docker tag my-app:1.0 registry.example.com/my-app:1.0
docker push registry.example.com/my-app:1.0
docker buildx imagetools inspect nginx:1.30-alpine         # digest and platforms without pulling
```

A tag is a mutable pointer. Pull by digest (`nginx@sha256:<digest>`) when a deployment must be reproducible.

## Dockerfile

Put instructions that rarely change first and those that change every commit last, so dependency layers stay cached. Copying the dependency manifest before the source (`COPY go.mod go.sum ./` then `COPY . .`) is the main technique.

```dockerfile
FROM golang:1.26-alpine AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download                          # cached until the manifests change
COPY . .
RUN CGO_ENABLED=0 go build -trimpath -o /out/server .

FROM alpine:3.24
RUN apk add --no-cache ca-certificates && adduser -S -u 10001 app
COPY --from=build /out/server /usr/local/bin/server
USER app
EXPOSE 8080
HEALTHCHECK --interval=30s --timeout=3s --retries=3 CMD wget -qO- http://127.0.0.1:8080/health || exit 1
ENTRYPOINT ["server"]
CMD ["--port", "8080"]
```

| Instruction | Behaviour |
| --- | --- |
| `ENTRYPOINT` vs `CMD` | `ENTRYPOINT` is the executable; `CMD` supplies default arguments, replaced by anything after the image name in `docker run` |
| Exec form `["x", "y"]` | No shell. The process is PID 1 and receives SIGTERM directly |
| Shell form `x y` | Wrapped in `/bin/sh -c`, so the shell is PID 1 and may not forward signals; `docker stop` then waits out the timeout |
| `EXPOSE` | Metadata only; it publishes nothing |
| `ARG` vs `ENV` | `ARG` exists during the build; `ENV` persists into the running container |
| `RUN apt-get update && apt-get install -y ...` | Keep both in one `RUN`, or a cached `update` layer feeds stale indexes to a later `install` |
| `COPY --from=build` | Copies artefacts out of an earlier stage and leaves the toolchain behind |
| `.dockerignore` | Excludes files from the build context; without it `.git` and local secrets are sent to the builder |

> [!WARNING] Build arguments are not secret
> Values passed through `ARG` or `ENV` are readable in the image history and config. Use a BuildKit secret mount, which is never written to a layer:
> `RUN --mount=type=secret,id=npm_token NPM_TOKEN="$(cat /run/secrets/npm_token)" npm ci` with `docker build --secret id=npm_token,env=NPM_TOKEN .`

## Multi-stage builds

Each `FROM` starts a stage. Only the last stage (or the one named with `--target`) becomes the image; earlier stages exist to produce files that later stages `COPY --from`. BuildKit builds stages in parallel where dependencies allow and skips stages nothing consumes, so a Dockerfile can carry a `test` or `lint` stage that costs nothing in a normal build.

```dockerfile
# syntax=docker/dockerfile:1
FROM --platform=$BUILDPLATFORM golang:1.26-alpine AS base   # run the compiler on the build host's architecture
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download

FROM base AS test
COPY . .
RUN go vet ./... && go test ./...

FROM base AS build
ARG TARGETOS TARGETARCH                        # set automatically by buildx per platform
COPY . .
RUN CGO_ENABLED=0 GOOS=$TARGETOS GOARCH=$TARGETARCH go build -trimpath -ldflags='-s -w' -o /out/server .

FROM gcr.io/distroless/static-debian12:nonroot AS runtime
COPY --from=build --chown=nonroot:nonroot /out/server /server
ENTRYPOINT ["/server"]

FROM runtime AS debug                          # same binary plus a shell, for docker exec
COPY --from=busybox:1.37-uclibc /bin/busybox /bin/busybox
```

```sh
docker build --target test .                        # run the tests stage only; fails the build on failure
docker build --target debug -t my-app:debug .       # image with a shell
docker build -t my-app:1.0 .                        # last stage: runtime
docker build --build-arg GOFLAGS=-mod=vendor .      # ARG values; unset ARGs take the Dockerfile default
docker buildx build --platform linux/amd64,linux/arm64 -t my-app:1.0 --push .   # cross-compile without QEMU thanks to $BUILDPLATFORM
docker build --no-cache-filter build .              # rebuild one stage, keep the cache for the others
docker build --progress=plain . 2>&1 | tail -50     # full RUN output instead of the collapsed view
```

`COPY --from` also accepts an image name, which is how a single binary from another image (a CLI, a CA bundle, `busybox`) gets into a distroless runtime. `COPY --link` copies into an independent layer that does not depend on earlier ones, so a base image update does not invalidate it; `COPY --chmod=755` sets permissions without a separate `RUN chmod`, which would duplicate the file into another layer. `ADD --checksum=sha256:... https://...` verifies a download.

## BuildKit cache mounts

The layer cache is all-or-nothing per instruction: a change to `go.sum` reruns `go mod download` from scratch. A cache mount is a persistent directory attached to a `RUN` only for its duration; it is never part of the image, survives across builds on the same builder, and makes package managers incremental.

```dockerfile
# syntax=docker/dockerfile:1
RUN --mount=type=cache,target=/go/pkg/mod \
    --mount=type=cache,target=/root/.cache/go-build \
    go build -o /out/server .

RUN --mount=type=cache,target=/var/cache/apt,sharing=locked \
    --mount=type=cache,target=/var/lib/apt,sharing=locked \
    apt-get update && apt-get install -y --no-install-recommends ca-certificates

RUN --mount=type=cache,target=/root/.npm npm ci --prefer-offline
RUN --mount=type=cache,target=/root/.cache/pip pip install -r requirements.txt   # do not use --no-cache-dir with a cache mount
RUN --mount=type=cache,target=/root/.cache/uv,uid=10001 uv sync --frozen         # uid for a non-root build user

RUN --mount=type=bind,source=go.sum,target=go.sum,readonly go mod verify   # file from the context without a COPY layer
RUN --mount=type=ssh git clone git@git.example.com:org/private.git          # with: docker build --ssh default .
RUN --mount=type=secret,id=netrc,target=/root/.netrc curl -fsS https://internal.example.com/artifact
```

`sharing=shared` (default) lets parallel builds write to the same cache concurrently, `locked` serialises them (required for `apt` and `dpkg`), `private` gives each build its own copy. `id=` names a cache independently of `target` when two paths should share one. Cache mounts live in the builder's storage, so they are lost by `docker builder prune` and do not exist on a fresh CI runner; export the layer cache for that.

```sh
docker buildx build --cache-to type=registry,ref=registry.example.com/my-app:buildcache,mode=max \
                    --cache-from type=registry,ref=registry.example.com/my-app:buildcache -t my-app:1.0 --push .
docker buildx build --cache-to type=gha,mode=max --cache-from type=gha .    # GitHub Actions cache backend
docker buildx build --cache-to type=local,dest=/var/cache/buildkit --cache-from type=local,src=/var/cache/buildkit .
docker buildx build --output type=local,dest=./out --target build .          # export files instead of an image
docker buildx du && docker buildx prune --keep-storage 10GB                  # what the builder holds
```

`mode=max` exports the cache for every stage, not only the layers in the final image; without it intermediate stages are rebuilt on the next runner. Layer cache from a registry does not restore cache mounts; a `--mount=type=cache` in CI helps only when the runner's builder persists between jobs.

## Healthchecks

A healthcheck runs a command inside the container on an interval and sets `.State.Health.Status` to `starting`, `healthy` or `unhealthy`. Docker itself only records the status; Compose (`depends_on: condition: service_healthy`, `up --wait`), Swarm and external tools act on it. A container that is `unhealthy` is not restarted by the daemon.

```dockerfile
HEALTHCHECK --interval=30s --timeout=5s --start-period=40s --start-interval=5s --retries=3 \
  CMD ["/server", "healthcheck"]                 # exec form: no shell needed in the image; exit 0 healthy, 1 unhealthy
HEALTHCHECK NONE                                  # disable one inherited from the base image
```

```sh
docker run -d --health-cmd 'curl -fsS http://127.0.0.1:8080/health || exit 1' --health-interval 10s --health-retries 3 --health-start-period 30s my-app
docker run -d --no-healthcheck my-app             # ignore the image's HEALTHCHECK
docker inspect -f '{{.State.Health.Status}}' my-app
docker inspect -f '{{range .State.Health.Log}}{{.End}} {{.ExitCode}} {{.Output}}{{end}}' my-app   # last five results with output
docker ps --filter health=unhealthy
docker events --filter event=health_status        # health_status: healthy / unhealthy transitions
```

`start-period` suppresses failures while the application boots but a success inside it flips the status to `healthy` immediately; `start-interval` (Engine 25+) probes faster during that window. Only exit codes 0 and 1 are meaningful; `2` is reserved. Distroless images have no `curl` or `wget`, so give the binary a health subcommand or copy a static probe such as `grpc-health-probe`. Keep the check cheap: it runs forever, and a check that hits the database turns a database blip into a fleet of unhealthy frontends.

## Logging drivers

The daemon captures PID 1's stdout and stderr through a logging driver chosen per container. `json-file` is the default and the only one, with `local`, that `docker logs` can read without extra configuration; the others ship logs away and `docker logs` fails unless dual logging (Engine 20.10+) caches a copy.

| Driver | Where the logs go | Notes |
| --- | --- | --- |
| `json-file` | `/var/lib/docker/containers/<id>/<id>-json.log` | No rotation unless `max-size` is set |
| `local` | Same directory, compressed, rotated by default (100 MB total) | Preferred default for new hosts |
| `journald` | The host journal, with `CONTAINER_NAME`, `CONTAINER_ID` fields | `journalctl CONTAINER_NAME=my-app -f`; rotation handled by journald |
| `syslog`, `fluentd`, `gelf`, `awslogs`, `splunk` | A remote collector | `docker logs` works only with dual logging |
| `none` | Discarded | For chatty containers whose logs are collected another way |

```sh
docker run -d --log-driver local --log-opt max-size=20m --log-opt max-file=5 my-app
docker run -d --log-driver journald --log-opt tag='{{.Name}}' my-app
docker run -d --log-opt mode=non-blocking --log-opt max-buffer-size=4m my-app   # drop logs rather than block the app if the driver stalls
docker inspect -f '{{.HostConfig.LogConfig.Type}} {{.HostConfig.LogConfig.Config}}' my-app
docker inspect -f '{{.LogPath}}' my-app                                     # the file json-file or local writes
sudo du -sh /var/lib/docker/containers/*/*-json.log | sort -h | tail        # who is filling the disk
```

Set the default for every new container in `/etc/docker/daemon.json` and restart the daemon; existing containers keep the driver they were created with until recreated.

```json
{
  "log-driver": "local",
  "log-opts": { "max-size": "20m", "max-file": "5" }
}
```

Applications that write to files inside the container are invisible to all of this. Point them at `/dev/stdout` and `/dev/stderr` (the nginx image symlinks its log files there) or a shared volume with a sidecar tailer.

## Resource limits

Limits are cgroup settings applied to the container's cgroup, visible under `/sys/fs/cgroup/system.slice/docker-<id>.scope/` on a systemd host. Without them a container can consume the whole machine; with a memory limit and no swap accounting, exceeding it is an OOM kill (`OOMKilled: true`, exit 137).

```sh
docker run -d --memory 512m --memory-swap 512m my-app     # hard limit, no swap (swap = memory-swap - memory)
docker run -d --memory 512m --memory-reservation 256m my-app   # soft target used under host pressure
docker run -d --cpus 1.5 my-app                           # CFS quota: 1.5 CPUs per period; throttled above that
docker run -d --cpu-shares 512 my-app                     # relative weight under contention only (default 1024)
docker run -d --cpuset-cpus 2-3 my-app                    # pin to CPUs 2 and 3
docker run -d --pids-limit 256 my-app                     # fork-bomb protection
docker run -d --ulimit nofile=65536:65536 --ulimit nproc=4096 my-app
docker run -d --shm-size 1g my-app                        # /dev/shm, 64 MB by default; browsers and PostgreSQL need more
docker run -d --device-read-bps /dev/sda:50mb --device-write-iops /dev/sda:1000 my-app
docker update --memory 1g --memory-swap 1g --cpus 2 my-app   # live change; memory-swap must be updated with memory
docker stats --no-stream --format 'table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}\t{{.MemPerc}}\t{{.PIDs}}'
docker exec my-app cat /sys/fs/cgroup/cpu.stat /sys/fs/cgroup/memory.max /sys/fs/cgroup/memory.events   # throttling and OOM counters from inside
```

`--cpus` throttles in 100 ms periods, so a container averaging 50% of its quota can still stall for tens of milliseconds in bursts; `nr_throttled` in `cpu.stat` shows it. JVMs and Go runtimes read the cgroup limit and size their heaps and `GOMAXPROCS` accordingly (Go 1.25+ for `GOMAXPROCS`); older runtimes see the host's CPUs and memory and need `-XX:MaxRAMPercentage` or `GOMAXPROCS` set explicitly. `--oom-kill-disable` exists and should not be used: the container then hangs instead of dying. Interpretation of the cgroup counters is covered in [Linux performance](https://www.wiki.jodisand.me/linux-performance/#cgroup-accounting).

## Security options

The default container has a root user, a default seccomp profile, a bounded capability set and a private network namespace. Each of the following narrows it further; apply them in the image where possible and in the run configuration otherwise.

```sh
docker run -d --user 10001:10001 my-app                    # never rely on the image's USER when you did not build it
docker run -d --read-only --tmpfs /tmp:rw,noexec,nosuid,size=64m -v my-data:/app/data my-app
docker run -d --cap-drop ALL --cap-add NET_BIND_SERVICE my-app
docker run -d --security-opt no-new-privileges:true my-app  # setuid binaries and file capabilities cannot raise privileges
docker run -d --security-opt seccomp=/etc/docker/seccomp/my-app.json my-app   # custom syscall allowlist; default profile blocks ~44 syscalls
docker run -d --security-opt apparmor=docker-default my-app # Debian/Ubuntu; SELinux hosts label containers as container_t automatically
docker run -d --security-opt label=type:my_container_t my-app   # custom SELinux type (Fedora/RHEL, needs a policy module)
docker run -d --security-opt label=disable -v /var/log:/host/log:ro my-app   # skip SELinux labelling for one container; prefer :z or :Z
docker run -d --userns=host my-app                          # opt out of daemon-level user namespace remapping for one container
docker run -d --pull always registry.example.com/my-app@sha256:...   # pin the digest and refuse a stale local copy
docker inspect -f '{{.HostConfig.Privileged}} {{.HostConfig.CapAdd}} {{.HostConfig.SecurityOpt}} {{.Config.User}}' my-app
docker container ls -q | xargs docker inspect -f '{{.Name}} privileged={{.HostConfig.Privileged}} user={{.Config.User}}'
```

Daemon-wide hardening in `/etc/docker/daemon.json`: `"userns-remap": "default"` maps container root to an unprivileged host UID range (breaks bind mounts owned by host users and needs a one-off migration of image storage), `"no-new-privileges": true` applies that option to every container, `"icc": false` stops containers on the default bridge talking to each other, and `"live-restore": true` keeps containers running across a daemon restart. Rootless Docker (`dockerd-rootless-setuptool.sh install`) runs the daemon itself without root and is the strongest option where it fits.

> [!WARNING] `--privileged` and the socket
> `--privileged` disables seccomp, AppArmor and SELinux confinement, grants every capability and exposes every host device; it is host root with extra steps. Mounting `/var/run/docker.sock` into a container has the same effect, since anything that can talk to the daemon can start a privileged container. Membership of the `docker` group is equivalent to root for the same reason.

## Volumes and bind mounts

Writes inside a container land in its writable layer and are removed with the container. A named volume is a directory the daemon manages under `/var/lib/docker/volumes`. A bind mount grafts a host path into the container with host ownership and permissions intact.

```sh
docker volume create my-data
docker run -d -v my-data:/var/lib/postgresql/data postgres:17   # named volume
docker run -v "$PWD/src:/app/src:ro,Z" my-app                   # bind mount, read-only, SELinux relabel
docker run --mount type=tmpfs,destination=/tmp my-app           # memory-backed scratch
docker volume inspect my-data                                   # Mountpoint on the host
docker volume ls -f dangling=true                               # not attached to any container
```

A new empty named volume is populated from the image's content at the mount path the first time it is mounted. A bind mount is never populated: it hides whatever the image had at that path.

Back up a named volume to a tarball in the current directory:

```sh
docker run --rm -v my-data:/data:ro -v "$PWD:/backup" alpine:3.24 tar czf /backup/my-data.tgz -C /data .
```

Stop the writing container first for databases, or use the database's own dump tool (see [PostgreSQL](https://www.wiki.jodisand.me/postgresql/)); a file copy of a live data directory can be inconsistent.

## Networks and published ports

On a user-defined network the daemon runs an embedded DNS resolver at `127.0.0.11`, so containers reach each other by container name or `--network-alias`. The default `bridge` network has no name resolution, which explains most "works in Compose, fails with `docker run`" reports.

```sh
docker network create my-net
docker run -d --network my-net --name db postgres:17
docker run --rm --network my-net alpine:3.24 getent hosts db   # prints the container IP
docker network connect my-net my-app                            # attach a running container
docker network inspect my-net -f '{{json .Containers}}'
docker port my-app                                              # published port mappings
```

| Mode | Behaviour |
| --- | --- |
| `bridge` (default) | Private network behind NAT; reachable from outside only through published ports; no DNS |
| User-defined bridge | Same, plus DNS by name and isolation from other networks |
| `host` | Shares the host network namespace; `-p` is ignored; the process binds host ports directly |
| `none` | Loopback only |
| `container:<name>` | Shares another container's network namespace (same IP and ports) |

Docker programs NAT and filter rules itself (iptables by default; an nftables backend exists and is selected with the daemon's `firewall-backend` option). Published ports are translated in the `nat` table before host firewall rules such as ufw see the packet, so a port published on `0.0.0.0` can be reachable even when the host firewall appears to block it. With firewalld, Docker adds a `docker` zone. Publish on `127.0.0.1` anything that should not leave the host. See [packet filtering and firewalls](https://docs.docker.com/engine/network/packet-filtering-firewalls/). For inspecting the resulting interfaces and routes, see [iproute2](https://www.wiki.jodisand.me/iproute2/).

## Logs and debugging

`docker logs` reads what the logging driver captured from PID 1's stdout and stderr. The default `json-file` driver does not rotate unless `max-size` is set in `/etc/docker/daemon.json`, so long-running chatty containers can fill `/var/lib/docker`.

```sh
docker logs -f --since 15m --timestamps my-app
docker exec -it my-app sh                           # bash on Debian/Ubuntu-based images
docker exec my-app ps -eo pid,comm,rss              # what is actually running
docker stats --no-stream                            # CPU, memory, network and block I/O per container
docker top my-app                                   # host-side view of the container's processes
docker diff my-app                                  # files added (A), changed (C) or deleted (D) since start
docker events --since 10m --filter container=my-app # die, oom, kill and health_status events
```

Debug a minimal image with no shell by joining its PID and network namespaces from a tools image:

```sh
docker run --rm -it --pid container:my-app --network container:my-app --cap-add SYS_PTRACE nicolaka/netshoot
```

Inside, `ps`, `ss -tlnp`, `curl` and `tcpdump` see the target container's processes and sockets.

## Registries and multi-platform builds

```sh
printf '%s' "$REGISTRY_TOKEN" | docker login registry.example.com -u alice --password-stdin
docker buildx build --platform linux/amd64,linux/arm64 -t registry.example.com/my-app:1.0 --push .
```

Without a credential helper (`credsStore` in `~/.docker/config.json`), credentials are stored in that file base64-encoded, not encrypted. `--password-stdin` keeps the token out of shell history and the process list.

A multi-platform build produces one manifest list with an image per platform. Building a foreign platform uses QEMU emulation unless the builder has native nodes, and is much slower.

## Cleaning up disk space

```sh
docker system df -v                                 # what is using the space, per object
docker container prune                              # stopped containers
docker image prune                                  # dangling (untagged) images
docker image prune -a --filter "until=168h"         # every image not used by a container, older than 7 days
docker builder prune --keep-storage 10GB            # trim BuildKit cache, often the largest item
docker system prune                                 # stopped containers, unused networks, dangling images, build cache
```

> [!WARNING] Volume prune deletes data
> `docker volume prune` and `docker system prune --volumes` remove unused anonymous volumes. `docker volume prune -a` (API 1.42+, Engine 23+) also removes unused **named** volumes, which is where databases usually live. "Unused" means no container, running or stopped, references it; a stack that is down with `docker compose down` has no containers. Run `docker system df -v` first.

## Troubleshooting

| Symptom | Cause | Check |
| --- | --- | --- |
| `Cannot connect to the Docker daemon` | Daemon not running, wrong context, or user not in the `docker` group | `systemctl status docker`, `docker context ls`, `id` |
| `permission denied ... docker.sock` | User lacks access to the socket | Group membership needs a new login; the `docker` group is root-equivalent |
| `port is already allocated` | Another container or host process holds the port | `docker ps --format '{{.Names}}\t{{.Ports}}'`, `ss -tlnp` |
| `no space left on device` during build | Build cache or images filled `/var/lib/docker` | `docker system df`, then prune the cache |
| `exec format error` | Image built for another architecture | `docker image inspect -f '{{.Architecture}}' my-app:1.0` |
| `manifest unknown` or `not found` on pull | Tag does not exist for this platform, or wrong registry path | `docker buildx imagetools inspect <image>` |
| Container cannot reach the internet | DNS from the host's resolver, or forwarding disabled | `docker run --rm alpine:3.24 nslookup example.com`, `sysctl net.ipv4.ip_forward` |
| `docker stop` always takes 10 s | PID 1 ignores SIGTERM (shell form, or no signal handler) | Use exec form or `--init` |
| Health status stuck at `starting` | Health command fails or the tool it calls is missing from the image | `docker inspect -f '{{json .State.Health}}' my-app` |
| `failed to solve: ... --mount option requires BuildKit` | Legacy builder in use | `DOCKER_BUILDKIT=1`, or upgrade; BuildKit is the default since Engine 23 |
| Cache mount empty on every CI run | Builder storage is not persisted between jobs | Use `--cache-to/--cache-from` for layers; accept that cache mounts are per builder |
| `COPY failed: file not found in build context` | File excluded by `.dockerignore`, or path outside the context | `docker build --progress=plain` shows the transferred context; check the ignore file |
| Build slow, "transferring context" takes minutes | `.git`, `node_modules` or data directories in the context | Add them to `.dockerignore`; check with `du -sh` in the context directory |
| Multi-platform build fails with `exec format error` inside `RUN` | Emulation not installed on the builder | `docker run --privileged --rm tonistiigi/binfmt --install all`, or cross-compile with `$BUILDPLATFORM` |
| `docker logs` prints nothing for a driver such as `fluentd` | Driver does not support reading back; dual logging off | Read at the collector, or enable `cache-disabled: false` in the driver options |
| `/var/lib/docker/containers` growing | `json-file` without rotation | Set `log-opts` `max-size` in `daemon.json`; recreate containers |
| Container killed at exactly the limit but `free` inside shows plenty | `free` reports host memory; the cgroup limit is what applies | `cat /sys/fs/cgroup/memory.max`, `memory.events` inside the container |
| Application sees all host CPUs and over-parallelises | Runtime not cgroup-aware (Java < 10, Go < 1.25 for `GOMAXPROCS`) | Set `GOMAXPROCS`, `-XX:ActiveProcessorCount` or the equivalent to match `--cpus` |
| `operation not permitted` after `--cap-drop ALL` | Process needs a capability (bind < 1024, chown, raw sockets) | Add capabilities back one at a time; `NET_BIND_SERVICE`, `CHOWN`, `SETUID`/`SETGID` are the usual ones; `journalctl -k -g audit` shows denials on SELinux hosts |
| `Read-only file system` with `--read-only` | Application writes to a path not covered by a volume or tmpfs | `docker diff` on a writable run to find the paths, then add `--tmpfs` or a volume |

For host-level CPU, memory and I/O pressure, see [Linux performance](https://www.wiki.jodisand.me/linux-performance/#the-first-minute). For the daemon itself, `journalctl -u docker` (see [systemd](https://www.wiki.jodisand.me/systemd/#a-failing-service)).

## Oneliners

```sh
# Stop every running container
docker stop $(docker ps -q)

# Remove every exited container
docker rm $(docker ps -aq -f status=exited)

# Every container's IP on every network
docker inspect -f '{{.Name}} {{range .NetworkSettings.Networks}}{{.IPAddress}} {{end}}' $(docker ps -q)

# Images sorted by size
docker image ls --format '{{.Size}}\t{{.Repository}}:{{.Tag}}' | sort -h -r | head

# Environment of a running container (may print secrets)
docker inspect -f '{{range .Config.Env}}{{println .}}{{end}}' my-app

# Which container publishes port 8080
docker ps --format '{{.Names}}\t{{.Ports}}' | grep ':8080->'

# Wait until a health check passes
until [ "$(docker inspect -f '{{.State.Health.Status}}' my-app)" = healthy ]; do sleep 1; done

# Containers with their restart policy and restart count (spot crash loops)
docker ps -a --format '{{.Names}}' | xargs docker inspect -f '{{.Name}} {{.HostConfig.RestartPolicy.Name}} restarts={{.RestartCount}} {{.State.Status}}'

# Containers running as root
docker ps -q | xargs docker inspect -f '{{.Name}} user={{.Config.User}}' | awk '$2 == "user=" || $2 == "user=root" || $2 == "user=0"'

# Containers with no memory limit
docker ps -q | xargs docker inspect -f '{{.Name}} {{.HostConfig.Memory}}' | awk '$2 == 0'

# Memory usage against limit, sorted
docker stats --no-stream --format '{{.MemPerc}}\t{{.MemUsage}}\t{{.Name}}' | sort -rn | head

# Containers with a mounted Docker socket (root-equivalent)
docker ps -q | xargs docker inspect -f '{{.Name}} {{range .Mounts}}{{.Source}} {{end}}' | grep docker.sock

# Digest a container is actually running, for comparison with the registry
docker inspect -f '{{.Image}}' my-app | xargs docker image inspect -f '{{index .RepoDigests 0}}'

# Log size per container
sudo sh -c 'for f in /var/lib/docker/containers/*/*-json.log; do printf "%s %s\n" "$(du -m "$f" | cut -f1)" "$(basename "$(dirname "$f")" | cut -c1-12)"; done' | sort -rn | head

# Resolve a short container ID to a name
docker ps -a --no-trunc --format '{{.ID}} {{.Names}}' | grep '^abc123'

# Export an image as a tarball and import it on an offline host
docker save my-app:1.0 | zstd -T0 > my-app-1.0.tar.zst;  zstd -dc my-app-1.0.tar.zst | docker load

# Flatten the filesystem of a container into a plain tar (no history, no metadata)
docker export my-app | tar -tvf - | head

# Build with full logs and keep the failing layer's state for inspection
docker build --progress=plain --target build . 2>&1 | tee build.log

# List the files a Dockerfile stage would produce, without creating an image
docker buildx build --target build --output type=tar,dest=- . | tar -tvf - | head -40

# Layers of an image with the command that created each, largest first
docker history --no-trunc --format '{{.Size}}\t{{.CreatedBy}}' my-app:1.0 | sort -h -r | head

# Compare the packages installed in two image versions (Debian-based)
diff <(docker run --rm my-app:1.0 dpkg-query -W) <(docker run --rm my-app:1.1 dpkg-query -W)

# Run a one-off command in the same network as a container, with a tools image
docker run --rm --network container:my-app nicolaka/netshoot curl -sS http://127.0.0.1:8080/health

# Copy a file into a running container without docker cp (works with read-only root if the target is a volume)
tar -cf - config.yaml | docker exec -i my-app tar -xf - -C /app/data

# Follow logs from several containers with names prefixed
docker compose logs -f 2>/dev/null || for c in web worker; do docker logs -f --tail 20 "$c" 2>&1 | sed "s/^/[$c] /" & done; wait

# Events for the last hour as JSON lines: OOMs, restarts, health flips
docker events --since 1h --until "$(date +%s)" --format '{{json .}}' | jq -r 'select(.status | IN("oom", "die", "health_status: unhealthy")) | [.time, .status, .Actor.Attributes.name] | @tsv'

# Dangling volumes with their size
docker system df -v | awk '/^VOLUME NAME/{f=1; next} f && $2 == 0 {print $1, $3}'

# Remove images older than 30 days that no container uses
docker image prune -a --filter 'until=720h' --force

# Check whether the daemon is on cgroup v2 and which storage driver it uses
docker info --format '{{.CgroupVersion}} {{.Driver}} {{.LoggingDriver}}'
```

## Scripts

Audit every running container for the risky settings that appear in incident reviews: privileged mode, host networking, mounted Docker socket, root user, no memory limit, missing healthcheck, and `json-file` logging without rotation. Read-only; one line per finding.

```sh
#!/usr/bin/env bash
# usage: docker-audit.sh [container...]   default: all running containers
set -euo pipefail

mapfile -t ids < <(if (( $# )); then printf '%s\n' "$@"; else docker ps -q; fi)
(( ${#ids[@]} )) || { echo 'no containers'; exit 0; }
findings=0
note() { printf '%-28s %s\n' "$1" "$2"; (( findings++ )) || true; }

for id in "${ids[@]}"; do
  j=$(docker inspect "$id")
  name=$(jq -r '.[0].Name | ltrimstr("/")' <<< "$j")
  jq -e '.[0].HostConfig.Privileged' <<< "$j" >/dev/null && note "$name" 'privileged'
  [[ $(jq -r '.[0].HostConfig.NetworkMode' <<< "$j") == host ]] && note "$name" 'host network'
  [[ $(jq -r '.[0].HostConfig.PidMode' <<< "$j") == host ]] && note "$name" 'host PID namespace'
  jq -e '.[0].Mounts[] | select(.Source | test("docker.sock$"))' <<< "$j" >/dev/null && note "$name" 'docker socket mounted'
  [[ $(jq -r '.[0].Config.User' <<< "$j") =~ ^(|root|0|0:0)$ ]] && note "$name" 'runs as root'
  [[ $(jq -r '.[0].HostConfig.Memory' <<< "$j") == 0 ]] && note "$name" 'no memory limit'
  [[ $(jq -r '.[0].HostConfig.PidsLimit // 0' <<< "$j") == 0 ]] && note "$name" 'no pids limit'
  jq -e '.[0].HostConfig.CapAdd // [] | index("SYS_ADMIN")' <<< "$j" >/dev/null && note "$name" 'CAP_SYS_ADMIN added'
  jq -e '.[0].State.Health == null' <<< "$j" >/dev/null && note "$name" 'no healthcheck'
  if [[ $(jq -r '.[0].HostConfig.LogConfig.Type' <<< "$j") == json-file ]] && ! jq -e '.[0].HostConfig.LogConfig.Config["max-size"]' <<< "$j" >/dev/null; then
    note "$name" 'json-file logging without max-size'
  fi
done
printf '%d findings across %d containers\n' "$findings" "${#ids[@]}"
(( findings == 0 ))
```

Back up every named volume on the host to compressed tarballs, skipping volumes attached to a running container unless `--stop` is given, in which case the containers are stopped for the copy and started again. Stops services when run with `--stop`.

```sh
#!/usr/bin/env bash
# usage: volume-backup.sh [--stop] DEST_DIR
set -euo pipefail

stop=0
[[ ${1:-} == --stop ]] && { stop=1; shift; }
dest=${1:?destination directory}
mkdir -p "$dest"
stamp=$(date +%Y%m%dT%H%M%S)

while IFS= read -r vol; do
  mapfile -t running < <(docker ps -q --filter "volume=$vol")
  if (( ${#running[@]} )); then
    if (( stop )); then docker stop "${running[@]}" >/dev/null; else printf 'skip %s: in use by running container(s)\n' "$vol" >&2; continue; fi
  fi
  out=$dest/$vol-$stamp.tar.zst
  docker run --rm -v "$vol:/data:ro" -v "$dest:/backup" alpine:3.24 \
    sh -c 'apk add --no-cache zstd >/dev/null && tar -C /data -cf - . | zstd -T0 -q > "/backup/$0"' "$(basename "$out")"
  printf '%s %s\n' "$(du -h "$out" | cut -f1)" "$out"
  if (( ${#running[@]} && stop )); then docker start "${running[@]}" >/dev/null; fi
done < <(docker volume ls -q --filter dangling=false; docker volume ls -q --filter dangling=true)
```

Report the image each running container was started from, whether the tag now points at a different digest in the registry, and the age of the running image. Read-only, needs network access to the registries.

```sh
#!/usr/bin/env bash
# usage: image-drift.sh
set -euo pipefail

printf '%-30s %-45s %-8s %s\n' CONTAINER IMAGE AGE STATUS
for id in $(docker ps -q); do
  name=$(docker inspect -f '{{.Name}}' "$id"); name=${name#/}
  ref=$(docker inspect -f '{{.Config.Image}}' "$id")
  local_digest=$(docker inspect -f '{{.Image}}' "$id" | xargs docker image inspect -f '{{if .RepoDigests}}{{index .RepoDigests 0}}{{end}}' | sed 's/.*@//')
  created=$(docker inspect -f '{{.Image}}' "$id" | xargs docker image inspect -f '{{.Created}}')
  age=$(( ($(date +%s) - $(date -d "$created" +%s)) / 86400 ))d
  if [[ $ref == *@sha256:* ]]; then status=pinned
  elif remote=$(timeout 20 docker buildx imagetools inspect "$ref" --format '{{json .Manifest.Digest}}' 2>/dev/null | tr -d '"'); then
    [[ $remote == "$local_digest" ]] && status=current || status="STALE (registry $remote)"
  else status='registry unreachable'; fi
  printf '%-30s %-45s %-8s %s\n' "$name" "$ref" "$age" "$status"
done
```

## Further reading

- [Dockerfile reference](https://docs.docker.com/reference/dockerfile/): every instruction, `--mount` types and the `# syntax` directive.
- [docker run reference](https://docs.docker.com/reference/cli/docker/container/run/): the complete flag list for resources, security and logging.
- [Build cache](https://docs.docker.com/build/cache/) and [cache backends](https://docs.docker.com/build/cache/backends/): invalidation rules and `--cache-to` types.
- [Logging drivers](https://docs.docker.com/engine/logging/configure/): options per driver, dual logging and `daemon.json` defaults.
- [Runtime metrics and resource constraints](https://docs.docker.com/engine/containers/resource_constraints/): how `--memory`, `--cpus` and the rest map to cgroup controls.
- [Docker security](https://docs.docker.com/engine/security/): capabilities, seccomp, AppArmor, user namespaces and rootless mode.


