Software Engineering WikiSE Wiki

Python

Write operational Python scripts that fail safely: subprocess, paths, logging, HTTP, configuration, concurrency, uv packaging and troubleshooting.

Reviewed MarkdownEdit

On this page

Cheatsheet#

TaskSnippet or command
Run a command, raise on failuresubprocess.run(cmd, check=True, capture_output=True, text=True, timeout=60)
Build a pathPath("/etc") / "app.conf"
Atomic file writewrite a temp file in the same directory, then os.replace(tmp, target)
Log with the tracebacklog.exception("sync failed") inside except
HTTP with a timeouthttpx.get(url, timeout=10)
Retry with backofftenacity.retry(wait=wait_exponential(), stop=stop_after_attempt(5))
Typed config from environmentpydantic_settings.BaseSettings
Temporary directorywith tempfile.TemporaryDirectory() as d:
Parse CLI argumentsargparse.ArgumentParser
Time a blockt = time.perf_counter(); ...; time.perf_counter() - t
Create a projectuv init --package my-app
Add a dependencyuv add httpx / uv add --dev pytest
Run in the project environmentuv run pytest
CI install, fail if lock is staleuv sync --locked
Upgrade one locked packageuv lock --upgrade-package httpx
Run a tool without installing ituvx ruff check .
Install a CLI tool globallyuv tool install ruff
Install a Python versionuv python install 3.14
Format and lintruff format . && ruff check --fix .
Type-checkmypy --strict src/
Run testspytest -x -q (see Testing)

A script worth keeping#

#!/usr/bin/env python3
"""Reconcile inventory against the API."""
from __future__ import annotations

import argparse
import logging
import sys
from pathlib import Path

log = logging.getLogger("reconcile")


def main() -> int:
    ap = argparse.ArgumentParser(description=__doc__)
    ap.add_argument("inventory", type=Path)
    ap.add_argument("--dry-run", action="store_true")
    ap.add_argument("-v", "--verbose", action="count", default=0)
    args = ap.parse_args()

    logging.basicConfig(
        level=logging.DEBUG if args.verbose else logging.INFO,
        format="%(asctime)s %(levelname)s %(name)s %(message)s",
        stream=sys.stderr,
    )

    if not args.inventory.is_file():
        log.error("inventory not found: %s", args.inventory)
        return 2

    log.info("reconciling %s dry_run=%s", args.inventory, args.dry_run)
    return 0


if __name__ == "__main__":
    raise SystemExit(main())

main returns an exit code, logs go to stderr, and stdout is reserved for data. That lets the script sit in a pipeline and makes failures visible in CI. Pass arguments to the logger (log.info("x=%s", x)) instead of an f-string so formatting only happens when the level is enabled.

A single-file script can declare its own dependencies (PEP 723). uv run script.py reads the block and runs it in a cached environment:

# /// script
# requires-python = ">=3.12"
# dependencies = ["httpx"]
# ///
uv add --script reconcile.py httpx     # writes the block for you
uv run reconcile.py inventory.csv

Running commands with subprocess#

import json, subprocess

res = subprocess.run(
    ["kubectl", "get", "pods", "-o", "json"],
    check=True, capture_output=True, text=True, timeout=30,
)
pods = json.loads(res.stdout)
ArgumentWhy
List, not a stringNo shell, so no quoting problems and no injection
check=TrueRaises CalledProcessError on a non-zero exit instead of continuing
capture_output=True, text=Truestdout and stderr as str
timeout=Kills the child and raises TimeoutExpired instead of hanging the job
cwd=, env=Explicit context. env= replaces the whole environment, so start from {**os.environ, ...}

shell=True is only safe with a literal string you wrote. With any interpolated value it is a command injection bug. If a shell is truly needed, quote each value with shlex.quote.

try:
    subprocess.run(cmd, check=True, capture_output=True, text=True, timeout=60)
except subprocess.CalledProcessError as e:
    log.error("command failed rc=%s stderr=%s", e.returncode, e.stderr.strip())
    raise
except subprocess.TimeoutExpired:
    log.error("command timed out: %s", shlex.join(cmd))
    raise

For long-running output, stream it instead of buffering everything in memory:

with subprocess.Popen(cmd, stdout=subprocess.PIPE, text=True) as p:
    for line in p.stdout:
        handle(line)
if p.returncode != 0:
    raise subprocess.CalledProcessError(p.returncode, cmd)

Paths and files#

from pathlib import Path

base = Path("/srv/app")
cfg = base / "conf" / "app.yaml"
cfg.exists(); cfg.stat().st_size; cfg.read_text(encoding="utf-8")
list(base.rglob("*.log"))
base.mkdir(parents=True, exist_ok=True)

Always pass encoding="utf-8" when reading or writing text. The default comes from the locale, so the same script can behave differently on another host (Python 3.15 changes the default to UTF-8 under PEP 686).

Write atomically so a crash cannot leave a half-written file where a complete one is expected:

import os, tempfile

def write_atomic(path: Path, data: str) -> None:
    fd, tmp = tempfile.mkstemp(dir=path.parent)     # same filesystem as the target
    try:
        with os.fdopen(fd, "w", encoding="utf-8") as f:
            f.write(data)
            f.flush()
            os.fsync(f.fileno())                     # data on disk before the rename
        os.replace(tmp, path)                        # atomic rename on POSIX
    except BaseException:
        os.unlink(tmp)
        raise

os.replace is only atomic within one filesystem, which is why the temp file is created in the target’s directory. mkstemp creates the file with mode 0600; chmod it if other users need to read it.

Logging#

import json, logging

class JsonFormatter(logging.Formatter):
    def format(self, record: logging.LogRecord) -> str:
        payload = {
            "ts": self.formatTime(record, "%Y-%m-%dT%H:%M:%S%z"),
            "level": record.levelname,
            "logger": record.name,
            "msg": record.getMessage(),
            **getattr(record, "fields", {}),
        }
        if record.exc_info:
            payload["exc"] = self.formatException(record.exc_info)
        return json.dumps(payload)

handler = logging.StreamHandler()
handler.setFormatter(JsonFormatter())
logging.basicConfig(level=logging.INFO, handlers=[handler])

log.info("deployed", extra={"fields": {"service": "api", "version": "1.4.2"}})

extra sets attributes on the LogRecord, so a key that collides with a built-in attribute (msg, args, name) raises KeyError. Nesting under one key avoids that. Call log.exception() inside an except block to record the traceback. Do not log secrets, tokens or full request bodies.

HTTP clients, timeouts and retries#

import httpx

with httpx.Client(timeout=10.0, headers={"user-agent": "reconcile/1.0"}) as client:
    r = client.get("https://api.example.com/items", params={"limit": 100})
    r.raise_for_status()
    items = r.json()

httpx applies a 5 second timeout by default. requests has no default timeout and can wait forever on a stalled connection, so always pass timeout=. Reuse one Client for many requests to keep connection pooling.

from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception

def transient(e: BaseException) -> bool:
    if isinstance(e, httpx.TransportError):          # timeouts, connection resets
        return True
    return isinstance(e, httpx.HTTPStatusError) and e.response.status_code in {429, 502, 503, 504}

@retry(
    stop=stop_after_attempt(5),
    wait=wait_exponential(multiplier=0.5, max=10),
    retry=retry_if_exception(transient),
    reraise=True,
)
def fetch(client: httpx.Client, url: str) -> dict:
    r = client.get(url)
    r.raise_for_status()
    return r.json()

Retry only transient failures. Retrying every HTTPStatusError also retries 400 and 404, which never succeed. Retry idempotent requests only: a retried POST can create two records. See HTTP for status code meanings.

Configuration and secrets#

from pydantic import SecretStr
from pydantic_settings import BaseSettings, SettingsConfigDict

class Settings(BaseSettings):
    model_config = SettingsConfigDict(env_prefix="APP_", env_file=".env")

    api_url: str
    api_token: SecretStr                          # printed as '**********'
    timeout: float = 10.0
    dry_run: bool = False

settings = Settings()                             # raises ValidationError if APP_API_URL is missing
token = settings.api_token.get_secret_value()     # explicit access to the real value

Validate configuration once at start-up and exit on failure. A missing variable found three hours into a batch job costs more than a crash on line one. Keep .env out of version control. See Vault for fetching secrets at runtime.

Data handling#

from dataclasses import dataclass

@dataclass(frozen=True, slots=True)
class Host:
    name: str
    ip: str
    zone: str
    tags: tuple[str, ...] = ()

from collections import Counter, defaultdict

counts = Counter(h.tags[0] for h in hosts if h.tags)
by_zone: dict[str, list[Host]] = defaultdict(list)
for h in hosts:
    by_zone[h.zone].append(h)

import csv
with open("hosts.csv", newline="", encoding="utf-8") as f:
    rows = list(csv.DictReader(f))

frozen=True makes instances hashable and blocks accidental mutation. slots=True removes the per-instance __dict__, which saves memory once there are hundreds of thousands of objects. Use a tuple, not a list, for fields of a frozen dataclass, or hashing fails.

Generators keep memory flat for large inputs:

def read_events(path: Path):
    with path.open(encoding="utf-8") as f:
        for line in f:                 # one line at a time, not the whole file
            yield json.loads(line)

Concurrency: threads, asyncio or processes#

WorkloadTool
Network I/O, tens of callsThreadPoolExecutor
Network I/O, thousands of callsasyncio with httpx.AsyncClient
CPU-bound pure PythonProcessPoolExecutor
Running other programsThreads; the GIL is released while waiting on the child
from concurrent.futures import ThreadPoolExecutor, as_completed

with ThreadPoolExecutor(max_workers=8) as pool:
    futures = {pool.submit(check_host, h): h for h in hosts}
    for fut in as_completed(futures):
        host = futures[fut]
        try:
            result = fut.result()
        except Exception:
            log.exception("check failed host=%s", host.name)
import asyncio, httpx

async def fetch_all(urls: list[str]) -> list[dict]:
    limit = asyncio.Semaphore(20)                  # at most 20 requests in flight
    async with httpx.AsyncClient(timeout=10) as client:
        async def one(u: str) -> dict:
            async with limit:
                r = await client.get(u)
                r.raise_for_status()
                return r.json()
        async with asyncio.TaskGroup() as tg:      # 3.11+: cancels the rest if one fails
            tasks = [tg.create_task(one(u)) for u in urls]
    return [t.result() for t in tasks]

results = asyncio.run(fetch_all(urls))

TaskGroup cancels sibling tasks when one raises and reports failures as an ExceptionGroup (catch with except*). asyncio.gather leaves the other tasks running after the first exception. Wrap a block in async with asyncio.timeout(30): to bound it.

Free-threaded CPython (no GIL) is a separate build, python3.14t, officially supported from 3.14 under PEP 779 and experimental in 3.13. The standard build still has the GIL. Check with python3 -c 'import sys; print(sys._is_gil_enabled())' on the interpreter you ship, and expect C extensions without free-threading support to re-enable the GIL.

Packaging with uv#

[project]
name = "reconcile"
version = "1.4.2"
requires-python = ">=3.12"
dependencies = ["httpx>=0.27", "pydantic-settings>=2"]

[project.scripts]
reconcile = "reconcile.cli:main"

[dependency-groups]
dev = ["pytest>=8", "mypy>=1.10", "ruff"]

[build-system]
requires = ["uv_build>=0.12,<0.13"]      # `uv init` writes a matching range
build-backend = "uv_build"

[tool.ruff]
line-length = 100

[tool.ruff.lint]
select = ["E", "F", "I", "B", "UP", "S"]

[tool.mypy]
strict = true

[tool.pytest.ini_options]
addopts = "-q --strict-markers"
uv sync                            # create .venv if needed, install exactly what uv.lock lists
uv run pytest                      # runs in .venv, syncing first
uv add httpx                       # updates pyproject.toml and uv.lock
uv sync --locked                   # CI: error if uv.lock does not match pyproject.toml
uv sync --frozen                   # use uv.lock as-is without checking it (Docker layers)
uv sync --no-dev                   # production install without the dev group
uv build                           # sdist and wheel into dist/

--locked fails when the lock file is stale, which is what CI should check. --frozen skips the check entirely and installs whatever the lock says. uv sync removes packages not in the lock file; uv run does not unless given --exact (uv sync docs).

Commit uv.lock for applications. Libraries should also commit it for reproducible CI, but consumers resolve from dependencies, so keep those ranges honest.

Ruff replaces flake8, isort and black. ruff check --fix applies only safe fixes; --unsafe-fixes opts into the rest. Use ruff rule B006 to read what a rule checks. See the ruff docs for rule codes.

uv workflows#

uv resolves once into uv.lock (a cross-platform universal lock), then installs from it. The commands below cover the day-to-day loop beyond sync and add; run them from the project root, where uv discovers pyproject.toml.

uv python pin 3.13                      # writes .python-version; uv run and sync use that interpreter
uv python list --only-installed         # interpreters uv knows about
uv lock --upgrade                       # re-resolve everything to the newest allowed versions
uv lock --upgrade-package httpx         # bump one package within its constraint
uv tree                                 # dependency tree from the lock file
uv tree --invert --package certifi      # who depends on certifi
uv add 'httpx>=0.28,<1' --optional cli  # optional extra; install with uv sync --extra cli
uv add --group lint ruff                # a named dependency group; uv sync --group lint
uv sync --all-groups                    # every group, for a full local environment
uv run --with rich python -c 'import rich'   # temporarily add a package without touching the lock
uv run --no-sync pytest                 # skip the sync step when the environment is known good
uv export --format requirements.txt --no-dev -o requirements.txt   # for tools that only read requirements files
uv pip install -r requirements.txt      # pip-compatible interface into the active or project venv
uv cache prune --ci                     # drop pre-built wheels; keeps what CI reuses
uv build && uv publish --token "$PYPI_TOKEN"   # build then upload; use trusted publishing in CI instead of a token

In a container, copy pyproject.toml and uv.lock first and run uv sync --frozen --no-install-project --no-dev so the dependency layer caches independently of the source. Set UV_COMPILE_BYTECODE=1 and UV_LINK_MODE=copy in the image so start-up is fast and the cache mount does not leave hard links behind. uv refuses to install into a system interpreter without --system, which is the right guard everywhere except a throwaway image.

Typing#

Type hints do nothing at runtime and everything at review time: a checker turns “this returns None sometimes” into an error before the script runs. Annotate function boundaries and module-level data; let inference handle locals. Python 3.10+ syntax replaces most of the typing module: list[str] not List[str], X | None not Optional[X], TypeAlias and type X = ... (3.12) for aliases.

from collections.abc import Iterable, Iterator, Mapping, Sequence, Callable
from typing import Literal, Protocol, TypedDict, NewType, TypeVar, overload, assert_never

def first[T](items: Sequence[T], default: T) -> T:        # 3.12 generic syntax; TypeVar before that
    return items[0] if items else default

Zone = Literal["au-east", "au-west"]                       # only these strings type-check

class Row(TypedDict, total=False):                         # shape of a dict from JSON or csv.DictReader
    name: str
    ip: str
    zone: Zone

class Fetcher(Protocol):                                   # structural: any object with this method satisfies it
    def get(self, url: str, /) -> bytes: ...

HostId = NewType("HostId", int)                            # distinct from int to the checker, an int at runtime

def handle(state: Literal["up", "down"]) -> str:
    match state:
        case "up": return "ok"
        case "down": return "alert"
        case _: assert_never(state)                        # mypy errors if a Literal member is unhandled

Prefer collections.abc types for parameters (Iterable, Mapping) so callers can pass any suitable object, and concrete types for return values. Protocol replaces an abstract base class when you want to type a dependency you do not own, such as an HTTP client, and it is what makes a test fake fit without inheritance. Never annotate with Any to silence an error you do not understand; reveal_type(x) in a scratch file shows what the checker believes.

mypy configuration lives in pyproject.toml. strict = true enables the checks that matter (disallow_untyped_defs, warn_return_any, no_implicit_optional); relax per module for third-party code without stubs rather than globally.

[tool.mypy]
python_version = "3.12"
strict = true
warn_unreachable = true
files = ["src", "tests"]

[[tool.mypy.overrides]]
module = ["somelib.*"]
ignore_missing_imports = true
uv run mypy                              # uses [tool.mypy] files
uv run mypy --install-types --non-interactive   # fetch stub packages mypy suggests
uv run mypy --strict --warn-unused-ignores src   # find # type: ignore comments that no longer do anything

Dataclasses and pydantic#

Dataclasses are for data you construct in code and trust; pydantic is for data that crosses a trust boundary (HTTP bodies, config files, environment, queue messages) and must be validated and coerced. Using pydantic for internal structs costs validation time on every construction; using a dataclass for external input skips the check that stops garbage reaching your database.

from dataclasses import dataclass, field, asdict, replace

@dataclass(frozen=True, slots=True, kw_only=True)         # kw_only: callers must name every field
class Deploy:
    service: str
    version: str
    replicas: int = 2
    tags: tuple[str, ...] = ()
    labels: dict[str, str] = field(default_factory=dict)   # never a mutable default

    def __post_init__(self) -> None:
        if self.replicas < 1:
            raise ValueError("replicas must be >= 1")

d = Deploy(service="api", version="1.4.2")
d2 = replace(d, replicas=4)                                # copy with changes; frozen instances are immutable
asdict(d2)                                                 # nested dict, for json.dumps
from pydantic import BaseModel, Field, field_validator, ValidationError, TypeAdapter

class Item(BaseModel, frozen=True, extra="forbid"):        # extra="forbid": unknown keys are an error
    id: int
    name: str = Field(min_length=1, max_length=64)
    price_cents: int = Field(ge=0)
    tags: list[str] = []

    @field_validator("tags")
    @classmethod
    def lowercase_tags(cls, v: list[str]) -> list[str]:
        return [t.lower() for t in v]

try:
    item = Item.model_validate_json(body)                  # bytes or str straight from the wire
except ValidationError as e:
    log.warning("bad payload: %s", e.errors())             # a list of {loc, msg, type}
    raise

items = TypeAdapter(list[Item]).validate_python(payload)   # validate a plain list without a wrapper model
item.model_dump(mode="json")                               # dict with JSON-safe types (datetimes as strings)

pydantic v2 coerces by default ("3" becomes 3 for an int field); pass strict=True to the model config or a field to refuse it. model_dump(exclude={"password"}) keeps secrets out of logs, and SecretStr (see Configuration and secrets) does it structurally.

asyncio#

The event loop runs one coroutine at a time and switches only at await. A blocking call (time.sleep, requests.get, a CPU loop, a synchronous database driver) freezes every other task, which is the cause of most “asyncio is slow” reports. Move blocking work to a thread with asyncio.to_thread and keep the loop free.

import asyncio

async def main() -> None:
    async with asyncio.timeout(60):                           # 3.11+: cancels the block when exceeded
        data = await asyncio.to_thread(read_big_file, path)   # blocking call in the default executor
        async with asyncio.TaskGroup() as tg:
            t1 = tg.create_task(fetch(data.url), name="fetch")
            t2 = tg.create_task(publish(data), name="publish")
    print(t1.result(), t2.result())

asyncio.run(main())
queue: asyncio.Queue[Job] = asyncio.Queue(maxsize=100)      # bounded: producers block when consumers lag

async def worker(name: str) -> None:
    while True:
        job = await queue.get()
        try:
            await process(job)
        finally:
            queue.task_done()                                 # pairs with queue.join()

async def run(jobs: list[Job]) -> None:
    async with asyncio.TaskGroup() as tg:
        workers = [tg.create_task(worker(f"w{i}")) for i in range(8)]
        for j in jobs:
            await queue.put(j)
        await queue.join()                                    # every put has had a task_done
        for w in workers:
            w.cancel()                                        # TaskGroup absorbs the CancelledError

Cancellation is cooperative: task.cancel() raises CancelledError at the task’s next await, so a coroutine that never awaits cannot be cancelled. Catch CancelledError only to clean up, then re-raise. asyncio.run(..., debug=True) (or PYTHONASYNCIODEBUG=1) logs coroutines that block the loop for more than 100 ms and tasks that were never awaited. asyncio.Semaphore bounds concurrency; asyncio.Lock protects a resource across awaits; neither is needed for plain attribute updates between awaits because no other task runs in between.

Logging configuration#

basicConfig is right for a script. A service needs per-logger levels, a handler per destination and a formatter, which logging.config.dictConfig sets up from one dict that can live in YAML or TOML.

import logging.config

logging.config.dictConfig({
    "version": 1,
    "disable_existing_loggers": False,                    # keep loggers created at import time
    "formatters": {
        "json": {"()": "myapp.logging.JsonFormatter"},    # "()" names a factory
        "plain": {"format": "%(asctime)s %(levelname)-8s %(name)s: %(message)s"},
    },
    "handlers": {
        "stderr": {"class": "logging.StreamHandler", "formatter": "json", "stream": "ext://sys.stderr"},
        "file": {
            "class": "logging.handlers.RotatingFileHandler",
            "filename": "/var/log/myapp/app.log",
            "maxBytes": 50_000_000, "backupCount": 5, "formatter": "plain",
        },
    },
    "loggers": {
        "httpx": {"level": "WARNING"},                    # quieten a chatty library
        "myapp.db": {"level": "DEBUG"},
    },
    "root": {"level": "INFO", "handlers": ["stderr", "file"]},
})

Loggers form a tree by dotted name, so logging.getLogger(__name__) in each module gives you per-package control for free. Records propagate up to the root’s handlers; set propagate: False on a logger only when it has its own handler and you want to stop duplicates. Under systemd or a container, log to stderr only and let the platform handle rotation. logging.captureWarnings(True) routes warnings.warn through logging. For a QueueHandler that keeps slow handlers off the request thread, 3.12 adds "respect_handler_level" and a queue key in dictConfig.

pytest fixtures and parametrize#

A fixture is a function whose return value is injected into any test that names it as a parameter; yield splits setup from teardown. Scope controls how often it runs (function default, module, session). Put shared fixtures in conftest.py at the level of the tests that use them, not one giant file at the root.

import pytest

@pytest.fixture
def tmp_config(tmp_path: Path) -> Path:                       # tmp_path is a built-in fixture
    p = tmp_path / "app.yaml"
    p.write_text("timeout: 5\n", encoding="utf-8")
    return p

@pytest.fixture(scope="session")
def db_url() -> Iterator[str]:
    container = start_postgres()                              # once for the whole run
    yield container.url
    container.stop()                                          # teardown after the last test

@pytest.fixture(autouse=True)
def no_network(monkeypatch: pytest.MonkeyPatch) -> None:      # applies to every test in scope
    monkeypatch.delenv("HTTP_PROXY", raising=False)

@pytest.mark.parametrize(
    ("raw", "expected"),
    [
        ("5s", 5.0),
        ("2m", 120.0),
        pytest.param("", None, id="empty"),
        pytest.param("bad", None, marks=pytest.mark.xfail(raises=ValueError, strict=True)),
    ],
)
def test_parse_duration(raw: str, expected: float | None) -> None:
    assert parse_duration(raw) == expected
uv run pytest -k 'parse and not slow'       # expression over test names and markers
uv run pytest --lf                          # only tests that failed last time
uv run pytest -x --pdb                      # drop into the debugger at the first failure
uv run pytest --fixtures tests/             # every fixture available, with docstrings
uv run pytest --setup-show tests/test_x.py  # print fixture setup and teardown order
uv run pytest -p no:cacheprovider -q        # no .pytest_cache, for read-only CI checkouts

monkeypatch reverts environment, attributes and sys.path changes after each test, unlike a bare os.environ[...] = .... capsys captures stdout and stderr; caplog captures log records and lets you assert on caplog.records. Mark slow or network tests (@pytest.mark.slow) and register the marker under [tool.pytest.ini_options] markers so --strict-markers catches typos. See Testing for what to mock and what to leave real.

Useful one-liners#

# Serve the current directory on localhost only
python3 -m http.server 8000 --bind 127.0.0.1

# Pretty-print JSON
python3 -m json.tool < data.json

# Time a snippet over many runs
python3 -m timeit -s 'import json' 'json.dumps({"a":1})'

# Profile a script, slowest cumulative first
python3 -m cProfile -s cumtime script.py | head -25

# Where a module is loaded from and its version
python3 -c 'import httpx, inspect; print(httpx.__version__, inspect.getfile(httpx))'

# Base64 and URL encoding
python3 -c 'import base64,sys; print(base64.b64encode(sys.stdin.buffer.read()).decode())'
python3 -c 'import urllib.parse,sys; print(urllib.parse.quote(sys.argv[1]))' 'a b&c'

# Epoch to ISO 8601 in UTC (datetime.UTC needs 3.11+)
python3 -c 'import sys,datetime as d; print(d.datetime.fromtimestamp(int(sys.argv[1]), d.UTC).isoformat())' 1700000000
2023-11-14T22:13:20+00:00
# Validate a YAML file (needs PyYAML)
uv run --with pyyaml python -c 'import sys,yaml; yaml.safe_load(open(sys.argv[1]))' config.yaml

# Which distribution provides an import name
python3 -c 'import importlib.metadata as m; print(m.packages_distributions()["yaml"])'

# Check a wheel's metadata before publishing
uvx twine check dist/*

# Every third-party import in the tree, to compare against pyproject dependencies
grep -rhoE '^(from|import) [a-zA-Z_]+' src | awk '{print $2}' | sort -u

# Run a module's tests with warnings turned into errors
uv run pytest -W error::DeprecationWarning tests/

# Dump every environment variable a Settings class would read
python3 -c 'from myapp.config import Settings; print(Settings.model_json_schema()["properties"].keys())'

# Byte-compile everything to catch syntax errors without importing
python3 -m compileall -q src/

# Show the resolved dependency tree from the lock file
uv tree --depth 1

# Start a REPL with the project importable and a variable pre-loaded
uv run python -i -c 'from myapp.config import Settings; s = Settings()'

Snippets#

Retry with exponential backoff and jitter without a library, for the case where tenacity is one dependency too many.

import random, time

def retry(fn, *, attempts: int = 5, base: float = 0.5, cap: float = 10.0, retry_on=(OSError,)):
    for i in range(attempts):
        try:
            return fn()
        except retry_on as e:
            if i == attempts - 1:
                raise
            delay = min(cap, base * 2**i) * random.uniform(0.5, 1.5)
            log.warning("attempt %d failed (%s), retrying in %.1fs", i + 1, e, delay)
            time.sleep(delay)

HTTP client with separate connect and read timeouts, limits and a retrying transport.

client = httpx.Client(
    base_url="https://api.example.com",
    timeout=httpx.Timeout(connect=3.0, read=15.0, write=10.0, pool=5.0),
    limits=httpx.Limits(max_connections=50, max_keepalive_connections=10),
    transport=httpx.HTTPTransport(retries=2),          # retries connection errors only, not HTTP status codes
    headers={"authorization": f"Bearer {token}"},
)

Graceful shutdown: finish the current item on SIGTERM, then exit.

import signal

stop = False
def _handle(signum, frame):
    global stop
    stop = True
    log.info("signal %s received, finishing current item", signal.Signals(signum).name)

signal.signal(signal.SIGTERM, _handle)
signal.signal(signal.SIGINT, _handle)
for item in work:
    if stop:
        break
    process(item)

The same for asyncio, cancelling the main task so finally blocks run.

async def main() -> None:
    loop = asyncio.get_running_loop()
    task = asyncio.current_task()
    for sig in (signal.SIGTERM, signal.SIGINT):
        loop.add_signal_handler(sig, task.cancel)
    try:
        await serve()
    except asyncio.CancelledError:
        await drain()                                   # bounded cleanup, then let the cancellation finish
        raise

Configuration layering: defaults, then a TOML file, then environment, with the environment winning.

import os, tomllib

DEFAULTS = {"timeout": 10.0, "workers": 4, "api_url": ""}

def load(path: Path) -> dict:
    cfg = dict(DEFAULTS)
    if path.is_file():
        with path.open("rb") as f:                      # tomllib needs binary mode
            cfg.update(tomllib.load(f))
    for k in cfg:
        if (v := os.getenv(f"APP_{k.upper()}")) is not None:
            cfg[k] = type(DEFAULTS[k])(v)               # coerce to the default's type
    return cfg

Worker pool with bounded queue and per-item error handling, using threads.

from concurrent.futures import ThreadPoolExecutor

def run_all(items, fn, workers: int = 8) -> tuple[int, int]:
    ok = failed = 0
    with ThreadPoolExecutor(max_workers=workers) as pool:
        for item, result in zip(items, pool.map(lambda i: _safe(fn, i), items)):
            if isinstance(result, Exception):
                failed += 1
                log.error("item %s failed: %s", item, result)
            else:
                ok += 1
    return ok, failed

def _safe(fn, item):
    try:
        return fn(item)
    except Exception as e:                              # captured, not raised, so the pool keeps going
        return e

Context manager that times a block and logs the duration once.

from contextlib import contextmanager
import time

@contextmanager
def timed(label: str):
    start = time.perf_counter()
    try:
        yield
    finally:
        log.info("%s took %.3fs", label, time.perf_counter() - start)

with timed("reconcile"):
    reconcile()

Structured log context that follows the request through every log call, using contextvars.

import contextvars, logging

request_id: contextvars.ContextVar[str] = contextvars.ContextVar("request_id", default="-")

class ContextFilter(logging.Filter):
    def filter(self, record: logging.LogRecord) -> bool:
        record.request_id = request_id.get()            # available to formatters as %(request_id)s
        return True

logging.getLogger().addFilter(ContextFilter())
token = request_id.set(incoming_id)                     # per request; asyncio tasks inherit a copy

Stream a large download to disk without holding it in memory, and verify the checksum.

import hashlib

def download(client: httpx.Client, url: str, dest: Path, sha256: str) -> None:
    h = hashlib.sha256()
    tmp = dest.with_suffix(dest.suffix + ".part")
    with client.stream("GET", url) as r, tmp.open("wb") as f:
        r.raise_for_status()
        for chunk in r.iter_bytes(1 << 20):
            f.write(chunk)
            h.update(chunk)
    if h.hexdigest() != sha256:
        tmp.unlink()
        raise ValueError(f"checksum mismatch for {url}")
    tmp.replace(dest)

Paginate an API until the cursor runs out, as a generator.

def iter_pages(client: httpx.Client, path: str, **params):
    cursor = None
    while True:
        r = client.get(path, params={**params, "cursor": cursor})
        r.raise_for_status()
        body = r.json()
        yield from body["items"]
        cursor = body.get("next_cursor")
        if not cursor:
            return

Table-driven test with a fake HTTP transport, no network and no mocking library.

@pytest.mark.parametrize(
    ("status", "body", "expected"),
    [(200, {"items": [1, 2]}, [1, 2]), (404, {}, []), (500, {}, None)],
)
def test_fetch_items(status: int, body: dict, expected):
    def handler(request: httpx.Request) -> httpx.Response:
        return httpx.Response(status, json=body)

    client = httpx.Client(transport=httpx.MockTransport(handler))
    if expected is None:
        with pytest.raises(httpx.HTTPStatusError):
            fetch_items(client)
    else:
        assert fetch_items(client) == expected

Ruff configuration that covers the rule families worth enforcing on operational code.

[tool.ruff]
line-length = 100
target-version = "py312"
src = ["src", "tests"]

[tool.ruff.lint]
select = ["E", "W", "F", "I", "B", "UP", "S", "PTH", "RUF", "SIM", "ASYNC", "LOG", "T20"]
ignore = ["S101"]                      # assert is fine in tests, and tests are where it appears
per-file-ignores = { "tests/**" = ["S", "T20"] }

[tool.ruff.lint.isort]
known-first-party = ["myapp"]

[tool.ruff.format]
docstring-code-format = true

Argument parser with subcommands, each dispatching to its own function.

def main(argv: list[str] | None = None) -> int:
    ap = argparse.ArgumentParser(prog="ops")
    sub = ap.add_subparsers(dest="cmd", required=True)
    p = sub.add_parser("sync", help="reconcile inventory"); p.add_argument("path", type=Path); p.set_defaults(fn=cmd_sync)
    p = sub.add_parser("report", help="print a summary"); p.add_argument("--json", action="store_true"); p.set_defaults(fn=cmd_report)
    args = ap.parse_args(argv)
    return args.fn(args)

Troubleshooting#

SymptomLikely causeCheck or fix
ModuleNotFoundError for an installed packageDifferent interpreter from the one you installed intopython3 -c 'import sys; print(sys.executable)'; use uv run or activate .venv
error: externally-managed-environment from pipDistro Python protects system packages (PEP 668)Use a venv, uv tool install, or uvx; do not pass --break-system-packages
uv sync --locked fails in CIpyproject.toml changed without re-lockingRun uv lock locally and commit uv.lock
Script hangs on a network callNo timeout (requests, raw sockets)Add timeout=; dump stacks with py-spy dump --pid <pid>
Script hangs on subprocess.runChild waits on stdin or never exitsPass stdin=subprocess.DEVNULL and timeout=
UnicodeDecodeError on one host onlyLocale-dependent default encodingPass encoding="utf-8", or run with PYTHONUTF8=1
RuntimeError: This event loop is already runningasyncio.run inside Jupyter or another loopawait the coroutine directly
coroutine ... was never awaited warningCalled an async def without awaitAdd await, or schedule it with a TaskGroup
Memory grows until the process is killedReading a whole file or response into memoryIterate lines or use client.stream(); profile with tracemalloc
Threads give no speed-upCPU-bound work held by the GILProcessPoolExecutor, or a free-threaded build
KeyError: "Attempt to overwrite 'name' in LogRecord"extra key collides with a built-in attributeRename the key or nest under one key
Task exception was never retrievedA fire-and-forget create_task raised and nothing awaited itKeep a reference and await it, or use TaskGroup
asyncio program is slow and one call dominatesA blocking call inside a coroutinePYTHONASYNCIODEBUG=1 logs slow callbacks; wrap in asyncio.to_thread
mypy: Incompatible types on a dict from JSONUntyped dict[str, Any] flowing into typed codeValidate with pydantic or a TypedDict at the boundary
mypy: Cannot find implementation or library stubPackage ships no types--install-types, or ignore_missing_imports in an override for that module
pydantic accepts "3" for an int fieldLax mode coerces by defaultmodel_config = ConfigDict(strict=True) or Field(strict=True)
ValueError: mutable default ... use default_factoryList or dict as a dataclass defaultfield(default_factory=list)
Fixture not foundconftest.py is in a sibling directory, or a plugin is disabledpytest --fixtures; move the fixture up the tree
Same test fails only when run with othersShared module-level state or fixture scope too widepytest -p randomly or --lf; narrow the scope to function
uv picks the wrong interpreterNo .python-version and several installed Pythonsuv python pin 3.13; check uv python find
Slow container start-upBytecode compiled on every launchUV_COMPILE_BYTECODE=1 at build time

For a live process that is stuck or slow, py-spy dump --pid <pid> prints every thread’s stack and py-spy top --pid <pid> samples hot functions, without restarting it. Both need permission to ptrace the process (root, or the same user with kernel.yama.ptrace_scope=0). python3 -X faulthandler script.py prints tracebacks on a crash or on SIGABRT.