Python
Write operational Python scripts that fail safely: subprocess, paths, logging, HTTP, configuration, concurrency, uv packaging and troubleshooting.
On this page
Cheatsheet#
| Task | Snippet or command |
|---|---|
| Run a command, raise on failure | subprocess.run(cmd, check=True, capture_output=True, text=True, timeout=60) |
| Build a path | Path("/etc") / "app.conf" |
| Atomic file write | write a temp file in the same directory, then os.replace(tmp, target) |
| Log with the traceback | log.exception("sync failed") inside except |
| HTTP with a timeout | httpx.get(url, timeout=10) |
| Retry with backoff | tenacity.retry(wait=wait_exponential(), stop=stop_after_attempt(5)) |
| Typed config from environment | pydantic_settings.BaseSettings |
| Temporary directory | with tempfile.TemporaryDirectory() as d: |
| Parse CLI arguments | argparse.ArgumentParser |
| Time a block | t = time.perf_counter(); ...; time.perf_counter() - t |
| Create a project | uv init --package my-app |
| Add a dependency | uv add httpx / uv add --dev pytest |
| Run in the project environment | uv run pytest |
| CI install, fail if lock is stale | uv sync --locked |
| Upgrade one locked package | uv lock --upgrade-package httpx |
| Run a tool without installing it | uvx ruff check . |
| Install a CLI tool globally | uv tool install ruff |
| Install a Python version | uv python install 3.14 |
| Format and lint | ruff format . && ruff check --fix . |
| Type-check | mypy --strict src/ |
| Run tests | pytest -x -q (see Testing) |
A script worth keeping#
#!/usr/bin/env python3
"""Reconcile inventory against the API."""
from __future__ import annotations
import argparse
import logging
import sys
from pathlib import Path
log = logging.getLogger("reconcile")
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("inventory", type=Path)
ap.add_argument("--dry-run", action="store_true")
ap.add_argument("-v", "--verbose", action="count", default=0)
args = ap.parse_args()
logging.basicConfig(
level=logging.DEBUG if args.verbose else logging.INFO,
format="%(asctime)s %(levelname)s %(name)s %(message)s",
stream=sys.stderr,
)
if not args.inventory.is_file():
log.error("inventory not found: %s", args.inventory)
return 2
log.info("reconciling %s dry_run=%s", args.inventory, args.dry_run)
return 0
if __name__ == "__main__":
raise SystemExit(main())main returns an exit code, logs go to stderr, and stdout is reserved for data. That lets the script sit in a pipeline and makes failures visible in CI. Pass arguments to the logger (log.info("x=%s", x)) instead of an f-string so formatting only happens when the level is enabled.
A single-file script can declare its own dependencies (PEP 723). uv run script.py reads the block and runs it in a cached environment:
# /// script
# requires-python = ">=3.12"
# dependencies = ["httpx"]
# ///uv add --script reconcile.py httpx # writes the block for you
uv run reconcile.py inventory.csvRunning commands with subprocess#
import json, subprocess
res = subprocess.run(
["kubectl", "get", "pods", "-o", "json"],
check=True, capture_output=True, text=True, timeout=30,
)
pods = json.loads(res.stdout)| Argument | Why |
|---|---|
| List, not a string | No shell, so no quoting problems and no injection |
check=True | Raises CalledProcessError on a non-zero exit instead of continuing |
capture_output=True, text=True | stdout and stderr as str |
timeout= | Kills the child and raises TimeoutExpired instead of hanging the job |
cwd=, env= | Explicit context. env= replaces the whole environment, so start from {**os.environ, ...} |
shell=True is only safe with a literal string you wrote. With any interpolated value it is a command injection bug. If a shell is truly needed, quote each value with shlex.quote.
try:
subprocess.run(cmd, check=True, capture_output=True, text=True, timeout=60)
except subprocess.CalledProcessError as e:
log.error("command failed rc=%s stderr=%s", e.returncode, e.stderr.strip())
raise
except subprocess.TimeoutExpired:
log.error("command timed out: %s", shlex.join(cmd))
raiseFor long-running output, stream it instead of buffering everything in memory:
with subprocess.Popen(cmd, stdout=subprocess.PIPE, text=True) as p:
for line in p.stdout:
handle(line)
if p.returncode != 0:
raise subprocess.CalledProcessError(p.returncode, cmd)Paths and files#
from pathlib import Path
base = Path("/srv/app")
cfg = base / "conf" / "app.yaml"
cfg.exists(); cfg.stat().st_size; cfg.read_text(encoding="utf-8")
list(base.rglob("*.log"))
base.mkdir(parents=True, exist_ok=True)Always pass encoding="utf-8" when reading or writing text. The default comes from the locale, so the same script can behave differently on another host (Python 3.15 changes the default to UTF-8 under PEP 686).
Write atomically so a crash cannot leave a half-written file where a complete one is expected:
import os, tempfile
def write_atomic(path: Path, data: str) -> None:
fd, tmp = tempfile.mkstemp(dir=path.parent) # same filesystem as the target
try:
with os.fdopen(fd, "w", encoding="utf-8") as f:
f.write(data)
f.flush()
os.fsync(f.fileno()) # data on disk before the rename
os.replace(tmp, path) # atomic rename on POSIX
except BaseException:
os.unlink(tmp)
raiseos.replace is only atomic within one filesystem, which is why the temp file is created in the target’s directory. mkstemp creates the file with mode 0600; chmod it if other users need to read it.
Logging#
import json, logging
class JsonFormatter(logging.Formatter):
def format(self, record: logging.LogRecord) -> str:
payload = {
"ts": self.formatTime(record, "%Y-%m-%dT%H:%M:%S%z"),
"level": record.levelname,
"logger": record.name,
"msg": record.getMessage(),
**getattr(record, "fields", {}),
}
if record.exc_info:
payload["exc"] = self.formatException(record.exc_info)
return json.dumps(payload)
handler = logging.StreamHandler()
handler.setFormatter(JsonFormatter())
logging.basicConfig(level=logging.INFO, handlers=[handler])
log.info("deployed", extra={"fields": {"service": "api", "version": "1.4.2"}})extra sets attributes on the LogRecord, so a key that collides with a built-in attribute (msg, args, name) raises KeyError. Nesting under one key avoids that. Call log.exception() inside an except block to record the traceback. Do not log secrets, tokens or full request bodies.
HTTP clients, timeouts and retries#
import httpx
with httpx.Client(timeout=10.0, headers={"user-agent": "reconcile/1.0"}) as client:
r = client.get("https://api.example.com/items", params={"limit": 100})
r.raise_for_status()
items = r.json()httpx applies a 5 second timeout by default. requests has no default timeout and can wait forever on a stalled connection, so always pass timeout=. Reuse one Client for many requests to keep connection pooling.
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception
def transient(e: BaseException) -> bool:
if isinstance(e, httpx.TransportError): # timeouts, connection resets
return True
return isinstance(e, httpx.HTTPStatusError) and e.response.status_code in {429, 502, 503, 504}
@retry(
stop=stop_after_attempt(5),
wait=wait_exponential(multiplier=0.5, max=10),
retry=retry_if_exception(transient),
reraise=True,
)
def fetch(client: httpx.Client, url: str) -> dict:
r = client.get(url)
r.raise_for_status()
return r.json()Retry only transient failures. Retrying every HTTPStatusError also retries 400 and 404, which never succeed. Retry idempotent requests only: a retried POST can create two records. See HTTP for status code meanings.
Configuration and secrets#
from pydantic import SecretStr
from pydantic_settings import BaseSettings, SettingsConfigDict
class Settings(BaseSettings):
model_config = SettingsConfigDict(env_prefix="APP_", env_file=".env")
api_url: str
api_token: SecretStr # printed as '**********'
timeout: float = 10.0
dry_run: bool = False
settings = Settings() # raises ValidationError if APP_API_URL is missing
token = settings.api_token.get_secret_value() # explicit access to the real valueValidate configuration once at start-up and exit on failure. A missing variable found three hours into a batch job costs more than a crash on line one. Keep .env out of version control. See Vault for fetching secrets at runtime.
Data handling#
from dataclasses import dataclass
@dataclass(frozen=True, slots=True)
class Host:
name: str
ip: str
zone: str
tags: tuple[str, ...] = ()
from collections import Counter, defaultdict
counts = Counter(h.tags[0] for h in hosts if h.tags)
by_zone: dict[str, list[Host]] = defaultdict(list)
for h in hosts:
by_zone[h.zone].append(h)
import csv
with open("hosts.csv", newline="", encoding="utf-8") as f:
rows = list(csv.DictReader(f))frozen=True makes instances hashable and blocks accidental mutation. slots=True removes the per-instance __dict__, which saves memory once there are hundreds of thousands of objects. Use a tuple, not a list, for fields of a frozen dataclass, or hashing fails.
Generators keep memory flat for large inputs:
def read_events(path: Path):
with path.open(encoding="utf-8") as f:
for line in f: # one line at a time, not the whole file
yield json.loads(line)Concurrency: threads, asyncio or processes#
| Workload | Tool |
|---|---|
| Network I/O, tens of calls | ThreadPoolExecutor |
| Network I/O, thousands of calls | asyncio with httpx.AsyncClient |
| CPU-bound pure Python | ProcessPoolExecutor |
| Running other programs | Threads; the GIL is released while waiting on the child |
from concurrent.futures import ThreadPoolExecutor, as_completed
with ThreadPoolExecutor(max_workers=8) as pool:
futures = {pool.submit(check_host, h): h for h in hosts}
for fut in as_completed(futures):
host = futures[fut]
try:
result = fut.result()
except Exception:
log.exception("check failed host=%s", host.name)import asyncio, httpx
async def fetch_all(urls: list[str]) -> list[dict]:
limit = asyncio.Semaphore(20) # at most 20 requests in flight
async with httpx.AsyncClient(timeout=10) as client:
async def one(u: str) -> dict:
async with limit:
r = await client.get(u)
r.raise_for_status()
return r.json()
async with asyncio.TaskGroup() as tg: # 3.11+: cancels the rest if one fails
tasks = [tg.create_task(one(u)) for u in urls]
return [t.result() for t in tasks]
results = asyncio.run(fetch_all(urls))TaskGroup cancels sibling tasks when one raises and reports failures as an ExceptionGroup (catch with except*). asyncio.gather leaves the other tasks running after the first exception. Wrap a block in async with asyncio.timeout(30): to bound it.
Free-threaded CPython (no GIL) is a separate build, python3.14t, officially supported from 3.14 under PEP 779 and experimental in 3.13. The standard build still has the GIL. Check with python3 -c 'import sys; print(sys._is_gil_enabled())' on the interpreter you ship, and expect C extensions without free-threading support to re-enable the GIL.
Packaging with uv#
[project]
name = "reconcile"
version = "1.4.2"
requires-python = ">=3.12"
dependencies = ["httpx>=0.27", "pydantic-settings>=2"]
[project.scripts]
reconcile = "reconcile.cli:main"
[dependency-groups]
dev = ["pytest>=8", "mypy>=1.10", "ruff"]
[build-system]
requires = ["uv_build>=0.12,<0.13"] # `uv init` writes a matching range
build-backend = "uv_build"
[tool.ruff]
line-length = 100
[tool.ruff.lint]
select = ["E", "F", "I", "B", "UP", "S"]
[tool.mypy]
strict = true
[tool.pytest.ini_options]
addopts = "-q --strict-markers"uv sync # create .venv if needed, install exactly what uv.lock lists
uv run pytest # runs in .venv, syncing first
uv add httpx # updates pyproject.toml and uv.lock
uv sync --locked # CI: error if uv.lock does not match pyproject.toml
uv sync --frozen # use uv.lock as-is without checking it (Docker layers)
uv sync --no-dev # production install without the dev group
uv build # sdist and wheel into dist/--locked fails when the lock file is stale, which is what CI should check. --frozen skips the check entirely and installs whatever the lock says. uv sync removes packages not in the lock file; uv run does not unless given --exact (uv sync docs).
Commit uv.lock for applications. Libraries should also commit it for reproducible CI, but consumers resolve from dependencies, so keep those ranges honest.
Ruff replaces flake8, isort and black. ruff check --fix applies only safe fixes; --unsafe-fixes opts into the rest. Use ruff rule B006 to read what a rule checks. See the ruff docs for rule codes.
uv workflows#
uv resolves once into uv.lock (a cross-platform universal lock), then installs from it. The commands below cover the day-to-day loop beyond sync and add; run them from the project root, where uv discovers pyproject.toml.
uv python pin 3.13 # writes .python-version; uv run and sync use that interpreter
uv python list --only-installed # interpreters uv knows about
uv lock --upgrade # re-resolve everything to the newest allowed versions
uv lock --upgrade-package httpx # bump one package within its constraint
uv tree # dependency tree from the lock file
uv tree --invert --package certifi # who depends on certifi
uv add 'httpx>=0.28,<1' --optional cli # optional extra; install with uv sync --extra cli
uv add --group lint ruff # a named dependency group; uv sync --group lint
uv sync --all-groups # every group, for a full local environment
uv run --with rich python -c 'import rich' # temporarily add a package without touching the lock
uv run --no-sync pytest # skip the sync step when the environment is known good
uv export --format requirements.txt --no-dev -o requirements.txt # for tools that only read requirements files
uv pip install -r requirements.txt # pip-compatible interface into the active or project venv
uv cache prune --ci # drop pre-built wheels; keeps what CI reuses
uv build && uv publish --token "$PYPI_TOKEN" # build then upload; use trusted publishing in CI instead of a tokenIn a container, copy pyproject.toml and uv.lock first and run uv sync --frozen --no-install-project --no-dev so the dependency layer caches independently of the source. Set UV_COMPILE_BYTECODE=1 and UV_LINK_MODE=copy in the image so start-up is fast and the cache mount does not leave hard links behind. uv refuses to install into a system interpreter without --system, which is the right guard everywhere except a throwaway image.
Typing#
Type hints do nothing at runtime and everything at review time: a checker turns “this returns None sometimes” into an error before the script runs. Annotate function boundaries and module-level data; let inference handle locals. Python 3.10+ syntax replaces most of the typing module: list[str] not List[str], X | None not Optional[X], TypeAlias and type X = ... (3.12) for aliases.
from collections.abc import Iterable, Iterator, Mapping, Sequence, Callable
from typing import Literal, Protocol, TypedDict, NewType, TypeVar, overload, assert_never
def first[T](items: Sequence[T], default: T) -> T: # 3.12 generic syntax; TypeVar before that
return items[0] if items else default
Zone = Literal["au-east", "au-west"] # only these strings type-check
class Row(TypedDict, total=False): # shape of a dict from JSON or csv.DictReader
name: str
ip: str
zone: Zone
class Fetcher(Protocol): # structural: any object with this method satisfies it
def get(self, url: str, /) -> bytes: ...
HostId = NewType("HostId", int) # distinct from int to the checker, an int at runtime
def handle(state: Literal["up", "down"]) -> str:
match state:
case "up": return "ok"
case "down": return "alert"
case _: assert_never(state) # mypy errors if a Literal member is unhandledPrefer collections.abc types for parameters (Iterable, Mapping) so callers can pass any suitable object, and concrete types for return values. Protocol replaces an abstract base class when you want to type a dependency you do not own, such as an HTTP client, and it is what makes a test fake fit without inheritance. Never annotate with Any to silence an error you do not understand; reveal_type(x) in a scratch file shows what the checker believes.
mypy configuration lives in pyproject.toml. strict = true enables the checks that matter (disallow_untyped_defs, warn_return_any, no_implicit_optional); relax per module for third-party code without stubs rather than globally.
[tool.mypy]
python_version = "3.12"
strict = true
warn_unreachable = true
files = ["src", "tests"]
[[tool.mypy.overrides]]
module = ["somelib.*"]
ignore_missing_imports = trueuv run mypy # uses [tool.mypy] files
uv run mypy --install-types --non-interactive # fetch stub packages mypy suggests
uv run mypy --strict --warn-unused-ignores src # find # type: ignore comments that no longer do anythingDataclasses and pydantic#
Dataclasses are for data you construct in code and trust; pydantic is for data that crosses a trust boundary (HTTP bodies, config files, environment, queue messages) and must be validated and coerced. Using pydantic for internal structs costs validation time on every construction; using a dataclass for external input skips the check that stops garbage reaching your database.
from dataclasses import dataclass, field, asdict, replace
@dataclass(frozen=True, slots=True, kw_only=True) # kw_only: callers must name every field
class Deploy:
service: str
version: str
replicas: int = 2
tags: tuple[str, ...] = ()
labels: dict[str, str] = field(default_factory=dict) # never a mutable default
def __post_init__(self) -> None:
if self.replicas < 1:
raise ValueError("replicas must be >= 1")
d = Deploy(service="api", version="1.4.2")
d2 = replace(d, replicas=4) # copy with changes; frozen instances are immutable
asdict(d2) # nested dict, for json.dumpsfrom pydantic import BaseModel, Field, field_validator, ValidationError, TypeAdapter
class Item(BaseModel, frozen=True, extra="forbid"): # extra="forbid": unknown keys are an error
id: int
name: str = Field(min_length=1, max_length=64)
price_cents: int = Field(ge=0)
tags: list[str] = []
@field_validator("tags")
@classmethod
def lowercase_tags(cls, v: list[str]) -> list[str]:
return [t.lower() for t in v]
try:
item = Item.model_validate_json(body) # bytes or str straight from the wire
except ValidationError as e:
log.warning("bad payload: %s", e.errors()) # a list of {loc, msg, type}
raise
items = TypeAdapter(list[Item]).validate_python(payload) # validate a plain list without a wrapper model
item.model_dump(mode="json") # dict with JSON-safe types (datetimes as strings)pydantic v2 coerces by default ("3" becomes 3 for an int field); pass strict=True to the model config or a field to refuse it. model_dump(exclude={"password"}) keeps secrets out of logs, and SecretStr (see Configuration and secrets) does it structurally.
asyncio#
The event loop runs one coroutine at a time and switches only at await. A blocking call (time.sleep, requests.get, a CPU loop, a synchronous database driver) freezes every other task, which is the cause of most “asyncio is slow” reports. Move blocking work to a thread with asyncio.to_thread and keep the loop free.
import asyncio
async def main() -> None:
async with asyncio.timeout(60): # 3.11+: cancels the block when exceeded
data = await asyncio.to_thread(read_big_file, path) # blocking call in the default executor
async with asyncio.TaskGroup() as tg:
t1 = tg.create_task(fetch(data.url), name="fetch")
t2 = tg.create_task(publish(data), name="publish")
print(t1.result(), t2.result())
asyncio.run(main())queue: asyncio.Queue[Job] = asyncio.Queue(maxsize=100) # bounded: producers block when consumers lag
async def worker(name: str) -> None:
while True:
job = await queue.get()
try:
await process(job)
finally:
queue.task_done() # pairs with queue.join()
async def run(jobs: list[Job]) -> None:
async with asyncio.TaskGroup() as tg:
workers = [tg.create_task(worker(f"w{i}")) for i in range(8)]
for j in jobs:
await queue.put(j)
await queue.join() # every put has had a task_done
for w in workers:
w.cancel() # TaskGroup absorbs the CancelledErrorCancellation is cooperative: task.cancel() raises CancelledError at the task’s next await, so a coroutine that never awaits cannot be cancelled. Catch CancelledError only to clean up, then re-raise. asyncio.run(..., debug=True) (or PYTHONASYNCIODEBUG=1) logs coroutines that block the loop for more than 100 ms and tasks that were never awaited. asyncio.Semaphore bounds concurrency; asyncio.Lock protects a resource across awaits; neither is needed for plain attribute updates between awaits because no other task runs in between.
Logging configuration#
basicConfig is right for a script. A service needs per-logger levels, a handler per destination and a formatter, which logging.config.dictConfig sets up from one dict that can live in YAML or TOML.
import logging.config
logging.config.dictConfig({
"version": 1,
"disable_existing_loggers": False, # keep loggers created at import time
"formatters": {
"json": {"()": "myapp.logging.JsonFormatter"}, # "()" names a factory
"plain": {"format": "%(asctime)s %(levelname)-8s %(name)s: %(message)s"},
},
"handlers": {
"stderr": {"class": "logging.StreamHandler", "formatter": "json", "stream": "ext://sys.stderr"},
"file": {
"class": "logging.handlers.RotatingFileHandler",
"filename": "/var/log/myapp/app.log",
"maxBytes": 50_000_000, "backupCount": 5, "formatter": "plain",
},
},
"loggers": {
"httpx": {"level": "WARNING"}, # quieten a chatty library
"myapp.db": {"level": "DEBUG"},
},
"root": {"level": "INFO", "handlers": ["stderr", "file"]},
})Loggers form a tree by dotted name, so logging.getLogger(__name__) in each module gives you per-package control for free. Records propagate up to the root’s handlers; set propagate: False on a logger only when it has its own handler and you want to stop duplicates. Under systemd or a container, log to stderr only and let the platform handle rotation. logging.captureWarnings(True) routes warnings.warn through logging. For a QueueHandler that keeps slow handlers off the request thread, 3.12 adds "respect_handler_level" and a queue key in dictConfig.
pytest fixtures and parametrize#
A fixture is a function whose return value is injected into any test that names it as a parameter; yield splits setup from teardown. Scope controls how often it runs (function default, module, session). Put shared fixtures in conftest.py at the level of the tests that use them, not one giant file at the root.
import pytest
@pytest.fixture
def tmp_config(tmp_path: Path) -> Path: # tmp_path is a built-in fixture
p = tmp_path / "app.yaml"
p.write_text("timeout: 5\n", encoding="utf-8")
return p
@pytest.fixture(scope="session")
def db_url() -> Iterator[str]:
container = start_postgres() # once for the whole run
yield container.url
container.stop() # teardown after the last test
@pytest.fixture(autouse=True)
def no_network(monkeypatch: pytest.MonkeyPatch) -> None: # applies to every test in scope
monkeypatch.delenv("HTTP_PROXY", raising=False)
@pytest.mark.parametrize(
("raw", "expected"),
[
("5s", 5.0),
("2m", 120.0),
pytest.param("", None, id="empty"),
pytest.param("bad", None, marks=pytest.mark.xfail(raises=ValueError, strict=True)),
],
)
def test_parse_duration(raw: str, expected: float | None) -> None:
assert parse_duration(raw) == expecteduv run pytest -k 'parse and not slow' # expression over test names and markers
uv run pytest --lf # only tests that failed last time
uv run pytest -x --pdb # drop into the debugger at the first failure
uv run pytest --fixtures tests/ # every fixture available, with docstrings
uv run pytest --setup-show tests/test_x.py # print fixture setup and teardown order
uv run pytest -p no:cacheprovider -q # no .pytest_cache, for read-only CI checkoutsmonkeypatch reverts environment, attributes and sys.path changes after each test, unlike a bare os.environ[...] = .... capsys captures stdout and stderr; caplog captures log records and lets you assert on caplog.records. Mark slow or network tests (@pytest.mark.slow) and register the marker under [tool.pytest.ini_options] markers so --strict-markers catches typos. See Testing for what to mock and what to leave real.
Useful one-liners#
# Serve the current directory on localhost only
python3 -m http.server 8000 --bind 127.0.0.1
# Pretty-print JSON
python3 -m json.tool < data.json
# Time a snippet over many runs
python3 -m timeit -s 'import json' 'json.dumps({"a":1})'
# Profile a script, slowest cumulative first
python3 -m cProfile -s cumtime script.py | head -25
# Where a module is loaded from and its version
python3 -c 'import httpx, inspect; print(httpx.__version__, inspect.getfile(httpx))'
# Base64 and URL encoding
python3 -c 'import base64,sys; print(base64.b64encode(sys.stdin.buffer.read()).decode())'
python3 -c 'import urllib.parse,sys; print(urllib.parse.quote(sys.argv[1]))' 'a b&c'
# Epoch to ISO 8601 in UTC (datetime.UTC needs 3.11+)
python3 -c 'import sys,datetime as d; print(d.datetime.fromtimestamp(int(sys.argv[1]), d.UTC).isoformat())' 17000000002023-11-14T22:13:20+00:00# Validate a YAML file (needs PyYAML)
uv run --with pyyaml python -c 'import sys,yaml; yaml.safe_load(open(sys.argv[1]))' config.yaml
# Which distribution provides an import name
python3 -c 'import importlib.metadata as m; print(m.packages_distributions()["yaml"])'
# Check a wheel's metadata before publishing
uvx twine check dist/*
# Every third-party import in the tree, to compare against pyproject dependencies
grep -rhoE '^(from|import) [a-zA-Z_]+' src | awk '{print $2}' | sort -u
# Run a module's tests with warnings turned into errors
uv run pytest -W error::DeprecationWarning tests/
# Dump every environment variable a Settings class would read
python3 -c 'from myapp.config import Settings; print(Settings.model_json_schema()["properties"].keys())'
# Byte-compile everything to catch syntax errors without importing
python3 -m compileall -q src/
# Show the resolved dependency tree from the lock file
uv tree --depth 1
# Start a REPL with the project importable and a variable pre-loaded
uv run python -i -c 'from myapp.config import Settings; s = Settings()'Snippets#
Retry with exponential backoff and jitter without a library, for the case where tenacity is one dependency too many.
import random, time
def retry(fn, *, attempts: int = 5, base: float = 0.5, cap: float = 10.0, retry_on=(OSError,)):
for i in range(attempts):
try:
return fn()
except retry_on as e:
if i == attempts - 1:
raise
delay = min(cap, base * 2**i) * random.uniform(0.5, 1.5)
log.warning("attempt %d failed (%s), retrying in %.1fs", i + 1, e, delay)
time.sleep(delay)HTTP client with separate connect and read timeouts, limits and a retrying transport.
client = httpx.Client(
base_url="https://api.example.com",
timeout=httpx.Timeout(connect=3.0, read=15.0, write=10.0, pool=5.0),
limits=httpx.Limits(max_connections=50, max_keepalive_connections=10),
transport=httpx.HTTPTransport(retries=2), # retries connection errors only, not HTTP status codes
headers={"authorization": f"Bearer {token}"},
)Graceful shutdown: finish the current item on SIGTERM, then exit.
import signal
stop = False
def _handle(signum, frame):
global stop
stop = True
log.info("signal %s received, finishing current item", signal.Signals(signum).name)
signal.signal(signal.SIGTERM, _handle)
signal.signal(signal.SIGINT, _handle)
for item in work:
if stop:
break
process(item)The same for asyncio, cancelling the main task so finally blocks run.
async def main() -> None:
loop = asyncio.get_running_loop()
task = asyncio.current_task()
for sig in (signal.SIGTERM, signal.SIGINT):
loop.add_signal_handler(sig, task.cancel)
try:
await serve()
except asyncio.CancelledError:
await drain() # bounded cleanup, then let the cancellation finish
raiseConfiguration layering: defaults, then a TOML file, then environment, with the environment winning.
import os, tomllib
DEFAULTS = {"timeout": 10.0, "workers": 4, "api_url": ""}
def load(path: Path) -> dict:
cfg = dict(DEFAULTS)
if path.is_file():
with path.open("rb") as f: # tomllib needs binary mode
cfg.update(tomllib.load(f))
for k in cfg:
if (v := os.getenv(f"APP_{k.upper()}")) is not None:
cfg[k] = type(DEFAULTS[k])(v) # coerce to the default's type
return cfgWorker pool with bounded queue and per-item error handling, using threads.
from concurrent.futures import ThreadPoolExecutor
def run_all(items, fn, workers: int = 8) -> tuple[int, int]:
ok = failed = 0
with ThreadPoolExecutor(max_workers=workers) as pool:
for item, result in zip(items, pool.map(lambda i: _safe(fn, i), items)):
if isinstance(result, Exception):
failed += 1
log.error("item %s failed: %s", item, result)
else:
ok += 1
return ok, failed
def _safe(fn, item):
try:
return fn(item)
except Exception as e: # captured, not raised, so the pool keeps going
return eContext manager that times a block and logs the duration once.
from contextlib import contextmanager
import time
@contextmanager
def timed(label: str):
start = time.perf_counter()
try:
yield
finally:
log.info("%s took %.3fs", label, time.perf_counter() - start)
with timed("reconcile"):
reconcile()Structured log context that follows the request through every log call, using contextvars.
import contextvars, logging
request_id: contextvars.ContextVar[str] = contextvars.ContextVar("request_id", default="-")
class ContextFilter(logging.Filter):
def filter(self, record: logging.LogRecord) -> bool:
record.request_id = request_id.get() # available to formatters as %(request_id)s
return True
logging.getLogger().addFilter(ContextFilter())
token = request_id.set(incoming_id) # per request; asyncio tasks inherit a copyStream a large download to disk without holding it in memory, and verify the checksum.
import hashlib
def download(client: httpx.Client, url: str, dest: Path, sha256: str) -> None:
h = hashlib.sha256()
tmp = dest.with_suffix(dest.suffix + ".part")
with client.stream("GET", url) as r, tmp.open("wb") as f:
r.raise_for_status()
for chunk in r.iter_bytes(1 << 20):
f.write(chunk)
h.update(chunk)
if h.hexdigest() != sha256:
tmp.unlink()
raise ValueError(f"checksum mismatch for {url}")
tmp.replace(dest)Paginate an API until the cursor runs out, as a generator.
def iter_pages(client: httpx.Client, path: str, **params):
cursor = None
while True:
r = client.get(path, params={**params, "cursor": cursor})
r.raise_for_status()
body = r.json()
yield from body["items"]
cursor = body.get("next_cursor")
if not cursor:
returnTable-driven test with a fake HTTP transport, no network and no mocking library.
@pytest.mark.parametrize(
("status", "body", "expected"),
[(200, {"items": [1, 2]}, [1, 2]), (404, {}, []), (500, {}, None)],
)
def test_fetch_items(status: int, body: dict, expected):
def handler(request: httpx.Request) -> httpx.Response:
return httpx.Response(status, json=body)
client = httpx.Client(transport=httpx.MockTransport(handler))
if expected is None:
with pytest.raises(httpx.HTTPStatusError):
fetch_items(client)
else:
assert fetch_items(client) == expectedRuff configuration that covers the rule families worth enforcing on operational code.
[tool.ruff]
line-length = 100
target-version = "py312"
src = ["src", "tests"]
[tool.ruff.lint]
select = ["E", "W", "F", "I", "B", "UP", "S", "PTH", "RUF", "SIM", "ASYNC", "LOG", "T20"]
ignore = ["S101"] # assert is fine in tests, and tests are where it appears
per-file-ignores = { "tests/**" = ["S", "T20"] }
[tool.ruff.lint.isort]
known-first-party = ["myapp"]
[tool.ruff.format]
docstring-code-format = trueArgument parser with subcommands, each dispatching to its own function.
def main(argv: list[str] | None = None) -> int:
ap = argparse.ArgumentParser(prog="ops")
sub = ap.add_subparsers(dest="cmd", required=True)
p = sub.add_parser("sync", help="reconcile inventory"); p.add_argument("path", type=Path); p.set_defaults(fn=cmd_sync)
p = sub.add_parser("report", help="print a summary"); p.add_argument("--json", action="store_true"); p.set_defaults(fn=cmd_report)
args = ap.parse_args(argv)
return args.fn(args)Troubleshooting#
| Symptom | Likely cause | Check or fix |
|---|---|---|
ModuleNotFoundError for an installed package | Different interpreter from the one you installed into | python3 -c 'import sys; print(sys.executable)'; use uv run or activate .venv |
error: externally-managed-environment from pip | Distro Python protects system packages (PEP 668) | Use a venv, uv tool install, or uvx; do not pass --break-system-packages |
uv sync --locked fails in CI | pyproject.toml changed without re-locking | Run uv lock locally and commit uv.lock |
| Script hangs on a network call | No timeout (requests, raw sockets) | Add timeout=; dump stacks with py-spy dump --pid <pid> |
Script hangs on subprocess.run | Child waits on stdin or never exits | Pass stdin=subprocess.DEVNULL and timeout= |
UnicodeDecodeError on one host only | Locale-dependent default encoding | Pass encoding="utf-8", or run with PYTHONUTF8=1 |
RuntimeError: This event loop is already running | asyncio.run inside Jupyter or another loop | await the coroutine directly |
coroutine ... was never awaited warning | Called an async def without await | Add await, or schedule it with a TaskGroup |
| Memory grows until the process is killed | Reading a whole file or response into memory | Iterate lines or use client.stream(); profile with tracemalloc |
| Threads give no speed-up | CPU-bound work held by the GIL | ProcessPoolExecutor, or a free-threaded build |
KeyError: "Attempt to overwrite 'name' in LogRecord" | extra key collides with a built-in attribute | Rename the key or nest under one key |
| Task exception was never retrieved | A fire-and-forget create_task raised and nothing awaited it | Keep a reference and await it, or use TaskGroup |
| asyncio program is slow and one call dominates | A blocking call inside a coroutine | PYTHONASYNCIODEBUG=1 logs slow callbacks; wrap in asyncio.to_thread |
mypy: Incompatible types on a dict from JSON | Untyped dict[str, Any] flowing into typed code | Validate with pydantic or a TypedDict at the boundary |
mypy: Cannot find implementation or library stub | Package ships no types | --install-types, or ignore_missing_imports in an override for that module |
pydantic accepts "3" for an int field | Lax mode coerces by default | model_config = ConfigDict(strict=True) or Field(strict=True) |
ValueError: mutable default ... use default_factory | List or dict as a dataclass default | field(default_factory=list) |
| Fixture not found | conftest.py is in a sibling directory, or a plugin is disabled | pytest --fixtures; move the fixture up the tree |
| Same test fails only when run with others | Shared module-level state or fixture scope too wide | pytest -p randomly or --lf; narrow the scope to function |
uv picks the wrong interpreter | No .python-version and several installed Pythons | uv python pin 3.13; check uv python find |
| Slow container start-up | Bytecode compiled on every launch | UV_COMPILE_BYTECODE=1 at build time |
For a live process that is stuck or slow, py-spy dump --pid <pid> prints every thread’s stack and py-spy top --pid <pid> samples hot functions, without restarting it. Both need permission to ptrace the process (root, or the same user with kernel.yama.ptrace_scope=0). python3 -X faulthandler script.py prints tracebacks on a crash or on SIGABRT.