Testing
Choose the right test layer, write fast deterministic tests, use fakes and fixtures, and run, isolate and de-flake tests in pytest, Go and Vitest.
On this page
Cheatsheet#
| Task | Command |
|---|---|
| pytest: stop at first failure, quiet | pytest -x -q |
| pytest: one test | pytest tests/test_api.py::test_create -v |
| pytest: tests matching a name | pytest -k 'expired and not slow' |
| pytest: rerun last failures only | pytest --lf |
| pytest: slowest tests | pytest --durations=10 |
| pytest: drop into the debugger on failure | pytest -x --pdb |
| pytest: run in parallel | pytest -n auto (pytest-xdist) |
| pytest: coverage with missing lines | pytest --cov=src --cov-report=term-missing (pytest-cov) |
| Go: all packages with the race detector | go test -race ./... |
| Go: one test, verbose | go test -run '^TestParse$' -v ./pkg |
| Go: bypass the test cache | go test -count=1 ./... |
| Go: repeat to expose flakes | go test -count=50 -run TestX ./pkg |
| Go: coverage summary | go test -coverprofile=c.out ./... && go tool cover -func=c.out |
| Go: fuzz one target for a minute | go test -fuzz=FuzzParse -fuzztime=60s ./pkg |
| Go: benchmarks | go test -bench=. -benchmem -run='^$' ./pkg |
| Vitest: one file, no watch | npx vitest run src/api.test.ts |
| Vitest: watch mode | npx vitest |
| Vitest: coverage | npx vitest run --coverage (needs @vitest/coverage-v8) |
| Real dependencies for integration tests | Testcontainers, or docker compose up -d --wait |
| Load test | k6 run script.js |
What each test layer is for#
| Layer | Answers | Typical cost | How many |
|---|---|---|---|
| Unit | Does this function behave for these inputs | Milliseconds | Many; no I/O |
| Integration | Do these components work together against real dependencies | Seconds | Enough to cover each seam (database, queue, HTTP client) |
| Contract | Does the provider still give consumers what they rely on | Seconds | One per consumer and provider pair |
| End-to-end | Does the critical user path work in a deployed system | Minutes, prone to flakes | A handful, on the paths that make money or lose data |
| Load | Does it hold latency and error targets under expected traffic | Minutes to hours | Before capacity or architecture decisions |
The test pyramid is about feedback speed. A test that takes ten minutes to report something a unit test could report in ten milliseconds is worse, even when it is more realistic. Push each check down to the cheapest layer that can catch the failure, and keep the slow layers for failures only they can see: wiring, configuration, real network and database behaviour.
Test behaviour through the public interface. Tests coupled to private functions or call order fail on every refactor and still miss real regressions. See Design for shaping code so the interface is testable.
What makes a good test#
def test_rejects_expired_token():
# arrange
token = make_token(expires_at=datetime(2020, 1, 1, tzinfo=UTC))
# act
result = verify(token, now=datetime(2024, 1, 1, tzinfo=UTC))
# assert
assert result == Err(TokenExpired)| Property | What it means in practice |
|---|---|
| Fast | Milliseconds; no network, no sleep, no real clock |
| Isolated | Passes in any order and in parallel; no shared mutable state |
| Deterministic | Time, randomness and IDs are injected, not read from the system |
| Self-checking | Asserts the outcome; never relies on someone reading output |
| Clear on failure | The failure message shows the input, the expected value and the actual value |
Name tests after the behaviour (test_rejects_expired_token, TestParse/rejects_empty_input), not the function (test_verify_2). The name is often the only context in a CI log.
In Go, table-driven tests with subtests give each case a name and let -run select one:
func TestParseDuration(t *testing.T) {
tests := []struct {
name string
in string
want time.Duration
}{
{"seconds", "1s", time.Second},
{"compound", "1h30m", 90 * time.Minute},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
got, err := ParseDuration(tt.in)
if err != nil {
t.Fatalf("ParseDuration(%q) error: %v", tt.in, err)
}
if got != tt.want {
t.Errorf("ParseDuration(%q) = %v, want %v", tt.in, got, tt.want)
}
})
}
}go test -run 'TestParseDuration/compound' ./pkg # one subtestSince Go 1.22 each loop iteration has its own tt, so capturing it in a parallel subtest is safe. See Go for the rest of the toolchain.
Test doubles#
| Double | What it does | Use when |
|---|---|---|
| Stub | Returns canned data | The test needs a dependency to answer but does not care how |
| Fake | Working lightweight implementation, such as an in-memory store | Many tests need realistic behaviour from a dependency |
| Mock | Asserts that specific calls happened | The call itself is the behaviour, such as sending an email |
| Spy | Records calls for inspection after the fact | As a mock, with the assertion written in the test |
Mock at a boundary you own, such as your UserStore interface, not the database driver three layers down. Mocking a third-party client’s internals produces tests that pass while the real integration is broken.
def test_sends_notification(monkeypatch):
sent = []
monkeypatch.setattr(notifier, "send", lambda msg: sent.append(msg))
deploy(version="1.4.2")
assert sent == ["deployed 1.4.2"]A fake of your own interface usually beats a mock. Every test that uses it exercises the same contract, and it breaks loudly when the contract changes. Run the fake and the real implementation through one shared test suite to keep them in agreement.
Fixtures and test data#
A pytest fixture provides setup and teardown to any test that names it as a parameter. The code after yield runs as teardown even if the test fails.
@pytest.fixture
def db(postgres_container): # depends on a session-scoped container
conn = connect(postgres_container.get_connection_url())
tx = conn.begin() # each test runs in a transaction
yield conn
tx.rollback() # nothing persists between tests
conn.close()Transaction rollback isolates tests cheaply, but it does not work when the code under test commits itself or uses a second connection. Use a fresh schema or database per test (or per xdist worker) in that case.
@pytest.mark.parametrize(
("raw", "expected"),
[("1s", 1), ("2m", 120), ("1h30m", 5400)],
ids=["seconds", "minutes", "compound"],
)
def test_parse_duration(raw, expected):
assert parse_duration(raw) == expectedBuild test data with a factory that takes overrides, so each test states only the fields it cares about:
def make_user(**over):
return User(**{"id": 1, "email": "alice@example.com", "active": True, **over})| pytest fixture scope | Created once per | Use for |
|---|---|---|
function (default) | Test | Anything mutable |
module | Test file | Read-only data shared by a file |
session | Test run (per xdist worker) | Containers, compiled assets, expensive clients |
Go equivalents: t.Cleanup registers teardown, t.TempDir() gives a directory removed after the test, t.Setenv restores the variable afterwards (not allowed in parallel tests), and t.Context() (Go 1.24+) is cancelled before cleanup runs.
Controlling time#
Tests that read the real clock or sleep are slow and flaky. Pass time in as a value or a clock interface, or use a tool that fakes it:
| Language | Approach |
|---|---|
| Python | Pass now as a parameter; time-machine or freezegun to patch datetime |
| Go | testing/synctest (Go 1.25+): inside synctest.Test, time uses a fake clock that advances only when every goroutine in the bubble is blocked |
| Vitest | vi.useFakeTimers(), vi.setSystemTime(date), vi.advanceTimersByTime(ms) |
func TestTimeout(t *testing.T) {
synctest.Test(t, func(t *testing.T) {
ctx, cancel := context.WithTimeout(t.Context(), 5*time.Second)
defer cancel()
<-ctx.Done() // returns immediately in real time; fake clock jumps 5 s
if ctx.Err() != context.DeadlineExceeded {
t.Fatal(ctx.Err())
}
})
}Integration tests#
Run the real dependency rather than an imitation of it. Containers make that cheap and repeatable. Testcontainers starts a container per session and removes it afterwards; it needs a Docker-compatible socket (Docker, or Podman with its socket enabled).
from testcontainers.postgres import PostgresContainer
@pytest.fixture(scope="session")
def postgres_container():
with PostgresContainer("postgres:17") as pg:
run_migrations(pg.get_connection_url())
yield pgWith Docker Compose, --wait blocks until services with health checks report healthy, which replaces sleep 30:
docker compose up -d --wait # start and wait for health checks
pytest tests/integration
docker compose down -v # stop and delete volumes (destroys test data)Test an HTTP API through the application object rather than a live socket where the framework allows it. The request goes through the same routing and middleware without a port to allocate:
from fastapi.testclient import TestClient
client = TestClient(app)
r = client.post("/items", json={"name": "widget"})
assert r.status_code == 201
assert r.json()["id"]The Go equivalent is httptest.NewRecorder() with the handler’s ServeHTTP, or httptest.NewServer(handler) when a real client is involved.
Contract tests#
The consumer’s tests record the requests it makes and the responses it relies on as a contract (a pact). The provider’s build replays those interactions against the real provider and fails if any response no longer matches. The break shows up in the provider’s CI instead of in staging.
# Consumer CI: publish pacts generated by the consumer's tests
pact-broker publish ./pacts \
--consumer-app-version "$(git rev-parse --short HEAD)" \
--branch "$(git branch --show-current)" \
--broker-base-url "$PACT_BROKER_URL" --broker-token "$PACT_BROKER_TOKEN"Provider verification runs inside the provider’s test suite using the Pact library for its language. See the Pact docs for the verifier setup.
A cheaper approximation is a schema check: generate the OpenAPI or protobuf schema in CI and fail on incompatible changes (for protobuf, buf breaking).
End-to-end tests#
test('user can check out', async ({ page }) => {
await page.goto('/cart');
await page.getByRole('button', { name: 'Checkout' }).click();
await expect(page.getByText('Order confirmed')).toBeVisible(); // retries until visible or timeout
});This is Playwright. Select elements by role and accessible name, not CSS classes, so tests survive styling changes and check accessibility as a side effect. Never use fixed sleeps: assert on the condition and let the framework retry. A flaky end-to-end test is usually a race between the test and the application, and often a race users hit too.
Coverage and mutation testing#
Coverage shows which lines ran, not whether any assertion checked their result. It is useful as a floor (uncovered code is definitely untested) and misleading as a target (100% coverage with weak assertions proves little). Branch coverage (--cov-branch in pytest-cov) catches untested if arms that line coverage marks as covered.
pytest --cov=src --cov-branch --cov-report=term-missing --cov-fail-under=80
go test -coverprofile=c.out ./... && go tool cover -func=c.out | tail -1
go tool cover -html=c.out # open annotated source in a browser
npx vitest run --coverage
mutmut run # Python mutation testingMutation testing changes the code (flips a comparison, removes a call) and reruns the tests. A mutant that survives means no test noticed the change. It is slow, so run it on the modules that matter rather than on every push.
Flakes#
A flaky test passes and fails on the same code. Find the cause before deciding what to do with it.
| Cause | Fix |
|---|---|
| Real clock or timezone | Inject time; fake timers; set TZ=UTC in CI |
| State shared between tests | Fresh fixtures, or a transaction rolled back per test |
| Test order dependence | Randomise order to expose it (pytest-randomly, go test -shuffle=on), then remove the shared state |
| Fixed sleeps | Poll the condition with a deadline |
| Unseeded randomness | Seed it and print the seed on failure |
| Parallel tests using one resource | Namespace per test: schema, key prefix, temp directory, port 0 |
| Data race | Run with go test -race; fix the unsynchronised access |
| External service | Replace with a fake or container; keep live calls out of unit tests |
pytest --count=20 -x tests/test_flaky.py # repeat each test 20 times (pytest-repeat)
pytest -p no:randomly # pytest-randomly shuffles by default; this turns it off
pytest --randomly-seed=1234 # replay the order from a failing run
go test -count=100 -race -run TestFlaky ./pkg # repeat with the race detector
go test -shuffle=on ./pkg # random order; prints the seed to reuse with -shuffle=<seed>pytest-randomly prints the seed it used at the start of each run; pass it back with --randomly-seed to reproduce the order.
Quarantining a flaky test often hides a real defect. Fix it, or delete it and cover the behaviour another way. A test nobody trusts costs CI time and teaches people to ignore red builds.
Shaping the pyramid#
The pyramid is a ratio, not a rule: many unit tests, fewer integration tests, a handful of end-to-end tests. Two inverted shapes show up in practice. The ice-cream cone has most checks in slow end-to-end suites because the code was never structured for unit testing, so every build takes 40 minutes and a failure points at nothing specific. The hourglass has a fat unit layer and a fat end-to-end layer with nothing in between, so the wiring between components (the SQL, the serialisation, the queue client) is exercised only by the slowest tests.
Push each check to the lowest layer that can see the failure: parsing and business rules to unit tests; the query, the migration and the HTTP contract to integration tests against a real container; the checkout flow to one end-to-end test. When an end-to-end test fails, write the unit or integration test that would have caught it, then decide whether the end-to-end case still earns its minutes. The unit layer should also cover error paths, which are where production incidents live and where end-to-end tests rarely go.
Code shape decides what is testable. A function that reads the clock, opens a connection and formats output in one body can only be tested end-to-end. Split it so the pure logic takes values and returns values, and a thin outer layer does the I/O; the pure part gets fast, exhaustive tests and the outer part gets one integration test. This is the same seam that makes a fake possible (see Test doubles).
Property-based testing#
Example-based tests check the cases you thought of. A property-based test states an invariant (round-trips, ordering, idempotency, equivalence to a slow reference) and the framework generates hundreds of inputs, then shrinks a failing input to the smallest one that still fails. It finds the empty string, the negative zero, the unicode combining character and the list of one element.
from hypothesis import given, strategies as st, settings
@given(st.text())
def test_encode_decode_roundtrip(s: str) -> None:
assert decode(encode(s)) == s
@given(st.lists(st.integers()))
def test_sort_is_idempotent_and_ordered(xs: list[int]) -> None:
once = my_sort(xs)
assert once == my_sort(once)
assert all(a <= b for a, b in zip(once, once[1:]))
assert sorted(xs) == once # equivalence to a trusted implementation
@settings(max_examples=500, deadline=None)
@given(st.builds(Order, quantity=st.integers(min_value=1, max_value=10_000)))
def test_total_never_negative(order: Order) -> None:
assert order.total() >= 0Hypothesis saves failing examples in .hypothesis/ and replays them first on the next run; commit that directory or use @example(...) to pin a regression. Go’s testing/quick is deprecated; use rapid (pgregory.net/rapid) for properties, or the built-in fuzzer for anything that parses bytes:
func FuzzParse(f *testing.F) {
f.Add("a=1") // seed corpus
f.Fuzz(func(t *testing.T, in string) {
cfg, err := Parse(in)
if err != nil { return } // invalid input is fine; a panic is not
if out := cfg.String(); out != in {
if _, err := Parse(out); err != nil { t.Fatalf("re-parse of %q failed: %v", out, err) }
}
})
}go test -fuzz=FuzzParse -fuzztime=60s ./pkg mutates inputs guided by coverage; failing inputs land in testdata/fuzz/FuzzParse/ and run as ordinary regression tests from then on. In TypeScript, fast-check provides the same model (fc.assert(fc.property(fc.string(), s => decode(encode(s)) === s))). Properties that pay off most: round-trip (serialise then parse), invariants after any operation sequence (a state machine test), and “the fast version equals the obvious slow version”.
Mocking guidance#
Mocks are a design tool as much as a testing tool: the need to mock something deep is a sign the code has no seam there. Rules that keep mocks from turning a suite into a mirror of the implementation:
| Rule | Reason |
|---|---|
| Mock only what you own | A mock of a third-party client encodes your guess about its behaviour; the real one differs on the edge case that matters |
| Prefer a fake over a mock | A fake is one implementation shared by every test; a mock is re-specified in each test and drifts |
| Never mock the unit under test | Partially mocked objects test the mock’s wiring, not the code |
| Assert outcomes, not calls, unless the call is the outcome | assert sent == [...] survives refactors; assert_called_once_with(...) breaks on every signature change |
| One mock per test, at the boundary | Three mocks means the test has become a re-statement of the code’s control flow |
| Verify the fake against the real thing | A shared contract test suite that runs against both keeps the fake honest |
# Fake: an in-memory implementation of the interface the code depends on
class FakeUserStore:
def __init__(self) -> None:
self.users: dict[int, User] = {}
def get(self, user_id: int) -> User | None:
return self.users.get(user_id)
def save(self, user: User) -> None:
self.users[user.id] = user
def test_deactivate_marks_user_inactive() -> None:
store = FakeUserStore()
store.save(make_user(id=7, active=True))
deactivate(store, user_id=7)
assert store.get(7).active is False # outcome, not "save was called with ..."# Contract suite: the same tests run against the fake and the real store
@pytest.fixture(params=["fake", "postgres"])
def store(request, postgres_container):
if request.param == "fake":
return FakeUserStore()
return PostgresUserStore(postgres_container.get_connection_url())
def test_get_returns_saved_user(store):
store.save(make_user(id=1))
assert store.get(1) == make_user(id=1)For HTTP, prefer a transport-level fake (httpx.MockTransport, Go’s httptest.NewServer, msw in Node) over patching the client library’s methods, because it exercises your serialisation and error handling. For time and randomness, inject a clock and a seed. If unittest.mock.patch needs a dotted path four modules deep, the code is telling you where the seam should be.
Contract tests in practice#
A consumer-driven contract test has three parts: the consumer writes a test against a mock provider that records the interaction; the resulting pact file is published; the provider’s CI replays it against the real provider and fails on mismatch. The consumer describes only the fields it reads, so the provider is free to add anything else.
# Consumer side (pact-python v3 API)
from pact import Pact
def test_get_user(pact: Pact) -> None:
(pact.upon_receiving("a request for user 7")
.given("user 7 exists")
.with_request("GET", "/users/7")
.will_respond_with(200)
.with_body({"id": 7, "email": "alice@example.com"}, content_type="application/json"))
with pact.serve() as srv:
user = UserClient(srv.url).get(7) # the real client code, against the mock provider
assert user.email == "alice@example.com"# Provider CI: verify every published pact for this provider, then record the result
pact-verifier --provider-base-url http://localhost:8080 \
--broker-url "$PACT_BROKER_URL" --broker-token "$PACT_BROKER_TOKEN" \
--provider my-api --provider-version "$(git rev-parse --short HEAD)" --publish-results
pact-broker can-i-deploy --pacticipant my-api --version "$(git rev-parse --short HEAD)" --to-environment productioncan-i-deploy answers the deployment question directly: has this provider version been verified against every consumer version currently in production. Provider states (given("user 7 exists")) map to setup hooks in the provider’s verification harness that seed the data. Contract tests replace the shared staging environment for API compatibility; they do not replace an integration test of the provider against its own database. For gRPC and event schemas, buf breaking --against '.git#branch=main' and schema-registry compatibility checks (Avro, Protobuf, JSON Schema) do the same job with less machinery.
Snapshot and golden tests#
A golden test compares output against a file checked in beside the test. It suits rendered templates, CLI output, generated code and serialised structures, where writing the expected value by hand is tedious and reviewing a diff is easy. The trap is accepting a changed snapshot without reading it, which turns the test into “the output is whatever it was last time”.
var update = flag.Bool("update", false, "rewrite golden files")
func TestRenderManifest(t *testing.T) {
got := RenderManifest(fixtureInput)
path := filepath.Join("testdata", "manifest.golden.yaml")
if *update {
if err := os.WriteFile(path, got, 0o644); err != nil { t.Fatal(err) }
}
want, err := os.ReadFile(path)
if err != nil { t.Fatal(err) }
if diff := cmp.Diff(string(want), string(got)); diff != "" {
t.Errorf("manifest differs (-want +got):\n%s\nrun: go test -update ./...", diff)
}
}pytest has syrupy (assert result == snapshot) and Vitest has toMatchSnapshot() and toMatchInlineSnapshot(), with -u to accept changes. Keep snapshots small and deterministic (no timestamps, sorted keys, stable IDs), and review snapshot diffs in pull requests as carefully as code.
CI gating#
A gate is a check that blocks a merge. Gates should be fast enough that people wait for them, strict enough that green means deployable, and few enough that each one is trusted. Put the fast, deterministic checks on every push and the slow ones on the merge queue or a nightly schedule, and make each gate report the failing test by name.
| Stage | Runs on | Gate | Budget |
|---|---|---|---|
| Lint, format, type-check | Every push | Required | Under 2 minutes |
| Unit tests | Every push | Required | Under 5 minutes |
| Integration tests (containers) | Every push or pull request | Required | Under 15 minutes |
| Contract verification | Pull request | Required for providers | Minutes |
| End-to-end | Merge queue, or nightly | Required on the queue; advisory nightly | Under 30 minutes |
| Load, mutation, fuzz | Nightly or weekly | Advisory, with an owner who reads the results | Hours |
# GitHub Actions: split the suite across runners and fail fast
jobs:
test:
strategy:
fail-fast: true
matrix: { shard: [1, 2, 3, 4] }
steps:
- run: pytest -q --splits 4 --group ${{ matrix.shard }} --durations-path .test_durations # pytest-split
- run: go test -race -count=1 -json ./... | tee test.json | go tool test2json >/dev/null # machine-readableBranch protection (or a merge queue) marks the jobs above as required checks so the merge button waits on them. Coverage gates (--cov-fail-under) should ratchet, not block: fail when coverage drops below the current value minus a small tolerance, so the number only moves up. Automatic retries of failed jobs hide flakes; if you must retry, retry only known-flaky tests through a plugin (pytest-rerunfailures, vitest retry) and track the rerun count as a metric with an owner, so the list shrinks. Cache dependency downloads, not test results; uv cache, the Go module and build caches, and ~/.npm are safe to restore, but a restored test cache can mask a change in an external input. See GitHub Actions for the workflow mechanics.
Running tests in CI#
Run the fast layers on every push and the slow layers before merge. Stop at the first failing layer, and make the log name the failing test and assertion so nobody needs to reproduce it locally to start.
ruff check . && mypy --strict src/ && pytest -q --cov=src --cov-fail-under=80
go vet ./... && go test -race -count=1 ./...
npm run lint && npx tsc --noEmit && npx vitest run --coverageGo caches successful test results and reuses them when the package and its inputs have not changed, printing (cached). Go reruns a cached test when files it opened inside the module or environment variables it read have changed, but it cannot see files outside the module, databases or network services. -count=1 disables the cache. go clean -testcache clears the cache.
Troubleshooting test runs#
| Symptom | Likely cause | Check or fix |
|---|---|---|
pytest fixture 'x' not found | Fixture defined outside a conftest.py the test can see | Move it to conftest.py in the test’s directory or a parent; pytest --fixtures lists what is visible |
pytest ModuleNotFoundError for your package | Package not installed in the environment | uv sync (installs the project), or set pythonpath = ["src"] under [tool.pytest.ini_options] |
pytest PytestUnknownMarkWarning or error with --strict-markers | Marker not registered | Add it under markers in the pytest config |
| Passes alone, fails in the full run | Shared state or order dependence | pytest -p no:randomly, bisect the run, check module-level globals and fixture scopes |
| Passes locally, fails in CI | Timezone, locale, CPU count, missing service, cached result | Compare environments; go test -count=1; set TZ=UTC locally |
Go test shows (cached) after an external change | Go only tracks files inside the module and environment variables | -count=1 |
cannot use -fuzz flag with multiple packages | -fuzz needs exactly one package and one matching target | Name one package, e.g. ./pkg, and an anchored regex such as '^FuzzParse$' |
| Fuzz failure in CI but not locally | Failing input saved under testdata/fuzz/FuzzXxx/ | Commit that file; it becomes a regression case run by plain go test |
| Testcontainers cannot connect to Docker | No Docker socket (for example with Podman) | Enable the Podman socket and set DOCKER_HOST to it |
Vitest --coverage asks to install a package | Coverage provider not installed | npm install -D @vitest/coverage-v8 |
Hypothesis Flaky or Unsatisfiable health check | Test depends on state between examples, or the strategy filters too hard | Make the test pure; replace .filter with a constructive strategy |
| Hypothesis finds a failure CI cannot reproduce | Example database not shared | Commit .hypothesis/ or add the input with @example(...) |
Pact verification fails with state not found | Provider state handler missing for a given(...) | Implement the state setup hook in the provider harness |
| Snapshot suite passes after a behaviour change | Snapshots were bulk-updated with -u | Review snapshot diffs in the PR; never update without reading |
| Mock asserts pass but production breaks | Mocked a third-party client instead of your own boundary | Replace with a transport-level fake or a container |
| Coverage gate fails on an unrelated PR | Global threshold above the current value | Ratchet from the current value; measure changed files only |
| CI job green but tests did not run | Test discovery found nothing (wrong path, missing __init__.py, glob mismatch) | Fail on zero tests: pytest --co -q | grep -c ::, vitest --passWithNoTests off |
| Sharded CI runs unbalanced | Durations file stale | Regenerate .test_durations on a schedule |