Software Engineering WikiSE Wiki

Testing

Choose the right test layer, write fast deterministic tests, use fakes and fixtures, and run, isolate and de-flake tests in pytest, Go and Vitest.

Reviewed MarkdownEdit

On this page

Cheatsheet#

TaskCommand
pytest: stop at first failure, quietpytest -x -q
pytest: one testpytest tests/test_api.py::test_create -v
pytest: tests matching a namepytest -k 'expired and not slow'
pytest: rerun last failures onlypytest --lf
pytest: slowest testspytest --durations=10
pytest: drop into the debugger on failurepytest -x --pdb
pytest: run in parallelpytest -n auto (pytest-xdist)
pytest: coverage with missing linespytest --cov=src --cov-report=term-missing (pytest-cov)
Go: all packages with the race detectorgo test -race ./...
Go: one test, verbosego test -run '^TestParse$' -v ./pkg
Go: bypass the test cachego test -count=1 ./...
Go: repeat to expose flakesgo test -count=50 -run TestX ./pkg
Go: coverage summarygo test -coverprofile=c.out ./... && go tool cover -func=c.out
Go: fuzz one target for a minutego test -fuzz=FuzzParse -fuzztime=60s ./pkg
Go: benchmarksgo test -bench=. -benchmem -run='^$' ./pkg
Vitest: one file, no watchnpx vitest run src/api.test.ts
Vitest: watch modenpx vitest
Vitest: coveragenpx vitest run --coverage (needs @vitest/coverage-v8)
Real dependencies for integration testsTestcontainers, or docker compose up -d --wait
Load testk6 run script.js

What each test layer is for#

LayerAnswersTypical costHow many
UnitDoes this function behave for these inputsMillisecondsMany; no I/O
IntegrationDo these components work together against real dependenciesSecondsEnough to cover each seam (database, queue, HTTP client)
ContractDoes the provider still give consumers what they rely onSecondsOne per consumer and provider pair
End-to-endDoes the critical user path work in a deployed systemMinutes, prone to flakesA handful, on the paths that make money or lose data
LoadDoes it hold latency and error targets under expected trafficMinutes to hoursBefore capacity or architecture decisions

The test pyramid is about feedback speed. A test that takes ten minutes to report something a unit test could report in ten milliseconds is worse, even when it is more realistic. Push each check down to the cheapest layer that can catch the failure, and keep the slow layers for failures only they can see: wiring, configuration, real network and database behaviour.

Test behaviour through the public interface. Tests coupled to private functions or call order fail on every refactor and still miss real regressions. See Design for shaping code so the interface is testable.

What makes a good test#

def test_rejects_expired_token():
    # arrange
    token = make_token(expires_at=datetime(2020, 1, 1, tzinfo=UTC))

    # act
    result = verify(token, now=datetime(2024, 1, 1, tzinfo=UTC))

    # assert
    assert result == Err(TokenExpired)
PropertyWhat it means in practice
FastMilliseconds; no network, no sleep, no real clock
IsolatedPasses in any order and in parallel; no shared mutable state
DeterministicTime, randomness and IDs are injected, not read from the system
Self-checkingAsserts the outcome; never relies on someone reading output
Clear on failureThe failure message shows the input, the expected value and the actual value

Name tests after the behaviour (test_rejects_expired_token, TestParse/rejects_empty_input), not the function (test_verify_2). The name is often the only context in a CI log.

In Go, table-driven tests with subtests give each case a name and let -run select one:

func TestParseDuration(t *testing.T) {
	tests := []struct {
		name string
		in   string
		want time.Duration
	}{
		{"seconds", "1s", time.Second},
		{"compound", "1h30m", 90 * time.Minute},
	}
	for _, tt := range tests {
		t.Run(tt.name, func(t *testing.T) {
			t.Parallel()
			got, err := ParseDuration(tt.in)
			if err != nil {
				t.Fatalf("ParseDuration(%q) error: %v", tt.in, err)
			}
			if got != tt.want {
				t.Errorf("ParseDuration(%q) = %v, want %v", tt.in, got, tt.want)
			}
		})
	}
}
go test -run 'TestParseDuration/compound' ./pkg     # one subtest

Since Go 1.22 each loop iteration has its own tt, so capturing it in a parallel subtest is safe. See Go for the rest of the toolchain.

Test doubles#

DoubleWhat it doesUse when
StubReturns canned dataThe test needs a dependency to answer but does not care how
FakeWorking lightweight implementation, such as an in-memory storeMany tests need realistic behaviour from a dependency
MockAsserts that specific calls happenedThe call itself is the behaviour, such as sending an email
SpyRecords calls for inspection after the factAs a mock, with the assertion written in the test

Mock at a boundary you own, such as your UserStore interface, not the database driver three layers down. Mocking a third-party client’s internals produces tests that pass while the real integration is broken.

def test_sends_notification(monkeypatch):
    sent = []
    monkeypatch.setattr(notifier, "send", lambda msg: sent.append(msg))

    deploy(version="1.4.2")

    assert sent == ["deployed 1.4.2"]

A fake of your own interface usually beats a mock. Every test that uses it exercises the same contract, and it breaks loudly when the contract changes. Run the fake and the real implementation through one shared test suite to keep them in agreement.

Fixtures and test data#

A pytest fixture provides setup and teardown to any test that names it as a parameter. The code after yield runs as teardown even if the test fails.

@pytest.fixture
def db(postgres_container):                       # depends on a session-scoped container
    conn = connect(postgres_container.get_connection_url())
    tx = conn.begin()                             # each test runs in a transaction
    yield conn
    tx.rollback()                                 # nothing persists between tests
    conn.close()

Transaction rollback isolates tests cheaply, but it does not work when the code under test commits itself or uses a second connection. Use a fresh schema or database per test (or per xdist worker) in that case.

@pytest.mark.parametrize(
    ("raw", "expected"),
    [("1s", 1), ("2m", 120), ("1h30m", 5400)],
    ids=["seconds", "minutes", "compound"],
)
def test_parse_duration(raw, expected):
    assert parse_duration(raw) == expected

Build test data with a factory that takes overrides, so each test states only the fields it cares about:

def make_user(**over):
    return User(**{"id": 1, "email": "alice@example.com", "active": True, **over})
pytest fixture scopeCreated once perUse for
function (default)TestAnything mutable
moduleTest fileRead-only data shared by a file
sessionTest run (per xdist worker)Containers, compiled assets, expensive clients

Go equivalents: t.Cleanup registers teardown, t.TempDir() gives a directory removed after the test, t.Setenv restores the variable afterwards (not allowed in parallel tests), and t.Context() (Go 1.24+) is cancelled before cleanup runs.

Controlling time#

Tests that read the real clock or sleep are slow and flaky. Pass time in as a value or a clock interface, or use a tool that fakes it:

LanguageApproach
PythonPass now as a parameter; time-machine or freezegun to patch datetime
Gotesting/synctest (Go 1.25+): inside synctest.Test, time uses a fake clock that advances only when every goroutine in the bubble is blocked
Vitestvi.useFakeTimers(), vi.setSystemTime(date), vi.advanceTimersByTime(ms)
func TestTimeout(t *testing.T) {
	synctest.Test(t, func(t *testing.T) {
		ctx, cancel := context.WithTimeout(t.Context(), 5*time.Second)
		defer cancel()
		<-ctx.Done()                  // returns immediately in real time; fake clock jumps 5 s
		if ctx.Err() != context.DeadlineExceeded {
			t.Fatal(ctx.Err())
		}
	})
}

Integration tests#

Run the real dependency rather than an imitation of it. Containers make that cheap and repeatable. Testcontainers starts a container per session and removes it afterwards; it needs a Docker-compatible socket (Docker, or Podman with its socket enabled).

from testcontainers.postgres import PostgresContainer

@pytest.fixture(scope="session")
def postgres_container():
    with PostgresContainer("postgres:17") as pg:
        run_migrations(pg.get_connection_url())
        yield pg

With Docker Compose, --wait blocks until services with health checks report healthy, which replaces sleep 30:

docker compose up -d --wait        # start and wait for health checks
pytest tests/integration
docker compose down -v             # stop and delete volumes (destroys test data)

Test an HTTP API through the application object rather than a live socket where the framework allows it. The request goes through the same routing and middleware without a port to allocate:

from fastapi.testclient import TestClient

client = TestClient(app)
r = client.post("/items", json={"name": "widget"})
assert r.status_code == 201
assert r.json()["id"]

The Go equivalent is httptest.NewRecorder() with the handler’s ServeHTTP, or httptest.NewServer(handler) when a real client is involved.

Contract tests#

The consumer’s tests record the requests it makes and the responses it relies on as a contract (a pact). The provider’s build replays those interactions against the real provider and fails if any response no longer matches. The break shows up in the provider’s CI instead of in staging.

# Consumer CI: publish pacts generated by the consumer's tests
pact-broker publish ./pacts \
  --consumer-app-version "$(git rev-parse --short HEAD)" \
  --branch "$(git branch --show-current)" \
  --broker-base-url "$PACT_BROKER_URL" --broker-token "$PACT_BROKER_TOKEN"

Provider verification runs inside the provider’s test suite using the Pact library for its language. See the Pact docs for the verifier setup.

A cheaper approximation is a schema check: generate the OpenAPI or protobuf schema in CI and fail on incompatible changes (for protobuf, buf breaking).

End-to-end tests#

test('user can check out', async ({ page }) => {
  await page.goto('/cart');
  await page.getByRole('button', { name: 'Checkout' }).click();
  await expect(page.getByText('Order confirmed')).toBeVisible();   // retries until visible or timeout
});

This is Playwright. Select elements by role and accessible name, not CSS classes, so tests survive styling changes and check accessibility as a side effect. Never use fixed sleeps: assert on the condition and let the framework retry. A flaky end-to-end test is usually a race between the test and the application, and often a race users hit too.

Coverage and mutation testing#

Coverage shows which lines ran, not whether any assertion checked their result. It is useful as a floor (uncovered code is definitely untested) and misleading as a target (100% coverage with weak assertions proves little). Branch coverage (--cov-branch in pytest-cov) catches untested if arms that line coverage marks as covered.

pytest --cov=src --cov-branch --cov-report=term-missing --cov-fail-under=80
go test -coverprofile=c.out ./... && go tool cover -func=c.out | tail -1
go tool cover -html=c.out                 # open annotated source in a browser
npx vitest run --coverage
mutmut run                                # Python mutation testing

Mutation testing changes the code (flips a comparison, removes a call) and reruns the tests. A mutant that survives means no test noticed the change. It is slow, so run it on the modules that matter rather than on every push.

Flakes#

A flaky test passes and fails on the same code. Find the cause before deciding what to do with it.

CauseFix
Real clock or timezoneInject time; fake timers; set TZ=UTC in CI
State shared between testsFresh fixtures, or a transaction rolled back per test
Test order dependenceRandomise order to expose it (pytest-randomly, go test -shuffle=on), then remove the shared state
Fixed sleepsPoll the condition with a deadline
Unseeded randomnessSeed it and print the seed on failure
Parallel tests using one resourceNamespace per test: schema, key prefix, temp directory, port 0
Data raceRun with go test -race; fix the unsynchronised access
External serviceReplace with a fake or container; keep live calls out of unit tests
pytest --count=20 -x tests/test_flaky.py        # repeat each test 20 times (pytest-repeat)
pytest -p no:randomly                           # pytest-randomly shuffles by default; this turns it off
pytest --randomly-seed=1234                     # replay the order from a failing run
go test -count=100 -race -run TestFlaky ./pkg   # repeat with the race detector
go test -shuffle=on ./pkg                       # random order; prints the seed to reuse with -shuffle=<seed>

pytest-randomly prints the seed it used at the start of each run; pass it back with --randomly-seed to reproduce the order.

Quarantining a flaky test often hides a real defect. Fix it, or delete it and cover the behaviour another way. A test nobody trusts costs CI time and teaches people to ignore red builds.

Shaping the pyramid#

The pyramid is a ratio, not a rule: many unit tests, fewer integration tests, a handful of end-to-end tests. Two inverted shapes show up in practice. The ice-cream cone has most checks in slow end-to-end suites because the code was never structured for unit testing, so every build takes 40 minutes and a failure points at nothing specific. The hourglass has a fat unit layer and a fat end-to-end layer with nothing in between, so the wiring between components (the SQL, the serialisation, the queue client) is exercised only by the slowest tests.

Push each check to the lowest layer that can see the failure: parsing and business rules to unit tests; the query, the migration and the HTTP contract to integration tests against a real container; the checkout flow to one end-to-end test. When an end-to-end test fails, write the unit or integration test that would have caught it, then decide whether the end-to-end case still earns its minutes. The unit layer should also cover error paths, which are where production incidents live and where end-to-end tests rarely go.

Code shape decides what is testable. A function that reads the clock, opens a connection and formats output in one body can only be tested end-to-end. Split it so the pure logic takes values and returns values, and a thin outer layer does the I/O; the pure part gets fast, exhaustive tests and the outer part gets one integration test. This is the same seam that makes a fake possible (see Test doubles).

Property-based testing#

Example-based tests check the cases you thought of. A property-based test states an invariant (round-trips, ordering, idempotency, equivalence to a slow reference) and the framework generates hundreds of inputs, then shrinks a failing input to the smallest one that still fails. It finds the empty string, the negative zero, the unicode combining character and the list of one element.

from hypothesis import given, strategies as st, settings

@given(st.text())
def test_encode_decode_roundtrip(s: str) -> None:
    assert decode(encode(s)) == s

@given(st.lists(st.integers()))
def test_sort_is_idempotent_and_ordered(xs: list[int]) -> None:
    once = my_sort(xs)
    assert once == my_sort(once)
    assert all(a <= b for a, b in zip(once, once[1:]))
    assert sorted(xs) == once                      # equivalence to a trusted implementation

@settings(max_examples=500, deadline=None)
@given(st.builds(Order, quantity=st.integers(min_value=1, max_value=10_000)))
def test_total_never_negative(order: Order) -> None:
    assert order.total() >= 0

Hypothesis saves failing examples in .hypothesis/ and replays them first on the next run; commit that directory or use @example(...) to pin a regression. Go’s testing/quick is deprecated; use rapid (pgregory.net/rapid) for properties, or the built-in fuzzer for anything that parses bytes:

func FuzzParse(f *testing.F) {
	f.Add("a=1")                                   // seed corpus
	f.Fuzz(func(t *testing.T, in string) {
		cfg, err := Parse(in)
		if err != nil { return }                   // invalid input is fine; a panic is not
		if out := cfg.String(); out != in {
			if _, err := Parse(out); err != nil { t.Fatalf("re-parse of %q failed: %v", out, err) }
		}
	})
}

go test -fuzz=FuzzParse -fuzztime=60s ./pkg mutates inputs guided by coverage; failing inputs land in testdata/fuzz/FuzzParse/ and run as ordinary regression tests from then on. In TypeScript, fast-check provides the same model (fc.assert(fc.property(fc.string(), s => decode(encode(s)) === s))). Properties that pay off most: round-trip (serialise then parse), invariants after any operation sequence (a state machine test), and “the fast version equals the obvious slow version”.

Mocking guidance#

Mocks are a design tool as much as a testing tool: the need to mock something deep is a sign the code has no seam there. Rules that keep mocks from turning a suite into a mirror of the implementation:

RuleReason
Mock only what you ownA mock of a third-party client encodes your guess about its behaviour; the real one differs on the edge case that matters
Prefer a fake over a mockA fake is one implementation shared by every test; a mock is re-specified in each test and drifts
Never mock the unit under testPartially mocked objects test the mock’s wiring, not the code
Assert outcomes, not calls, unless the call is the outcomeassert sent == [...] survives refactors; assert_called_once_with(...) breaks on every signature change
One mock per test, at the boundaryThree mocks means the test has become a re-statement of the code’s control flow
Verify the fake against the real thingA shared contract test suite that runs against both keeps the fake honest
# Fake: an in-memory implementation of the interface the code depends on
class FakeUserStore:
    def __init__(self) -> None:
        self.users: dict[int, User] = {}
    def get(self, user_id: int) -> User | None:
        return self.users.get(user_id)
    def save(self, user: User) -> None:
        self.users[user.id] = user

def test_deactivate_marks_user_inactive() -> None:
    store = FakeUserStore()
    store.save(make_user(id=7, active=True))
    deactivate(store, user_id=7)
    assert store.get(7).active is False            # outcome, not "save was called with ..."
# Contract suite: the same tests run against the fake and the real store
@pytest.fixture(params=["fake", "postgres"])
def store(request, postgres_container):
    if request.param == "fake":
        return FakeUserStore()
    return PostgresUserStore(postgres_container.get_connection_url())

def test_get_returns_saved_user(store):
    store.save(make_user(id=1))
    assert store.get(1) == make_user(id=1)

For HTTP, prefer a transport-level fake (httpx.MockTransport, Go’s httptest.NewServer, msw in Node) over patching the client library’s methods, because it exercises your serialisation and error handling. For time and randomness, inject a clock and a seed. If unittest.mock.patch needs a dotted path four modules deep, the code is telling you where the seam should be.

Contract tests in practice#

A consumer-driven contract test has three parts: the consumer writes a test against a mock provider that records the interaction; the resulting pact file is published; the provider’s CI replays it against the real provider and fails on mismatch. The consumer describes only the fields it reads, so the provider is free to add anything else.

# Consumer side (pact-python v3 API)
from pact import Pact

def test_get_user(pact: Pact) -> None:
    (pact.upon_receiving("a request for user 7")
         .given("user 7 exists")
         .with_request("GET", "/users/7")
         .will_respond_with(200)
         .with_body({"id": 7, "email": "alice@example.com"}, content_type="application/json"))
    with pact.serve() as srv:
        user = UserClient(srv.url).get(7)          # the real client code, against the mock provider
        assert user.email == "alice@example.com"
# Provider CI: verify every published pact for this provider, then record the result
pact-verifier --provider-base-url http://localhost:8080 \
  --broker-url "$PACT_BROKER_URL" --broker-token "$PACT_BROKER_TOKEN" \
  --provider my-api --provider-version "$(git rev-parse --short HEAD)" --publish-results
pact-broker can-i-deploy --pacticipant my-api --version "$(git rev-parse --short HEAD)" --to-environment production

can-i-deploy answers the deployment question directly: has this provider version been verified against every consumer version currently in production. Provider states (given("user 7 exists")) map to setup hooks in the provider’s verification harness that seed the data. Contract tests replace the shared staging environment for API compatibility; they do not replace an integration test of the provider against its own database. For gRPC and event schemas, buf breaking --against '.git#branch=main' and schema-registry compatibility checks (Avro, Protobuf, JSON Schema) do the same job with less machinery.

Snapshot and golden tests#

A golden test compares output against a file checked in beside the test. It suits rendered templates, CLI output, generated code and serialised structures, where writing the expected value by hand is tedious and reviewing a diff is easy. The trap is accepting a changed snapshot without reading it, which turns the test into “the output is whatever it was last time”.

var update = flag.Bool("update", false, "rewrite golden files")

func TestRenderManifest(t *testing.T) {
	got := RenderManifest(fixtureInput)
	path := filepath.Join("testdata", "manifest.golden.yaml")
	if *update {
		if err := os.WriteFile(path, got, 0o644); err != nil { t.Fatal(err) }
	}
	want, err := os.ReadFile(path)
	if err != nil { t.Fatal(err) }
	if diff := cmp.Diff(string(want), string(got)); diff != "" {
		t.Errorf("manifest differs (-want +got):\n%s\nrun: go test -update ./...", diff)
	}
}

pytest has syrupy (assert result == snapshot) and Vitest has toMatchSnapshot() and toMatchInlineSnapshot(), with -u to accept changes. Keep snapshots small and deterministic (no timestamps, sorted keys, stable IDs), and review snapshot diffs in pull requests as carefully as code.

CI gating#

A gate is a check that blocks a merge. Gates should be fast enough that people wait for them, strict enough that green means deployable, and few enough that each one is trusted. Put the fast, deterministic checks on every push and the slow ones on the merge queue or a nightly schedule, and make each gate report the failing test by name.

StageRuns onGateBudget
Lint, format, type-checkEvery pushRequiredUnder 2 minutes
Unit testsEvery pushRequiredUnder 5 minutes
Integration tests (containers)Every push or pull requestRequiredUnder 15 minutes
Contract verificationPull requestRequired for providersMinutes
End-to-endMerge queue, or nightlyRequired on the queue; advisory nightlyUnder 30 minutes
Load, mutation, fuzzNightly or weeklyAdvisory, with an owner who reads the resultsHours
# GitHub Actions: split the suite across runners and fail fast
jobs:
  test:
    strategy:
      fail-fast: true
      matrix: { shard: [1, 2, 3, 4] }
    steps:
      - run: pytest -q --splits 4 --group ${{ matrix.shard }} --durations-path .test_durations   # pytest-split
      - run: go test -race -count=1 -json ./... | tee test.json | go tool test2json >/dev/null    # machine-readable

Branch protection (or a merge queue) marks the jobs above as required checks so the merge button waits on them. Coverage gates (--cov-fail-under) should ratchet, not block: fail when coverage drops below the current value minus a small tolerance, so the number only moves up. Automatic retries of failed jobs hide flakes; if you must retry, retry only known-flaky tests through a plugin (pytest-rerunfailures, vitest retry) and track the rerun count as a metric with an owner, so the list shrinks. Cache dependency downloads, not test results; uv cache, the Go module and build caches, and ~/.npm are safe to restore, but a restored test cache can mask a change in an external input. See GitHub Actions for the workflow mechanics.

Running tests in CI#

Run the fast layers on every push and the slow layers before merge. Stop at the first failing layer, and make the log name the failing test and assertion so nobody needs to reproduce it locally to start.

ruff check . && mypy --strict src/ && pytest -q --cov=src --cov-fail-under=80
go vet ./... && go test -race -count=1 ./...
npm run lint && npx tsc --noEmit && npx vitest run --coverage

Go caches successful test results and reuses them when the package and its inputs have not changed, printing (cached). Go reruns a cached test when files it opened inside the module or environment variables it read have changed, but it cannot see files outside the module, databases or network services. -count=1 disables the cache. go clean -testcache clears the cache.

Troubleshooting test runs#

SymptomLikely causeCheck or fix
pytest fixture 'x' not foundFixture defined outside a conftest.py the test can seeMove it to conftest.py in the test’s directory or a parent; pytest --fixtures lists what is visible
pytest ModuleNotFoundError for your packagePackage not installed in the environmentuv sync (installs the project), or set pythonpath = ["src"] under [tool.pytest.ini_options]
pytest PytestUnknownMarkWarning or error with --strict-markersMarker not registeredAdd it under markers in the pytest config
Passes alone, fails in the full runShared state or order dependencepytest -p no:randomly, bisect the run, check module-level globals and fixture scopes
Passes locally, fails in CITimezone, locale, CPU count, missing service, cached resultCompare environments; go test -count=1; set TZ=UTC locally
Go test shows (cached) after an external changeGo only tracks files inside the module and environment variables-count=1
cannot use -fuzz flag with multiple packages-fuzz needs exactly one package and one matching targetName one package, e.g. ./pkg, and an anchored regex such as '^FuzzParse$'
Fuzz failure in CI but not locallyFailing input saved under testdata/fuzz/FuzzXxx/Commit that file; it becomes a regression case run by plain go test
Testcontainers cannot connect to DockerNo Docker socket (for example with Podman)Enable the Podman socket and set DOCKER_HOST to it
Vitest --coverage asks to install a packageCoverage provider not installednpm install -D @vitest/coverage-v8
Hypothesis Flaky or Unsatisfiable health checkTest depends on state between examples, or the strategy filters too hardMake the test pure; replace .filter with a constructive strategy
Hypothesis finds a failure CI cannot reproduceExample database not sharedCommit .hypothesis/ or add the input with @example(...)
Pact verification fails with state not foundProvider state handler missing for a given(...)Implement the state setup hook in the provider harness
Snapshot suite passes after a behaviour changeSnapshots were bulk-updated with -uReview snapshot diffs in the PR; never update without reading
Mock asserts pass but production breaksMocked a third-party client instead of your own boundaryReplace with a transport-level fake or a container
Coverage gate fails on an unrelated PRGlobal threshold above the current valueRatchet from the current value; measure changed files only
CI job green but tests did not runTest discovery found nothing (wrong path, missing __init__.py, glob mismatch)Fail on zero tests: pytest --co -q | grep -c ::, vitest --passWithNoTests off
Sharded CI runs unbalancedDurations file staleRegenerate .test_durations on a schedule