Performance & Optimization

QA Can Benchmark an Inherited Codebase Without a Rewrite

You inherited a codebase with few tests, frequent regressions, and no appetite for a rewrite. I would not begin by measuring coverage or automating the oldest modules, because neither tells you which failure will hurt the next release. Start with one observable behavior that has already broken, make its test repeatable, and let that evidence determine where the next test belongs.

The incident log is a better starting map than the module tree

The article Add test automation to legacy code without a rewrite QA playbook offers a useful menu of techniques, but I would resist turning it into a repository-wide rollout plan because the code most in need of protection may occupy only a small part of the tree. An inherited system usually has more possible test targets than a QA engineer can investigate before the next release. Recent failures make that choice less speculative.

Take a working sample of 12 recent bug reports; treat that count as an adjustable starting size, not a statistical claim. For each report, record the user action, the input that mattered, the visible wrong result, and whether the behavior still exists. Link the report to the Git commit that fixed it when possible. A report saying “orders failed” is too broad to automate, while “an order with an expired discount returned the undiscounted total but retained the discount label” describes an assertion you can check.

Rank candidates by the cost of another failure and by how reliably you can recreate their inputs. A severe incident is not automatically the first test: if it requires an unavailable partner service and unknowable data, a smaller reproducible failure may establish a usable test path sooner. Keep that trade-off visible in the ticket rather than disguising it as a risk score with invented precision.

Before writing a test, run the behavior manually against a local build and note the environment: application revision, database schema revision, feature flags, timezone, and external-service responses. Those details matter because an assertion against yesterday’s data is not a stable description of today’s code. Use a sanitized fixture or a disposable PostgreSQL database where practical; do not copy production records into a test repository, because stable tests do not justify exposing customer data.

I would also avoid a coverage target at this stage because JaCoCo 0.8.x or coverage.py can report executed lines without showing whether the test detects the incident you care about. Coverage can later reveal an unexercised branch inside a chosen path. It cannot choose that path for you.

Protect the boundary before changing the code behind it

For a legacy service, the first useful assertion often sits at an existing HTTP boundary because it checks behavior without requiring you to untangle internal dependencies. Pick a read-only route backed by controlled data. Assert a field that expresses the reported failure, not merely that the server returns HTTP 200; a successful status can accompany a wrong order state.

The following Bash probe runs with curl and jq against an existing JSON route. Set BASE_URL, GET_PATH, and EXPECTED_STATE to a known fixture before running it; the route must return an object with a state field. The temporary file keeps the response available for inspection during the check and removes it afterward.

#!/usr/bin/env bash
set -euo pipefail
: "${BASE_URL:?Set BASE_URL}"
: "${GET_PATH:?Set GET_PATH}"
: "${EXPECTED_STATE:?Set EXPECTED_STATE}"
body="$(mktemp)"
trap 'rm -f "$body"' EXIT
code="$(curl -sS -o "$body" -w '%{http_code}' "$BASE_URL$GET_PATH")"
test "$code" = "200"
jq -e --arg expected "$EXPECTED_STATE" \
  '.state == $expected' "$body" >/dev/null

This is a thin characterization check, not proof that the entire order flow works: it observes one response for one fixture. Run it against the current revision before a fix, then record whether it fails in the way the bug report predicts. If the bug is already fixed, run it against a known failing Git revision when that is safe and feasible; otherwise, describe the missing failure evidence in the test review. A test that has never distinguished wrong behavior from right behavior needs more scrutiny than its green result suggests.

For state-changing paths, a read-only probe is insufficient because the defect may appear only after a write. Use an isolated database and assert both the response and the resulting record. Docker Compose can start a local dependency stack, while Testcontainers can provision dependencies for automated runs; either is preferable to relying on a shared staging record because another test or person can change that record between assertions. Keep fixture creation explicit so a failed run can be reproduced without knowing which staging account was used.

Do not add a production code seam merely to make the first test elegant. An awkward boundary-level test can be acceptable temporarily if it observes an important regression and leaves the application unchanged. Introduce an injectable clock, repository interface, or smaller unit seam when the boundary test shows exactly which uncontrolled dependency makes it slow or unreliable; that evidence makes the code change easier to review.

API checks should usually precede browser checks, but not replace them

Compare pytest with Requests against Playwright 1.x for the same failing scenario. Pytest with Requests wins when the defect is visible in an HTTP response or persisted state, because it avoids browser rendering and can set up fixtures directly. Its cost is that it will miss a broken form, client-side validation, or a button that never sends the request. Playwright wins when the failure depends on those browser interactions, because it exercises the path a user takes. Its cost is more setup and more opportunities for timing or selector failures.

That comparison is a choice about the failure mechanism, not a ladder every test must climb. If an API returns the wrong total, I would first assert the total through Requests and add a browser check only if the interface has separately failed to display or submit it. If the API is correct but a save button is disabled for a valid input, an API test cannot protect the defect, however quickly it runs.

The companion article How QA can add safety nets to legacy code without rewriting makes a fair case for several layers of protection, but I would give each new test one reason to exist because duplicated assertions at every layer increase maintenance without necessarily detecting a different failure. Write that reason in the test name or review description: “expired discount remains visibly labeled” says more than “order page works.”

For an initial CI budget, try a tunable ceiling of 2 minutes for the small set of release-blocking checks. That is a planning constraint, not a measured property of your system; change it after observing runs. Export pytest results as JUnit XML so GitHub Actions can show which case failed, and keep diagnostic output that identifies the fixture and revision without printing secrets. If a Playwright check fails intermittently, investigate its waits and selectors before making it blocking. Repeatedly rerunning a flaky check can produce a green build while leaving the cause untouched.

A test earns release authority only after it survives routine change

New tests should first run alongside the existing build without blocking merges, because an inherited environment may have undocumented dependencies that only CI exposes. Review failures against the same commit and fixture before promoting a check to a gate. A suggested observation window is 20 consecutive CI runs; use it as a threshold to tune, not as a guarantee of reliability. Record failures as application defects, fixture defects, environment defects, or unknown rather than counting every red run as a product regression.

Promote a check when its failure gives a developer an actionable next step. A useful failure identifies the scenario, expected behavior, actual behavior, and relevant response or log reference. OpenTelemetry trace IDs can help connect an HTTP failure to service activity if the application already emits them; adding tracing solely for one test may cost more than keeping a concise local log. Likewise, a Pact contract is worth considering when a provider change has repeatedly broken a consumer, but it is overhead for a single service whose boundary test already observes the behavior.

Keep the gate narrow enough that people trust it. When a test fails after an intentional behavior change, require a reviewer to check the new expectation against the incident and product decision before updating the assertion. Otherwise, “fixing the test” can quietly remove the protection it was introduced to provide. When a test fails because its fixture depends on mutable shared data, repair the fixture before adding a retry; retries hide that dependency rather than removing it.

A proposed review metric is the share of blocking-test failures that identify a real, reproducible problem. Measure it from your own CI history instead of borrowing a benchmark, because environments and failure classifications differ. This metric is more useful for deciding whether to keep a gate than a rising test count: a small suite that people investigate protects releases better than a large suite they routinely bypass.

Start with one failure you can make happen twice

Choose one recent regression with a known input and a visible wrong result. Reproduce it on a disposable environment, save the Git revision and fixture, then write the smallest boundary check that fails for that reason. Run it again after the fix and in CI. Only then choose the next incident; you will have a working route into the codebase rather than a rewrite plan.