Inherited Untested Code? Scale Will Expose the Seams

An inherited codebase can survive a traffic increase while quietly returning the wrong customer’s order, charging twice, or showing a status that never existed. For a QA engineer, those failures…

An inherited codebase can survive a traffic increase while quietly returning the wrong customer’s order, charging twice, or showing a status that never existed. For a QA engineer, those failures matter more than a slower p95 response. My position is that the first scaling risk to investigate is a broken business invariant under concurrency, not insufficient capacity. A faster system that makes more wrong decisions is not an upgrade.

The first scaling failure may return HTTP 200

Latency dashboards are easy to watch; incorrect outcomes are harder to notice because the request can succeed. A retry can create two orders, a cached lookup can cross a tenant boundary, and two workers can both accept the last available item. Each response may be fast and individually plausible. In an untested codebase, you cannot assume the path exercised by a load test is the path that enforces the business rule.

Start with one invariant that would be costly to violate: an order belongs to exactly one tenant, a payment request creates at most one charge, or stock never falls below zero. Trace it through the HTTP handler, authorization check, database transaction, cache, and background worker. This is narrower than mapping the entire application because it gives you a result you can assert even when the implementation is unfamiliar.

Scope an Architecture Upgrade Without Rewriting the Product puts a useful boundary around product work, but I would put the observable invariant ahead of the component boundary because a neatly limited change can still multiply an existing correctness defect. I would not begin by adding replicas or increasing a Kubernetes autoscaling/v2 HorizontalPodAutoscaler target: extra workers increase the number of simultaneous paths through shared state, so they can make races more frequent before they improve safety.

Look particularly for assumptions that held when requests rarely overlapped. PostgreSQL 16 documents Read Committed as its default isolation level; two statements in one transaction can therefore observe different committed data, which matters if the code reads a value, makes a decision, and later writes without an appropriate constraint or lock. A Redis 7 cache introduces a different question: whether its keys include the tenant and authorization context required by the lookup. Neither PostgreSQL nor Redis is inherently the bug; the missing assertion between a business rule and its storage behavior is.

Write down the expected result before running a heavier test. For a tenant-bound order read, the owner should receive HTTP 200, while another tenant should receive HTTP 403 or 404, according to the application’s published contract. RFC 9110 defines those HTTP status codes, but it cannot tell you whether the application selected the correct tenant. That distinction is why a successful status counter alone is weak evidence.

A small concurrent probe is worth more than a broad green suite

Inherited test suites often prove that routes respond, not that two callers cannot interfere. Build a disposable fixture first: create an order owned by tenant A, obtain valid tokens for tenants A and B, and make its identifier available to the probe. Use a nonproduction environment with representative authorization and cache configuration. A mock that bypasses either layer would make this particular result less useful.

The following Python 3 probe uses only the standard library. Set BASE_URL, ORDER_ID, TOKEN_A, and TOKEN_B in the environment before running it; TOKEN_B must belong to a different tenant. Its eight workers, 32 requests, and five-second timeout are initial probe settings to adjust for your environment, not a claim that this volume establishes safety.

python3 - <<'PY'
import os, urllib.request, urllib.error
from concurrent.futures import ThreadPoolExecutor
def fetch(token):
    request = urllib.request.Request(os.environ["BASE_URL"].rstrip("/") + "/api/orders/" + os.environ["ORDER_ID"], headers={"Authorization": "Bearer " + token})
    try:
        with urllib.request.urlopen(request, timeout=5) as response: return response.status
    except urllib.error.HTTPError as error: return error.code
with ThreadPoolExecutor(max_workers=8) as pool:
    codes = list(pool.map(fetch, [os.environ["TOKEN_A"], os.environ["TOKEN_B"]] * 16))
assert codes[::2] == [200] * 16, codes
assert all(status in (403, 404) for status in codes[1::2]), codes
print("32 tenant checks passed")
PY

A pass establishes only that this fixture survived this run. It does not establish that every route is isolated, because the probe uses one order, two identities, and a short burst. A failure is more informative: preserve the response status, request ID, tenant IDs in a secure test log, and the deployment revision so a developer can reproduce the boundary that failed. Do not put bearer tokens into the bug report.

Next, adapt the same method to a state-changing invariant rather than merely turning up concurrency. For an order-creation endpoint, submit the same idempotency key during retries and query the resulting records or payment provider test account afterward. Define the expected number of durable effects before sending requests. If a staging run records six duplicate records from 2,000 replayed requests, report those as observations from that run, along with the fixture and revision; a percentage without the replay conditions hides the failure mechanism.

pytest 8 can turn these probes into repeatable assertions, while k6 can generate longer overlapping request sequences after the assertions exist. A k6 run with –vus 8 and –duration 2m uses tunable load settings, not a universal traffic target. Keep the oracle outside the response under test: if the endpoint says “created,” check the database or a separately exposed read path as well, because a handler can acknowledge work that its background worker later loses.

Tracing should explain a violated invariant, not replace its assertion

Once a probe can detect a wrong result, instrument the path that produced it. OpenTelemetry 1.x spans can connect an incoming request to a database call and queued job; include an opaque request or order identifier so a failed assertion can be followed across processes. Avoid tenant names, tokens, and raw order payloads in span attributes because trace storage is not an authorization boundary. If sampling drops the interesting request, reproduce it with targeted sampling or a controlled test trace rather than treating an empty trace search as proof of correctness.

Prometheus histograms remain useful, but attach them to the assertion question. Track request duration and, separately, counters for rejected cross-tenant access, idempotency-key conflicts, and worker retries. A p99 latency improvement does not excuse a rise in duplicate effects because the measurements describe different outcomes. Keep labels bounded—route templates rather than raw order IDs—because a unique identifier on every time series can make the monitoring system itself expensive to operate.

For database-backed decisions, pg_stat_statements can identify queries whose execution time or call count changed during the probe. It cannot prove that an update was logically correct, because it aggregates query behavior rather than validating order ownership or final stock. Capture the affected row count, relevant constraint violations, and a before-and-after fixture query alongside it. If a failed assertion coincides with a lock wait, that is a lead to investigate; it is not yet evidence that removing the lock is safe.

Set a correctness gate separately from a performance target. Zero cross-tenant disclosures is an acceptance requirement because even one disclosed order violates the stated boundary; the number of requests you run to seek counterexamples should grow with risk and available test time. A provisional p95 target of 300 milliseconds, by contrast, is a value to tune against your measured baseline and user expectations. Do not trade the former for the latter merely because the latency result is easier to graph.

Scope Architecture Changes Without Rewriting the Codebase is a useful planning frame, but a change list without a before-and-after invariant probe cannot distinguish a safe move from a faster regression. The QA deliverable is therefore not a blanket sign-off on the new diagram; it is a reproducible claim about what remained true while the system was stressed.

The right test boundary depends on the failure you need to catch

Two reasonable options for making that claim are Pact v4 consumer contract tests and Playwright browser journeys. Pact wins when a service boundary is changing and you need to show that a provider still meets the requests its consumer actually makes. Its cost is maintaining consumer examples, provider verification, and realistic states; a passing contract will not catch a race inside the provider if the examples never create one. Playwright wins when authorization depends on cookies, navigation, or browser-visible state that an HTTP-only contract misses. Its cost is slower setup and more environmental sensitivity, especially when tests share accounts or data.

For an inherited codebase, I would choose the option that reaches the threatened invariant with the fewest unverified substitutes. If the defect appears only when a browser session changes tenant context, use Playwright with separate sessions and isolated fixtures. If a new worker or service consumes an order event, use Pact for the message shape, then add a state assertion for duplicate delivery. Neither option should claim to prove exactly-once processing: consumers can receive a message again, so the durable effect must be checked after replay.

OpenAPI 3.1 descriptions can help identify documented response shapes, but they are a starting inventory rather than a correctness oracle because a schema cannot express every ownership or idempotency rule. Compare the description with observed behavior and record discrepancies as questions before freezing them into tests. Otherwise, characterization tests can preserve a defect simply because the current application consistently exhibits it.

Make the test environment adversarial in one controlled way at a time. Delay a worker after it receives an event, repeat a request with the same idempotency key, or run two authorized callers against the same record. Then reset the fixture and repeat the experiment. Changing traffic, cache configuration, worker count, and data simultaneously may reproduce a failure, but it leaves little evidence about which boundary broke first.

The first useful upgrade artifact is a failing assertion

Tomorrow, choose one high-impact invariant and create its fixture in staging: an owned order, two tenant identities, and an expected unauthorized result are enough to start. Run the concurrent probe, save the revision and request IDs, and repeat after the proposed change. If it fails, you have a specific defect to fix before scaling. If it passes, expand the conditions rather than calling the architecture safe.