Scaling an existing product rarely starts with a clean architectural decision: it starts with a roadmap, a release process, and code that still serves customers. My position is that a product manager should fund one reversible change at a time, tied to a measured constraint. A target architecture can guide those changes, but funding the target state before proving the first migration step makes the scope misleading.
The first budget should buy evidence, not a new platform
Architecture diagrams make change look cheaper than it is because they show components rather than the contracts, data ownership, and release habits that connect them. Read Modern Software Architecture Patterns for Scalable Systems as a menu of possible destinations, not as a sprint backlog. The incident log and product roadmap should determine which destination, if any, deserves work now.
Start with a decision brief for one customer-visible problem. State the affected journey, the current failure, its frequency, and the cost of leaving it alone. For example, suppose OpenTelemetry traces from your own application measure a checkout p95 of 780 milliseconds over the last 14 days. That is a useful baseline only if the traces separate database time, application time, and calls to other systems; otherwise, a new service could preserve the same bottleneck behind an additional network hop. Ask engineering to pair that trace with PostgreSQL’s pg_stat_statements and an EXPLAIN (ANALYZE, BUFFERS) result for the expensive query before estimating a split.
Give the team a short, explicit discovery allocation rather than an open-ended “architecture phase.” As a planning assumption, reserve two engineers for one week to identify the slow path, its callers, its data writes, and its release dependencies. The deliverable is a choice among a query change, an internal module boundary, or a separately deployed service, with evidence for each. This is a scope-control step because the cheapest effective intervention may not require a new runtime at all.
I would not begin by moving all endpoints behind a new gateway, because that changes a broad request path before the team has shown that routing is the constraint. I would also avoid promising a throughput multiplier from a diagram, because actual capacity depends on contention, query plans, and downstream limits that the diagram cannot measure. A product manager can approve a problem and a testable outcome without approving a preferred pattern in advance.
A narrow seam makes migration estimable
Once the constraint is known, choose a seam around behavior that can be routed or replaced without changing every caller. In a checkout example, that might be price calculation rather than the whole order flow: its inputs can be specified, its output can be compared, and the existing application can remain responsible for taking payment. Define the request and response in OpenAPI 3.1, then record the less visible rules: rounding, tax precedence, error responses, and which system owns the final amount. A boundary is not ready for estimation if those rules live only in the original code.
Keep the first implementation inside the existing deployable where possible. An internal module with a narrow interface lets engineers test the boundary before paying for separate deployment, networking, authentication, and operational ownership. Use Pact consumer tests where independently released callers need contract protection; use ordinary integration tests where callers ship together, because maintaining a contract broker for a single coordinated release adds work without providing release independence.
Plan for a parallel check before switching customers. For read-only or deterministic calculations, the existing path can return the answer while a candidate implementation calculates a second answer for comparison. Record mismatches in Prometheus and inspect them in Grafana, but do not send duplicate payment or inventory writes: a “shadow” write can still have a real effect. The comparison needs an agreed tolerance. Exact equality may be appropriate for a price in minor currency units, while a derived recommendation may need a different rule because its output is not an accounting record.
A small rollout flag makes the eventual switch reversible. This Python 3.11 example assigns the same account to the same cohort on every run; the 5% value is a starting setting to tune, not a claim that 5% is universally safe:
import hashlib
def enabled(account_id, percent):
digest = hashlib.sha256(account_id.encode()).digest()
bucket = int.from_bytes(digest[:4], "big") % 100
return bucket < percent
for account_id in ("acct-17", "acct-21", "acct-35"):
print(account_id, enabled(account_id, 5))
In production, store the percentage in an audited flag system such as Unleash, and define whether the unit of assignment is an account, user, or request. Account-level assignment is often easier to support because everyone investigating one account sees the same behavior. The flag is a routing control, not a substitute for a rollback plan: the old path must remain compatible with new data until switching back is demonstrably safe.
Compatibility and rollback consume more work than extraction
The code moved into a new module or service may be the smallest line item because existing callers still expect the old behavior. Scope the migration as distinct work packages: document the contract, make data ownership explicit, build the replacement, compare outputs, release to a cohort, monitor it, and retire the old path. Give each package an exit condition. “New service deployed” is not an exit condition for the migration if the original application still owns the decisions and the new service cannot be disabled safely.
Data is usually the hardest boundary to reverse because two writers can disagree even when their APIs match. Prefer one authoritative writer during the first migration. If another component needs updates, a transactional outbox in PostgreSQL 16 can commit the business change and an outgoing event in one database transaction; a worker can then deliver events with retries. The consumer must tolerate duplicate delivery because a worker can crash after sending an event but before recording success. Do not add Kafka solely to move the first event: a broker’s operational cost is difficult to justify until the product needs its retention, replay, or multiple independent consumers.
Include database behavior in the acceptance criteria rather than treating it as an implementation footnote. PostgreSQL 16 documentation gives statement_timeout a default of 0, meaning disabled; an overloaded query may therefore wait longer than a customer request unless the team sets an appropriate limit. That published default is a reason to inspect production configuration, not a recommendation to copy one timeout everywhere. Set a request deadline and a database timeout together so that abandoned requests do not continue consuming capacity without a useful response.
Write down the rollback boundary before launch. If the candidate only reads data, turning off a flag may be enough. If it introduces a new field or write format, the old code must read that format before the new writer is enabled. If it changes ownership of a table, rollback may require a reconciliation procedure rather than a toggle. This difference should appear in the estimate because “reversible in minutes” and “recoverable through a data repair” are materially different product risks.
For the first release, agree on a short watch period and named decision makers. A proposed trigger might be a checkout error rate above 1% for five minutes, but that threshold should be tuned against the product’s normal variance and traffic volume. Report both the rollout cohort and the unaffected cohort; a global dashboard can hide a regression affecting a small flagged group. The product manager’s job here is to protect time for the watch, comparison, and possible reversal rather than counting deployment as completion.
A modular monolith usually wins the first round
The choice between a modular monolith and an extracted service should be explicit. The modular monolith wins when the code shares a release cadence and data transaction, because a module boundary can reduce coupling without adding network failure modes; its cost is continued shared deployment and the discipline to keep imports inside the boundary. An extracted service wins when a team needs independent releases or isolation of a proven workload, because separate deployment can then remove a real coordination or capacity constraint; its cost includes API versioning, authentication, on-call ownership, observability, and failure handling.
That is where I would push back on a pattern-led plan. Scalable Software Architecture Patterns for Modern Systems can help name a plausible end state, but the existing codebase has to earn each new runtime through an operational benefit. If two teams still require a synchronized release for every change, separate services may simply turn an in-process dependency into a distributed one without giving either team autonomy.
Make the comparison visible in the project estimate. As an illustrative budgeting assumption, allocate three weeks to extracting and testing the first internal boundary, then price a service as an additional option rather than hiding it inside that estimate. The service option should include deployment configuration, service-to-service credentials, alerts, runbooks, and a rollback rehearsal. If the team uses Kubernetes, a successful rollout status is evidence that pods became available, not that the new behavior preserved customer outcomes; the cohort metrics and contract checks still decide acceptance.
Ask for a decision date, not a standing mandate to “modernize.” At that date, engineering should present the observed bottleneck, the cost of the internal boundary, and the incremental cost of extraction. Product can then compare that spend with the feature work it displaces. This makes an architecture decision a priced product trade-off instead of an unbounded prerequisite.
Approve the first reversible change before the program
Put one affected customer journey on the next planning agenda and ask for its baseline trace, owning team, and rollback boundary. Fund the smallest change that can improve that journey, with a date for reviewing the result. If the evidence then supports a service split, approve that split with its operational costs visible. If it does not, keep the boundary and spend the remaining budget elsewhere.
