An existing codebase rarely needs a new architecture everywhere. It needs one costly boundary made safer to change. For a product manager, the hard part is resisting a migration plan that sounds coherent but has no release-sized first step. I would fund a narrow change tied to a product outcome, keep the rest of the system in place, and require evidence before expanding the work.
The first deliverable should be a boundary decision, not a target diagram
Start with a product change that the current codebase makes unusually expensive: adding a pricing rule, changing checkout eligibility, or releasing a new account setting. Ask engineering to trace that change through code ownership, synchronous calls, database writes, and deployment steps. That exercise produces a scope you can fund; a diagram of the desired future system does not, because it omits the compatibility work between the current and proposed designs.
The reference Modern Software Architecture Patterns for Scalable Systems can help name possible boundaries, but the current call graph should decide which boundary to attempt first. A pattern is a destination rather than a work estimate, because existing transactions, shared tables, and release habits determine how difficult it is to reach.
Make the team bring evidence from one representative workflow. OpenTelemetry 1.x instrumentation and W3C Trace Context headers can show where a request crosses process boundaries; PostgreSQL’s pg_stat_statements can expose queries repeatedly invoked by that workflow. Those tools answer different questions: traces show the route through services, while query statistics help identify database coupling that a service diagram may hide. Neither proves that a split is desirable, because a costly call may be easier to fix in place.
For example, suppose the last 14 days of observed releases show that a pricing change touches three teams and usually waits for a coordinated deployment. That is a stronger migration candidate than an isolated module with untidy code, because the coordination delay affects the product schedule. If Prometheus reports a measured 320 ms p95 for the same workflow, record it as a guardrail rather than automatically making latency the goal: the proposed boundary may reduce release coupling without making requests faster.
The initial decision record should fit on one page. It should name the product change, the code that owns its rules today, the callers and data it shares, the release dependency to remove, and one explicit non-goal. It should also identify who can accept a changed contract. If that owner is missing, discovery is not finished, because an interface cannot be made stable by the implementing team alone.
Compatibility work is the migration, not an afterthought
A useful first slice leaves old and new behavior able to coexist. For an in-process boundary, that may mean moving pricing decisions behind a single interface while keeping the same deployment and PostgreSQL database. For a network boundary, it may mean introducing a versioned HTTP contract and routing only a small class of requests to the new implementation. Both approaches demand tests against existing behavior, because customers encounter the transition rather than the architecture diagram.
Ask engineering to list every contract that must remain valid: API responses, error codes, database semantics, scheduled jobs, and administrative operations. An OpenAPI 3.1 specification can make an HTTP contract reviewable, while consumer-driven checks with Pact can catch a caller that depends on a field the team intended to remove. Neither replaces an integration test for the old transaction, because a contract test cannot establish that two writes still succeed or fail together.
Data ownership deserves its own line in the plan. If two parts of the application update the same row in one PostgreSQL transaction, extracting either part into a service changes failure behavior. A remote call cannot simply inherit the database transaction’s rollback guarantee. I would not add Kafka to that workflow as the first move, because introducing asynchronous delivery before defining ownership creates retry, ordering, and reconciliation work alongside the original change. Kafka may be appropriate later if the product can tolerate eventual consistency and a team owns the resulting recovery process.
Scope compatibility as visible work rather than burying it under “implementation.” Include characterization tests for current behavior, a migration path for stored data if ownership changes, dashboards for both routes, and a rollback that has been rehearsed. PostgreSQL documents lock_timeout = 0 as the default, meaning no lock timeout; a team planning a data migration should choose a nonzero setting appropriate to its workload rather than assume a schema change will fail quickly when blocked. The right value must be tested, because a timeout that protects a busy system can also interrupt a valid migration.
The second reference, Scalable Software Architecture Patterns for Modern Systems, is useful for comparing end states, but it cannot settle who owns a shared customer record next Tuesday. Put that ownership decision in the scope document, because unresolved data responsibility is likely to reappear as an integration dependency.
A modular monolith usually wins the first round against a new service
There are two credible options for the first boundary. Option A: a modular monolith wins when the product needs independent code ownership but can still deploy the application together. Its cost is interface design, tests, and enforcement inside the repository; its limitation is that releases and runtime resources remain shared. A team can enforce import rules with ArchUnit for Java or dependency-cruiser for JavaScript, then review violations in CI. That constraint is valuable only if someone owns exceptions, because a rule ignored under deadline pressure is not a boundary.
Option B: a strangler-style service wins when a narrow workflow genuinely needs an independent release or runtime profile and its data can have one clear owner. Its cost includes an HTTP contract, authentication between processes, routing, observability, deployment support, and handling partial failure. NGINX proxy_pass or Envoy weighted routing can send a defined slice of traffic to the new path, but routing does not solve data ownership; it only gives the team a controlled way to expose the change.
For an existing product, I would default to Option A until the team can show that shared deployment is the obstacle. That position is contestable: an independent service can be justified immediately when a small, well-owned component must ship on a different cadence. But paying for distributed failure modes before proving the need is an expensive way to reorganize code, because each call now needs timeouts, retries, monitoring, and an accountable operator.
Keep the comparison tied to the chosen product change. If pricing rules can be isolated behind an interface and released with the application, the modular monolith buys a reversible step. If pricing has a separately owned data model and repeated release conflicts remain after that isolation, a service becomes a more defensible next investment. This sequencing avoids pretending that the first boundary must also be the final deployment topology.
Fund the slice with exit criteria and a separate uncertainty allowance
A realistic plan has discovery, construction, and validation, each with a decision point. Discovery ends when the team can identify callers and shared writes. Construction ends when the old and new paths pass the same business examples. Validation ends when the release can be rolled back and the product owner accepts the observed result. Without those exits, “architecture work” can consume a quarter while producing no decision about whether to continue.
Here is a small, runnable Python 3 estimate worksheet. The numbers are planning assumptions, not measured delivery rates: replace the person-days after discovery. The output separates known work from an uncertainty allowance so the latter cannot quietly become a promised feature.
from math import ceil
days = {
"trace workflow": 3,
"define contract": 4,
"build first slice": 8,
"test and rehearse rollback": 5,
}
known = sum(days.values())
allowance = ceil(known * 0.30)
capacity_per_week = 10
print(f"Known person-days: {known}")
print(f"Uncertainty allowance: {allowance}")
print(f"Indicative weeks: {ceil((known + allowance) / capacity_per_week)}")
Here, 30% is a reserve to tune after inspecting the codebase, not a universal architecture multiplier. The assumed capacity of 10 person-days per week needs a reality check against support work and meetings, because two named engineers rarely spend every hour on the migration. The result is a conversation starter, not a date commitment: discovery may reveal a shared write that changes the option altogether.
Use a feature flag through OpenFeature, or an existing flag service such as LaunchDarkly, to control exposure if the old and new paths can safely run side by side. An initial 5% traffic share is a tunable release gate, not evidence that the design is safe. Watch error rate, p95 latency, and business outcomes for the affected workflow; compare them with the old path over a period that includes ordinary usage peaks. Turn the flag off if the agreed threshold is breached, because a rollback criterion decided during an incident is usually too late to be useful.
Set a stop condition as well as a success condition. If the first slice still requires coordinated changes across the same three teams, do not automatically fund a second extraction: the boundary may be in the wrong place. If it removes that dependency without changing customer-visible behavior, decide whether another slice is worth its cost. That keeps architecture investment answerable to the product constraint that justified it.
The next step is to commission a trace, not a migration
Choose one recently delayed product change and ask for its callers, shared writes, deployment dependencies, and current release time. Give engineering a short discovery window and require a boundary recommendation with both options priced. Do not approve a service, event stream, or repository split yet. Approve the smallest reversible slice only after the team can explain what it will preserve and how it will retreat.
