Skip to main content

How to audit a software architecture before a rebuild

A rebuild is rarely decided on facts. It is decided because the team has had enough, because a new manager arrived, or because one feature took three months. An audit turns that feeling into a decision you can defend.

Christophe Bellec ·

Short answer

An architecture audit before a rebuild runs in five steps: map the system as it actually runs, list pain points that are measured rather than felt, find where the cost of change concentrates, present two or three costed options including doing nothing, and produce an incremental plan. It typically takes two to ten days depending on system size, and it often concludes that a full rebuild is not needed.

Why audit before, not during

A rebuild commits the team for months, freezes the product roadmap, and in one case out of two reproduces the very problems it claimed to solve, because those problems were not where everyone assumed.

The audit is not there to rubber-stamp a decision already made. It is there to find out what the problem actually is, and it has to be able to conclude that no rebuild is warranted. An audit that can only end in “yes, rebuild” is not an audit.

Step 1: Map the system as it really is

Not the diagram in the documentation: the system as it runs. The two always diverge, and the gap is itself information.

  • The deployed components, and what actually calls them in production.
  • The data flows between them, including the forgotten integrations: a nightly export, a script on a server nobody restarts.
  • External dependencies and what happens when one is unavailable.
  • The environments, and how closely they genuinely resemble production.
  • What is no longer used but still deployed, often a significant share of the system.

A useful starting signal: ask three people to draw the architecture. If you get three different diagrams, the main problem may be shared understanding rather than technology.

Step 2: Collect the pain, then measure it

Start by listening: developers, product, support, operations. Each sees a different face of the system, and support often sees what engineering does not.

Then look for the factual trace of each complaint, because that is what ranks them:

  • Lead time between a feature being ready and it reaching production.
  • Share of time spent fixing rather than building.
  • Frequency and duration of incidents, and which modules are involved.
  • Time it takes a new joiner to ship their first change.
  • The modules everyone avoids touching, and why.

The gap between the pain people describe and the pain the numbers show is regularly the most useful output of the audit.

Step 3: Locate the cost of change

An architecture is not judged on elegance but on what it costs to evolve. So the central question is: when something has to change, where does the difficulty appear?

  1. Take the last five significant changes and retrace what they genuinely required touching.
  2. Spot the components modified every single time, whatever the feature: those are the real bottlenecks.
  3. Cross-reference with the repository history: files that change often and change together reveal coupling the diagram does not show.
  4. Identify what prevents deploying one part without deploying everything.

That cross-referencing almost always narrows the diagnosis to two or three precise areas, very rarely to “the whole architecture”.

Step 4: Put costed options on the table

An audit that produces a single recommendation does not help anyone decide. You need two or three, with their consequences:

Change nothing
Always include it, even when it is obviously bad. It provides the baseline: what inaction costs over twelve months, in velocity and in incidents.
Fix the bottlenecks
Address the two or three areas found in step 3 and leave the rest alone. In most cases this is the best effect-to-risk ratio available.
Rebuild a bounded scope
Rewrite one specific component behind a stable interface, with a progressive cutover. Expensive, but reversible and shippable in stages.
Rebuild everything
Reserved for cases where the technology is no longer supported, where nobody left knows how to evolve the system, or where the business model has fundamentally changed. It then has to be owned as a project in its own right, with its budget and its roadmap freeze.

Step 5: Produce a plan that ships in pieces

A useful plan is ordered by value against risk and cut into increments that each deliver an observable benefit. If the first benefit arrives after six months, the plan will not survive the first shift in priorities.

  • First, whatever reduces immediate risk: backups, observability, tests on the critical paths.
  • Then whatever unblocks the team: reliable environments, automated deployment, decoupling the main bottleneck.
  • Finally the structural changes, one at a time, each behind a stable interface.
  • At every step, an observable criterion: lead time, incident count, build time.

What the audit has to leave behind

  • A diagram of the real system, readable by a non-specialist.
  • A prioritised list of problems, each with its measured effect.
  • The options, their costs and their risks, including inaction.
  • A plan cut into steps, with success criteria.
  • The decisions written down, so they are not re-debated in three months.

If the audit only produces a report nobody knows how to apply, it failed. The test: after reading it, a team should be able to start the first step the following week.

Frequently asked questions

  • How long does an architecture audit take?

    Between two and ten days depending on the size of the project and the depth of analysis expected. Two days are enough for a diagnosis and the quick wins on a mid-sized system; beyond that you move into detailed code and data-flow analysis.

  • Should development stop during the audit?

    No, and it is better that it does not: watching the team ship during the audit yields information no interview provides, particularly about the real friction in the development cycle.

  • Can the audit conclude that no rebuild is needed?

    That is a frequent conclusion, and often the most useful one. Fixing the two or three identified bottlenecks costs a fraction of a rebuild and delivers most of the expected benefit.

  • Who should be involved?

    The developers working on the system day to day, product, and someone from operations or support. The three viewpoints do not overlap, and support often holds the best list of symptoms.

How I can help on this

Get in touch

More articles