AI-assisted engineering
The economics of AI-assisted development sit beyond code generation
Cheaper code generation can improve delivery, but the economic result depends on review, integration, correction and what the organisation does with released capacity.
The price of a coding tool is easy to identify. The cost of a completed software change is harder: it includes understanding the requirement, implementing it, reviewing it, integrating it, verifying it and supporting the result.
AI can alter several of those activities. It can also create additional work when the output is plausible but incorrect. An economic assessment that counts only generation time is therefore answering a narrower question than the business usually needs answered.
I would measure the cost of an accepted outcome and the capacity of the whole delivery process. That approach is useful whether the team is building a small internal tool or maintaining a consequential enterprise application.
Name the unit being purchased
A generated function, a pull request and a deployed feature are different units. Comparing their costs without distinguishing them can make a delivery process appear dramatically cheaper while leaving the business waiting just as long.
Choose an outcome with a clear acceptance boundary: a verified defect fix, a production-ready workflow or a documented infrastructure change. Include the evidence required to accept it. The outcome should remain the same when comparing assisted and unassisted work.
For each outcome, record human effort across clarification, implementation, review, correction and verification. Add tool usage and relevant execution costs. Keep the arithmetic transparent enough that another person can see what was counted and what was excluded.
Do not convert speculative defect reduction or future revenue into a saving without an evidence basis. Those effects may matter, but they belong as assumptions or separately measured outcomes, rather than adjustments that make the comparison look favourable.
Separate elapsed time, effort and cash
A change can reach deployment sooner without consuming fewer engineering hours. Several people may work in parallel, or an agent may reduce waiting while review effort remains unchanged. Faster delivery is valuable, but it is not the same claim as lower labour cost.
Likewise, reducing engineering effort does not automatically reduce expenditure. If staff costs remain fixed, the benefit may be additional capacity. That capacity becomes economically useful when the organisation assigns it to work that matters, such as reducing a backlog, improving reliability or handling more demand.
An avoided cost has its own evidence requirement. If AI assistance removes the need for a planned external engagement, record what was actually avoided. If the budget has not changed, describe capacity gained rather than money saved.
This distinction prevents a common accounting mistake: treating every minute released as an immediate cash saving while also claiming the value of the additional work completed in that same time.
Look for the constraint in the delivery process
Suppose implementation becomes quicker but the same reviewer must still assess every change. More generated work can arrive at the review queue without increasing the rate of accepted work. The team has changed the arrival rate, not necessarily the delivery capacity.
Measure queue time as well as active effort. Is work waiting for a requirement decision, security review, test environment or deployment window? Those delays tell you where further investment might help.
Sometimes the best economic change is better repository guidance or a stronger verification harness. It can reduce repeated explanation and make generated changes easier to assess. In other cases, the constraint is a business decision that a faster model cannot resolve.
Do not solve a review bottleneck by removing review where it carries necessary assurance. Improve the size, clarity and evidence of the changes so that the reviewer can reach a reliable decision with less effort.
Use a transparent worked comparison
Here is an illustrative arithmetic example, not a measured result. Assume a comparable change needs eight hours of human effort without AI assistance: four for implementation and four for clarification, review, correction and verification.
| Illustrative case | Implementation | Other human work | Total human effort |
|---|---|---|---|
| Unassisted baseline | 4 hours | 4 hours | 8 hours |
| Assisted, more review | 2 hours | 5 hours | 7 hours |
| Assisted, same review | 2 hours | 4 hours | 6 hours |
| Assisted, substantial correction | 2 hours | 7 hours | 9 hours |
The second row releases one hour of capacity, before tool and execution costs. A claim that development became twice as fast would describe only the implementation portion. The last row consumes more human effort despite quicker implementation. These are alternative assumptions, not forecasts.
The example shows which variables to measure. The decision should use actual task records and an agreed cost basis, with a separate account of delivery time and post-release quality. It should not depend on choosing the most flattering denominator.
Be careful when transferring evidence
METR's early-2025 study examined experienced developers working in familiar open-source repositories. It found a difference between expected benefit and measured completion time. That is evidence about a particular population, task setting and generation of tools.
It should neither be ignored nor used as a universal conclusion about 2026 engineering. A new repository, a repetitive transformation and a mature product change can have different review and context demands. Tool capability and working practices also change.
Use published research to challenge assumptions, then measure the workflow actually being considered. Include unsuccessful attempts and tasks where the team decides not to use AI. Excluding those cases selects only the favourable part of the experience.
Count the operating consequences
Generated software creates future obligations: dependency updates, incidents, security review, documentation and support. Lower implementation effort does not make those obligations disappear.
Track corrective work after release, changes reopened during review and defects discovered outside the initial acceptance run. Avoid crude comparisons based on defect counts alone when changes differ substantially in risk and complexity. The purpose is to understand whether the assisted process is producing supportable outcomes.
Also consider the opportunity created by better coverage. As discussed in the existing platform-engineering article, some useful analysis may become practical that was previously skipped. Its value may appear as better decisions or clearer evidence rather than fewer hours spent on an identical task.
The economic case should state which benefit is being sought: lower effort, shorter lead time, greater capacity or improved assurance. AI-assisted development earns its place when that benefit survives the full delivery and operating account, with the assumptions visible and the accepted outcome held constant.