The software factory I want to build has two central concepts: Change and Impact.
Change
Change is the work the factory performs.
Impact
Impact is what that work achieves for customers, employees, and the business.
Between them sits a feedback system that determines what the factory should do next—and how the factory itself should improve.
My argument is that AI can become the main actor in this system: investigating problems, proposing solutions, implementing changes, validating them, and gathering evidence of their effects. Making that work requires us to manage both sides deliberately. Change needs an operational process. Impact needs measurement. The inputs that connect them need ownership and rules.
That is a bigger ambition than giving every developer an AI coding assistant. It changes how an organisation decides what to build, who can initiate delivery, and how it evaluates the economics of software.
A factory that learns from every change
Change and Impact sit at the centre. Each feedback loop improves the next decision.
Factory inputs need management too
Closing a loop means turning evidence into a decision. Collecting more signals is only the beginning.
An input needs provenance, relevant context, an accountable owner, and a route through the factory. Repeated reports need consolidation. Conflicting requests need resolution. An urgent incident needs a different admission policy from a speculative feature idea.
The factory also needs to record when it declines work, defers it, or asks for more evidence. Otherwise, an automated intake system can create an ever-growing queue faster than delivery can clear it.
That result should update the next decision. It should also improve future estimates of impact and confidence. This is how a factory starts learning from its work.
Integrations and process setup are part of the factory
“Managed inputs” covers a substantial integration layer. Signals arrive through several categories, each with its own context and operating requirements:
- Product analytics: adoption, conversion, retention, and experiment results from tools such as Amplitude and PostHog.
- System logs and observability: errors, alerts, and diagnostic context from platforms such as Datadog.
- Input forms: structured customer reports, employee improvement requests, and proposed experiments, with fields that capture the problem, affected users, and intended outcome.
- Chats and messaging: requests and reports through Slack, Microsoft Teams, WhatsApp, or a Telegram bot, where a conversation often needs clarification before it becomes actionable.
These are source categories; the feedback loops describe how the organisation uses their evidence. A customer conversation can inform Product strategy, while an analytics finding can expose an operational problem. The factory must preserve that context when routing an input.
From scattered signals to actionable work
Integrations bring in evidence. The operating process turns it into a qualified input.
A connector supplies access to a source. Making that source useful also requires mapping its data to the factory's context, associating signals with the relevant customer or service, consolidating duplicates, assigning owners, and deciding what evidence permits a change to proceed. The process also needs to return a decision and outcome to the person or channel that raised the problem.
Connecting the customer's systems
My view is that platform success depends on resourcing this work explicitly. Forward-deployed implementation engineers should work alongside customers to connect systems, configure workflows, validate the first complete feedback loop, and turn repeatable patterns into reusable integrations.
Configuring the organisation's factory
Internal AI Software Factory engineers drive implementation inside the organisation. Their role is similar to internal SAP consultants configuring an ERP system: they translate the organisation's processes into an operating setup. They connect internal systems, map data and context, configure the change taxonomy and workflow routing, implement guardrails and escalation paths, and maintain that configuration as the organisation changes. Product, engineering, and operations owners supply the priorities, constraints, and accountability that these engineers make executable.
The implementation plan should fund both integration engineering and process design, with ongoing ownership for connector health and workflow changes. These costs also belong in the factory's economics: setup costs can be allocated over a declared period, while maintenance and internal operating effort remain recurring costs.
A practical starting point is one source, one class of change, one accountable owner, and one measured outcome. For example, a recurring error from Datadog enters a defined repair workflow, the factory delivers a validated fix, and follow-up observation determines whether the problem was resolved. That proves the full loop before the organisation expands the integration footprint.
Change needs an operational process
A change might fix an error, improve onboarding, automate an employee workflow, or test a new product idea. Each has a different purpose and risk profile, but all need a traceable path from an input to an outcome.
In the factory, that path becomes explicit:
AI can perform much of the work along this path. It can investigate an alert, assemble context from customer reports, propose acceptance criteria, modify software, run checks, and prepare a controlled rollout. Humans set the objectives, define decision boundaries, and remain accountable for the result.
The important unit is a coherent change tied to an intended outcome. A pull request is one implementation artifact; a single change might require several of them, plus documentation, configuration, and an experiment.
Every admitted change should therefore carry a few essentials: where it came from, what problem it addresses, what improvement is expected, what evidence would count as success, and who owns the decision after measurement.
For an onboarding change, that could mean reducing abandonment at a particular step, measuring completion for an eligible customer cohort, and checking that support requests do not increase. For a reliability fix, it could mean eliminating a recurring error without degrading latency.
This definition lets the factory evaluate the work it produces. A successful deployment begins the outcome measurement; it does not complete it.
Inside Change: workflows follow input and change type
Change needs a taxonomy and a collection of executable workflows. A bug fix, an optimisation, an internal automation, and a product experiment have different evidence requirements. The source of the input matters too: a diagnostic alert and an ambiguous customer message may describe the same bug, but require different investigation steps.
The factory should select a workflow using input category × change type, then apply the relevant guardrails for the affected service and scope. Several combinations can share a workflow; the taxonomy helps select and configure it without requiring a bespoke process for every combination.
| Input category | Change type | Example workflow |
|---|---|---|
| System logs or alerts | Bug fix | Reproduce, repair, review, and release within approved limits |
| Customer form or chat | Bug fix | Clarify the report, then enter the repair workflow when reproducible |
| Product analytics | Product optimisation | Form a hypothesis, run a controlled experiment, and evaluate outcomes |
| Employee request | Internal automation | Validate the process, test with its users, and roll out within agreed scope |
Practitioner framework—illustrative. These are example routing rules, not a complete taxonomy.
Example: a bug-fixing workflow
Take a recurring application error reported through logs or a customer channel. The workflow begins with triage across complexity, impact, and risk. Complexity asks whether the investigation and repair fit the factory's capabilities. Impact assesses the problem's severity, affected users, and urgency. Risk assesses the consequences of changing the system: the affected services, permissions, data, and reversibility.
A high-impact problem can still have a small, low-risk repair. The workflow should evaluate those dimensions separately.
Automatic delivery needs an explicit stopping path
A bug-fixing workflow progresses only while its scope, risk, and evidence remain eligible.
The workflow has six explicit stages
Decide · prove · fix · review · release · observe
Triage and decide
Decision gateIf complexity or risk exceeds the approved limits, or the evidence is too uncertain, create or update an issue and stop autonomous execution.
Escalation and ownership
Attach the diagnosis, impact assessment, evidence, and reason for escalation. Assign it to an accountable owner.
Prove the failure
EvidenceCreate a regression test that fails on the buggy revision for the reported reason.
Before fix: same test failsAfter fix: same test passesWhat counts as proof
A failed test caused by missing credentials or a broken environment is not evidence that the bug was reproduced. If reproduction is unsuccessful, return to the issue-and-stop path.
Fix and open the PR
ImplementationImplement the repair and link the pull request to the original signal and issue.
Validation checks
The regression test should now pass, alongside the other checks required for the affected code.
Run the review cycle
Review loopReview findings return to the fixing step, followed by another review and check run.
Retry budget and escalation
The cycle has a time or retry budget. Unresolved findings, exhausted budgets, or newly discovered risk route the work to its owner rather than allowing an endless autonomous loop.
Evaluate the release gate
Release gateRequire a clean review, the failing-before/passing-after evidence, and all mandatory checks for the exact revision being released.
Scope and release policy
Reassess the actual patch against the guardrails: triage's initial permission does not authorise an expanded change. If it remains eligible, merge and release automatically under the configured release policy.
Notify and observe
Outcome feedbackAfter a successful release, notify the corresponding customer, service, or operations group—and the original reporter where appropriate—with what changed and links to the issue, PR, test evidence, and release.
Monitoring and recovery
Then observe whether the error and its customer effect improved. If deployment fails, report that failure to the owner; if post-release checks fail, follow the configured rollback or escalation process.
I would describe the release condition as inside an approved low-risk envelope, with no unresolved risks. No software change can be shown to have literally zero risk. The practical objective is to define which changes the organisation is willing to release automatically, and to require evidence that each candidate still meets that policy.
This is also where forward-deployed engineers and internal AI Software Factory engineers work together. They configure the taxonomy, routing and risk thresholds, required tests and reviews, escalation paths, and notification groups with the accountable business and service owners. Like an ERP implementation, the work turns organisational rules into an executable operating process.
The bug-fixing example shows how the factory can handle a complete change autonomously while preserving an explicit stopping path. Its outcome evidence then feeds the system and customer loops; its cost, retries, and elapsed time feed the engineering loop.
Impact creates several feedback loops
My sketch shows several loops returning to Change. They operate at different speeds and answer different questions.
| Feedback loop | Signals and measured outcomes | What it changes next |
|---|---|---|
| System | Alerts, errors, recurring incidents, reliability | Repairs, prevention, and operational safeguards |
| Engineering | Cost per change, engineering cycle time, rework, delivery failures | Factory workflows, tooling, architecture, and capacity |
| Product | Customer insights, adoption, retention, experiment results | Strategy, roadmap, and ideas for further investment |
| Customer | Direct requests, problem reports, and feedback after delivery | Clarified needs, fixes, and candidate improvements |
| Employee | Workflow friction, operational cycle time, manual work, error rates | Internal tools and process improvements |
Practitioner framework—illustrative. Adapted from the author's software factory sketch; these are proposed loops, not measured research results.
A system alert may justify action in minutes. An engineering bottleneck may require weeks of observation. A product hypothesis may need a longer window before its effect is visible. The factory must accommodate those differences.
Customer input also arrives through two distinct routes. A direct report provides a specific problem to investigate. Research, behavioural data, and conversations reveal broader patterns that Product can translate into strategy. Both matter, but a request alone does not establish its priority or prove that the requested solution is right.
Employee feedback deserves the same treatment. If an internal change reduces the elapsed time of an approval process, we should also ask whether it reduced active labour, mistakes, or customer waiting. Faster operations can create value, but elapsed time saved and money saved are different measurements.
The engineering loop has a particular role: it improves the factory itself. Expensive retries, slow reviews, or frequent rollbacks become inputs for changes to the delivery process. The factory should be able to improve both the product and its own ability to produce useful changes.
This emphasis on the surrounding system is consistent with DORA's 2025 research, which describes AI as an amplifier of existing organisational strengths and weaknesses. My proposed factory is one way to make that system explicit.
The economics move towards impact per cost
I expect cost per change to become a central operating metric for AI software delivery—and impact per cost to become its strategic measure of success.
Cost per change tells us how efficiently the factory operates. Its cost base should include staff time, compute, LLM usage, and shared tooling. Staff cost includes investigation, review, supervision, and rework. Compute includes execution and testing; LLM cost includes failed attempts and retries. Shared costs need a consistent allocation rule, and bundled charges should be counted once.
For comparable changes over a defined period:
Operating efficiencyCost per change = total allocated factory cost ÷ completed changes.
The denominator matters. Splitting one improvement into ten tickets should not make the factory appear more productive. Failed attempts and abandoned work remain costs even when they produce no completed change. Comparisons should therefore use consistent change definitions and separate work with substantially different complexity or risk.
Impact per cost asks a further question: what did that spending achieve?
Strategic outcomeImpact per cost = measured outcome improvement ÷ total cost associated with achieving it.
For example, the numerator might be additional completed customer tasks, a reduction in operational errors, or incremental contribution margin. Each needs a baseline, an observation period, and a credible way to distinguish the change's effect from other influences. Controlled rollouts can help; where evidence is weaker, the factory should preserve that uncertainty. The denominator should include delivery and incremental operating costs over the declared period: a cheap implementation can still be expensive to run.
These measures will often belong at the initiative or portfolio level. Several changes can contribute to one outcome, and allocating the same benefit to each would overstate impact. Revenue, reliability, and employee time also have different units; they cannot be added into one universal score without explicit assumptions.
My proposed north star is therefore impact per cost within clearly defined outcome domains, supported by cost per change as an operating measure. Reliability, security, and customer experience remain constraints on that optimisation.
Cycle time and parallel delivery matter alongside the ratios. The factory should measure elapsed time from admitted input to rollout, the time until useful outcome evidence arrives, and the number of changes in flight. More parallel work helps only while review, integration, and measurement can keep up.
The CTO's goal becomes democratising delivery
If AI can execute much of the delivery process, the CTO has an opportunity to make that capability accessible across the organisation.
A support employee should be able to introduce a customer problem. An operations manager should be able to propose a workflow improvement. A product team should be able to initiate an experiment. They need a supported route from their problem to a controlled change.
The CTO's responsibility is to make those routes dependable: providing context, permissions, validation, rollout controls, observability, and appropriate human decision points. Access can widen as the factory demonstrates that it handles particular classes of work reliably.
That gives democratisation a concrete meaning. More people can initiate useful changes, while the factory applies consistent standards and keeps responsibility visible.
Product owns the strategy governing the work
This model also changes where Product spends its attention.
Modern product thinking already emphasises outcomes and discovery. SVPG's product model centres empowered teams and experimentation aimed at achieving outcomes. The factory extends that direction by making more of the routine coordination executable.
My view is that Product should increasingly own the strategy that governs the backlog: which customers and problems matter, which outcomes deserve investment, what evidence is required, and how capacity is divided between established products and new opportunities.
AI can then help maintain the operational queue. It can consolidate requests, assemble evidence, estimate delivery cost, flag dependencies, and propose rankings. RICE offers one familiar starting point: reach, impact, confidence, and effort. Its original guidance also recognises that dependencies and business needs can justify working outside score order.
Automating that work does not make the underlying judgements disappear. Product still defines the goal behind “impact,” challenges confidence estimates, and decides when a strategically important opportunity deserves investment despite weak historical evidence. Confidence should come from customer research and observed outcomes. Scores should retain their evidence and assumptions so those decisions can be examined.
Product's attention can shift towards customer understanding, opportunity discovery, positioning, and investment choices—including cross-sell, up-sell, and entirely new propositions. The factory can execute more of the routine delivery decisions within that strategy.
Established products and innovation need different paths
Improve what already works
For an established product, many improvements have repeatable patterns: repairs, usability refinements, operational automation, and optimisation against known goals. My ambition is to make this work economical enough that maintaining and improving the product consumes less scarce human coordination.
That creates room to invest in uncertain opportunities. It does not imply that established products should stop innovating; customer needs and markets still change.
Learn before scaling something new
New ideas need a discovery path through the factory. A customer problem may lead to a prototype, a proof of concept, a limited rollout, and then a decision to expand, revise, or stop. A proof of concept can establish technical feasibility; a customer experiment is needed to test whether the proposition is useful and viable.
These stages should have explicit questions, budgets, and decision criteria. An early experiment can create value by showing that a larger investment should be avoided. Its immediate commercial impact may be small, while the decision it informs is significant.
This is why a single automated ranking across maintenance, optimisation, and discovery would be inadequate. Product needs to allocate capacity across these purposes, then evaluate work within each using suitable evidence. Otherwise, high-confidence incremental improvements can continually crowd out uncertain opportunities.
Start with one complete feedback loop
The first implementation should connect one useful input to one controlled workflow and one measured outcome. A recurring bug with a reproducible failure is a good candidate for the repair workflow described here.
Give an AI Software Factory engineer responsibility for the setup. Agree the admission criteria, release boundaries, escalation owner, and notification group with the relevant service and business owners. Record the baseline and define what would count as a resolved problem before the workflow starts.
Run the first changes with close human oversight. Use the results to improve routing, tests, and decision rules, then enable automatic release for the class of changes that meets the agreed policy. Track the full cost, elapsed time, rework, and resulting outcome before expanding to another input or change type.
That creates a practical foundation for the ambition in this article: delivery accessible to more people, AI executing more of the work, and Product directing investment through strategy and evidence. The factory earns broader autonomy by showing that its changes produce useful impact at an acceptable cost.