Skip to the content
Software Made Clear Diagrams that show the mechanism About

Who owns the factory

ANSWER

A factory is operated, not installed. Every role it creates — the instruction set, the runners, the escalations, the stop control, its own on-call — needs a name against it, because the ones nobody claims fail during an incident rather than before one.

IN PLAIN TERMS

Think of a new machine bolted to a shop floor. Somebody has to load it, service it, hold the key to its emergency stop and answer for what comes out. Buying it settled none of that, and the night it jams is a poor time to work out whose job it was.

A software factory gets designed as though the machine were the hard part. Choose the orchestrator, write the instruction set, scope the credentials, decide where a person stands at the gate. All of that is real work and none of it runs on its own, because a factory is operated rather than installed — and operating one is a set of jobs that existed in nobody’s team last year. The jobs nobody claims do not surface quietly during a design review. They surface during an incident.

The work a factory creates#

Automation is usually discussed as though it only subtracts: this task goes away, that step stops being typed by hand. It also adds, and the additions are easy to miss because they arrive without a title. Somebody now maintains the shared instruction set every agent reads and decides which proposed edit to it is an improvement. Somebody is woken when the orchestrator stops halfway through a run. Somebody answers the question an agent raises when it will not guess. Somebody decides whether a pilot that went well on one service may go near a second. None of that appeared on anyone’s objectives, and none of it happens by default.

Each one needs a person rather than a department. The platform team owns it is the same statement as nobody owns it, delivered with more confidence: a team cannot be paged, cannot be argued with, and cannot notice that it has left the company. A name can do all three. The useful exercise is to read the rows below and try to fill each with a person you could ring today.

Situation Take Because
Factory product owner Value and roadmap Without a name here the factory becomes whichever automations somebody found interesting, and nobody can say what it is for or when it is finished.
Platform operator Runners and tokens Queued runs drain slowly and nobody notices; a token expires on a Friday; the fleet is available until it is not, and the outage belongs to everyone equally.
Harness owner Rules and evals The shared instruction set grows a sentence per incident, never shrinks, and no edit to it is replayed against already-solved work before the whole fleet reads it.
Service owner, per service One codebase Agent changes land in code someone already answers for at night. Held centrally it becomes a promise about services the central team has never read.
Escalation answerer Product questions An agent stops, asks and gets no reply. Items then sit in a status that looks like progress and is not, which is worse than a visible queue.
Risk and data approver The boundaries Residual risk stays unaccepted, so it stays unowned, so somebody meets it for the first time during an audit or an incident.
Factory on-call The machine itself The orchestrator stops at three in the morning, the engineer paged is on call for the product, has no runbook for the builder and is unsure they may halt it.

Two of those rows resist being tidied into a central function, and both are usually tidied anyway. Service accountability stays where it already sits — an agent’s change lands in code that a team answers for at night, and moving that to whoever runs the factory is a promise about services the factory team has never opened. The last row is not the incident process you already have, either. The incidents a factory processes are faults in the product; the incidents a factory has are its own, and confusing the two is close to universal because it only becomes visible at the moment the second kind happens and the first kind’s rota answers the phone.

The questions with no default answer#

Start with the one that has the shortest useful answer. Who can stop all automation, right now, and can they do it at three in the morning without asking anybody’s permission first? The stop itself is a mechanism, platform-enforced, with more than one setting. A mechanism with no name against it is not a control though — it is a feature somebody built. The name, a phone number that reaches them, and standing authority to use it without convening anyone are what turn one into the other.

Then the questions that get deferred because none of them blocks a demo. Who is on call for the orchestrator itself, on the same terms as any other production system? Who accepts the residual risk, in writing, given that a factory described as having none has not been described honestly and an unaccepted risk is simply one that will surprise somebody later? Who decides when the pilot widens, on what evidence, and who is allowed to say no? And who pays for the model usage, which is rarely the same person as whoever watches the spend week by week — a budget with no watcher is a budget read in arrears.

Underplayed almost everywhere is what happens to the job that already exists. When agents produce most of the changes, the remaining human work shifts towards reading output, answering escalations and keeping the instructions current, and review capacity stops being a background cost and becomes the limit on what the whole system can deliver. That is a different job with different skills: less writing, more judging; less uninterrupted stretch, more arriving queue. Presenting it as the same job with a multiplier attached is how a team ends up doing work nobody agreed to, was never trained for, and is then measured on.

Which raises an ownership question that the measurement side of this can set out but cannot answer: not which measures are safe to report per person, but who decides that, and when. The honest answer to “when” is before anything is switched on rather than after somebody objects — in many jurisdictions tooling that can function as performance monitoring carries obligations that include consulting employee representatives ahead of go-live, and the consultation is not a formality to be completed afterwards. The answer to “who” is the part with a name against it: somebody has to be accountable for that decision, and it cannot be the person who built the dashboard, because they are the last person able to see it from the outside. Late is a legal problem and the quickest way to lose the goodwill the whole programme runs on.

Where ownership goes wrong#

The instruction set with no owner is the slowest failure and the most common. Every incident produces one more sentence, each reasonable on the day it was written. Nothing removes one, because removal is nobody’s task and defending a removal earns nothing. A year on, the files contradict themselves in two places, the agents behave in ways nobody can account for, and an investigation has nowhere to start.

A stop control that has never been used fails in the opposite way — instantly, at the worst moment. It exists, it is documented, and because nobody has ever pulled it, nobody knows whether it works, how long it takes to take effect, or whether they are the person permitted to use it. Under pressure that uncertainty reads as an argument for waiting to see whether the problem gets worse. The repair is dull: pull it deliberately on a quiet afternoon, with the people who might have to pull it for real standing there.

Routing failures are quieter still, by construction. An escalation queue points at a team that was reorganised out of existence eight months ago; nothing errors, no alert fires, and the items wait in a status that resembles progress. Or the factory’s own alert reaches the product rota, so the engineer woken at three has never heard of the orchestrator, has no runbook for it, and is unsure whether stopping it is within their authority — a fifteen-minute problem that takes four hours, of which the first is spent establishing whose problem it is.

The version that passes every audit is the worst of them. Everything is owned by the platform team on paper. In practice it is owned by whoever touched it most recently, and during the fortnight that person is on leave it is owned by no one at all. Ownership recorded without a named person is documentation of an intention.

None of this needs resolving before you start, and saying otherwise would be its own kind of dishonesty. A first advisory agent running on real changes needs a repository, an instruction set and somebody willing to read what it writes; demanding a complete accountability map before that is a way of never beginning. What the roles gate is the step after — widening past the pilot, granting write access, letting a run proceed unattended overnight. The failure mode worth naming is the reverse of the one people fear: not the pilot that fails loudly, but the pilot that succeeds so quietly that it spreads by inertia, and the first time anyone asks who owns the thing is the first time it breaks.

That list of names is the working form of the commitment the whole arrangement rests on — that people stay in control of it. Control is not a posture or a sentence in a policy. It is a set of people, each attached to something that can be pulled, answered, approved or refused, and each able to say what happens next when the machine does something nobody intended.

IF YOU REMEMBER ONE THING

A control with nobody named against it is not a control. Write down the roles the factory creates, put one person on each, and pay attention to the rows you could not fill — those are the ones that fail during an incident.

Questions people also ask

4 QUESTIONS
Who should own an AI software factory?

Split it rather than handing it to one group. One person owns what the factory is for and when it widens. One owns the runners, the credentials and its availability. One owns the shared instruction set and the gate that guards changes to it. Service accountability stays with the teams that already hold it, because agent changes land in code somebody else answers for at night. Every row gets a name, not a department.

Who should be on call when an agent orchestrator fails?

Somebody who knows what the orchestrator is, on the same rota terms as any other production system. This is not the product rota. The incidents a factory processes are faults in the product; the incidents a factory has are its own — a runner pool that will not scale, a credential rotation that broke every workflow, an agent reopening the same change forty times. Paging the product engineer for the second kind costs hours before anyone starts diagnosing.

How does a developer's job change when agents write most of the code?

It moves from producing changes to judging them, answering escalations and maintaining the instructions the agents read. Review capacity then becomes the limit on everything downstream, so the constraint moves from typing speed to reading attention. That is a different job with different skills and a different rhythm — more interruption, less flow — and treating it as the same job with a multiplier attached is how people end up doing work they never agreed to.

What has to be decided before an agent pilot is extended?

Who may stop all automation without asking permission, and can actually reach the control on a Sunday night. Who is on call for the machine itself. Who accepts the residual risk in writing. Who approves the data and security boundary. Who pays for the model usage and who watches the spend weekly. And which measures will never be reported per person — decided with the people affected, before anything is switched on.