AI Factory
An agent earns the right to move a change forward by leaving a receipt.
An AI Factory is the system a coding agent runs inside: every stage names the artifact it produces and the evidence it leaves behind. Yours goes live in 4 weeks, ending with one real change landing through it unattended.
What a factory is made of
Seven stages, and each one leaves evidence
Trust is not granted to the agent. Every stage below names the artifact it produces, the evidence that proves it happened, and the person accountable for it.
| Stage | What it produces | Receipt | Accountable |
|---|---|---|---|
| Intent | A written intent and a spec, in your repository | Approval committed to version control before any code exists | Product owner |
| Plan | An implementation plan, reviewed before work starts | The plan is committed, so the diff can later be compared against it | Engineer |
| Build | The change, its tests and its logs | Evidence bound to the exact base and candidate commit | Back to the loop, nothing is pushed |
| Verify | The check that decides whether the work is correct | Break the behavior on purpose and only its own tests fail, nothing else | A human reads the diff |
| Land | The update to your trunk | A protected check on the exact commit, outside the agent authority | The human becomes the checker |
| Release | The deploy or the feature flag | Whoever lands a change cannot release it alone | Release manager |
| Operate | Monitoring, and a diagnosis when a band is breached | The landed commit is watched, and an ordinary revert restores a green build | Service owner |
AI Factory
What earns a change the right to land without a human
Scored on 11 pillars, from build and testing through to agent tooling and interfaces and multi-agent coordination.
Loopone change
Gather context, act, check, repeat until a condition.
Harnessone session
Sandbox, tools, memory between runs, and the gate that defines done.
Factorya stream
Many harnessed loops, a queue, an enforced update protocol, a human owning the whole.
Name the level before you tune the model. Nine gates decide the landing event, which is the trunk update, not the merge.
G1 Scope
Passes when
One routine bounded task with a checkable completion condition
Otherwise
Split it or give it to a human
G2 Maturity
Passes when
Repository at L3 or above on the applicable pillars
Otherwise
Human landing required
G3 Risk
Passes when
Not auth, billing, a schema or data migration, or a public contract
Otherwise
Leave the lights on
G4 Oracle
Passes when
The check is cheap, frequent and unfakeable at once
Otherwise
A human reads the diff
G5 Harness
Passes when
Restricted tools, isolated candidate state, recorded base SHA
Otherwise
Do not run unattended
G6 Evidence
Passes when
Diff, checks, logs and explanation bound to base and candidate SHA
Otherwise
Back to the inner loop, do not push
G7 Landing check
Passes when
A protected check with an external source of truth on the exact SHA
Otherwise
The human becomes the checker
G8 Exposure authority
Passes when
Whoever lands on trunk cannot alone deploy or flip a flag
Otherwise
Treat the landing as a release
G9 Outcome and recovery
Passes when
The landed SHA is observed; failure stops the line and reverts to green
Otherwise
No unattended landing
The change does not land unattended.
Nine gates, nine named fallbacks. Not one of them is a percentage.
What this model does not measure
Comprehension debt: the gap between how much code exists and how much anyone understands. A dark factory accrues it as fast as it can, with the suite green throughout. The score does not cover it, so track it separately before widening unattended work.
The ceiling is local, not a book number
Cap unattended landings at the size three quarters of your merged changes already stay under. Across three repositories of one team that came out at 5, 6 and 12 files. One number would have been wrong for all three.
A loop's right to run unattended comes from properties of the oracle, not from the quality of the model.
L3 Standardized is necessary and not sufficient. Autonomy is granted to a class of task, never to a repository.
The build
Your factory goes live in 4 weeks
There are two separate timelines. This one builds the factory. Shipping features through it runs on its own schedule afterward.
Week 1
Measure and build one honest check
We measure the repository against the readiness model and build a single check that survives a negative control: break the behavior on purpose and only its own tests fail.
Week 2
Contain the agent and set the ceiling
Restricted tools, isolated candidate state, a recorded base commit. The change-size ceiling comes from your own merge history rather than from a policy number.
Week 3
Wire the gates and rehearse the revert
The gates below go in, and the team runs the revert drill until restoring a green build is routine rather than an incident.
Week 4
Prove it on one real change
One real change from your backlog lands through the factory unattended, with the receipts from every stage attached to it.
Who builds it
Forward Deployed Engineers, inside your team
The factory is built by senior engineers embedded in your repositories, your standups and your review queue, not handed over as a document. That delivery model, the pod behind it and the engagement tracks are on the Forward Deployed Engineering page.
See how the pod worksQuestions
What an engineering leader asks before letting an agent near the repository
What can an agent do on its own in our repository?
It has no route to push to your trunk. Every change arrives as a pull request behind your existing branch protection, and the principal that lands a change cannot release it. Autonomy is granted to a class of task, never to a repository, and each class is bounded by a size ceiling taken from your own merge history rather than a policy number.
How do you stop the agent setup itself from drifting?
We version the configuration that controls agent behavior and test any change to it against an evaluation set built from your own recent work. Relevant failures become regression cases before autonomy expands.
How long does the factory take to stand up?
There are two separate timelines. The factory itself goes live in 4 weeks. Shipping features through it runs on its own schedule after that. The 4 weeks assume the agreed class of change already has a reliable check and the applicable pillars meet L3, which is what gate G2 requires. If either condition is missing, we scope remediation first.
How do you protect our code from LLM exposure?
We use enterprise endpoints with documented data-use and retention terms, or deploy inside your environment. We do not opt client code into model training, and the data boundary and retention settings are agreed before access is granted.
Can you work on our own infrastructure and models?
Yes. We deploy against your cloud, your VPC, and your model endpoints, including open-weight models hosted on your own infrastructure via vLLM.
Who owns the code and IP?
You own the code, prompts, evals, runbooks and project artifacts we create for you from the first commit. Pre-existing components, open-source libraries and third-party models stay under their existing licenses, with usage rights set out in the agreement.
Bring us one class of change
Tell us which routine change you want agents to handle. Fifteen minutes is enough to decide whether an assessment of the repository is the right next step, and no access is needed to have that conversation.
Get in touch
Start with one problem worth solving
Tell us the outcome you need and the systems involved. In 15 minutes we can say whether there is a useful first step, or why there is not. No pitch either way.
Email us directly
sales@3alica.com