AI Factory

An agent earns the right to move a change forward by leaving a receipt.

An AI Factory is the system a coding agent runs inside: every stage names the artifact it produces and the evidence it leaves behind. Yours goes live in 4 weeks, ending with one real change landing through it unattended.

Request a 15-minute callNo repository access needed for the first conversation.
Anthropic PartnerBuilding software since 201140+ engineers

What a factory is made of

Seven stages, and each one leaves evidence

Trust is not granted to the agent. Every stage below names the artifact it produces, the evidence that proves it happened, and the person accountable for it.

StageWhat it producesReceiptAccountable
IntentA written intent and a spec, in your repositoryApproval committed to version control before any code existsProduct owner
PlanAn implementation plan, reviewed before work startsThe plan is committed, so the diff can later be compared against itEngineer
BuildThe change, its tests and its logsEvidence bound to the exact base and candidate commitBack to the loop, nothing is pushed
VerifyThe check that decides whether the work is correctBreak the behavior on purpose and only its own tests fail, nothing elseA human reads the diff
LandThe update to your trunkA protected check on the exact commit, outside the agent authorityThe human becomes the checker
ReleaseThe deploy or the feature flagWhoever lands a change cannot release it aloneRelease manager
OperateMonitoring, and a diagnosis when a band is breachedThe landed commit is watched, and an ordinary revert restores a green buildService owner

AI Factory

What earns a change the right to land without a human

Scored on 11 pillars, from build and testing through to agent tooling and interfaces and multi-agent coordination.

Loopone change

Gather context, act, check, repeat until a condition.

Harnessone session

Sandbox, tools, memory between runs, and the gate that defines done.

Factorya stream

Many harnessed loops, a queue, an enforced update protocol, a human owning the whole.

Name the level before you tune the model. Nine gates decide the landing event, which is the trunk update, not the merge.

G1 Scope

Passes when

One routine bounded task with a checkable completion condition

Otherwise

Split it or give it to a human

G2 Maturity

Passes when

Repository at L3 or above on the applicable pillars

Otherwise

Human landing required

G3 Risk

Passes when

Not auth, billing, a schema or data migration, or a public contract

Otherwise

Leave the lights on

G4 Oracle

Passes when

The check is cheap, frequent and unfakeable at once

Otherwise

A human reads the diff

G5 Harness

Passes when

Restricted tools, isolated candidate state, recorded base SHA

Otherwise

Do not run unattended

G6 Evidence

Passes when

Diff, checks, logs and explanation bound to base and candidate SHA

Otherwise

Back to the inner loop, do not push

G7 Landing check

Passes when

A protected check with an external source of truth on the exact SHA

Otherwise

The human becomes the checker

G8 Exposure authority

Passes when

Whoever lands on trunk cannot alone deploy or flip a flag

Otherwise

Treat the landing as a release

G9 Outcome and recovery

Passes when

The landed SHA is observed; failure stops the line and reverts to green

Otherwise

No unattended landing

The change does not land unattended.

Nine gates, nine named fallbacks. Not one of them is a percentage.

What this model does not measure

Comprehension debt: the gap between how much code exists and how much anyone understands. A dark factory accrues it as fast as it can, with the suite green throughout. The score does not cover it, so track it separately before widening unattended work.

The ceiling is local, not a book number

Cap unattended landings at the size three quarters of your merged changes already stay under. Across three repositories of one team that came out at 5, 6 and 12 files. One number would have been wrong for all three.

A loop's right to run unattended comes from properties of the oracle, not from the quality of the model.

L3 Standardized is necessary and not sufficient. Autonomy is granted to a class of task, never to a repository.

The build

Your factory goes live in 4 weeks

There are two separate timelines. This one builds the factory. Shipping features through it runs on its own schedule afterward.

Week 1

Measure and build one honest check

We measure the repository against the readiness model and build a single check that survives a negative control: break the behavior on purpose and only its own tests fail.

Week 2

Contain the agent and set the ceiling

Restricted tools, isolated candidate state, a recorded base commit. The change-size ceiling comes from your own merge history rather than from a policy number.

Week 3

Wire the gates and rehearse the revert

The gates below go in, and the team runs the revert drill until restoring a green build is routine rather than an incident.

Week 4

Prove it on one real change

One real change from your backlog lands through the factory unattended, with the receipts from every stage attached to it.

Who builds it

Forward Deployed Engineers, inside your team

The factory is built by senior engineers embedded in your repositories, your standups and your review queue, not handed over as a document. That delivery model, the pod behind it and the engagement tracks are on the Forward Deployed Engineering page.

See how the pod works

Questions

What an engineering leader asks before letting an agent near the repository

What can an agent do on its own in our repository?

It has no route to push to your trunk. Every change arrives as a pull request behind your existing branch protection, and the principal that lands a change cannot release it. Autonomy is granted to a class of task, never to a repository, and each class is bounded by a size ceiling taken from your own merge history rather than a policy number.

How do you stop the agent setup itself from drifting?

We version the configuration that controls agent behavior and test any change to it against an evaluation set built from your own recent work. Relevant failures become regression cases before autonomy expands.

How long does the factory take to stand up?

There are two separate timelines. The factory itself goes live in 4 weeks. Shipping features through it runs on its own schedule after that. The 4 weeks assume the agreed class of change already has a reliable check and the applicable pillars meet L3, which is what gate G2 requires. If either condition is missing, we scope remediation first.

How do you protect our code from LLM exposure?

We use enterprise endpoints with documented data-use and retention terms, or deploy inside your environment. We do not opt client code into model training, and the data boundary and retention settings are agreed before access is granted.

Can you work on our own infrastructure and models?

Yes. We deploy against your cloud, your VPC, and your model endpoints, including open-weight models hosted on your own infrastructure via vLLM.

Who owns the code and IP?

You own the code, prompts, evals, runbooks and project artifacts we create for you from the first commit. Pre-existing components, open-source libraries and third-party models stay under their existing licenses, with usage rights set out in the agreement.

Bring us one class of change

Tell us which routine change you want agents to handle. Fifteen minutes is enough to decide whether an assessment of the repository is the right next step, and no access is needed to have that conversation.

Get in touch

Start with one problem worth solving

Tell us the outcome you need and the systems involved. In 15 minutes we can say whether there is a useful first step, or why there is not. No pitch either way.

Email us directly

sales@3alica.com