Too little control
Early-stage research
Assurance for AI agents doing real work.
Loom examines the workflow around an agent—what it can do, what its tools can prove, and where its authority ends. Teams use that evidence to choose controls and test a change before release.
- Focus
- Consequential tool use
- Method
- Bounded comparisons
- Status
- Controlled research
Selected What can the team observe?
01 / Release pressure
Every added control changes the workflow.
Agents can update records, issue refunds, write code, contact users, and operate internal systems. Teams surround them with layers of instructions, permissions, approvals, evaluators, and retries.
- Prompts
- Tools
- Permissions
- Approvals
- Evaluators
- Retries
- Guardrails
Too much control
Friction rises. Task success falls.
For any given workflow, the job is to decide which controls are necessary, what evidence supports them, and whether the agent can still finish the task.
02 / The full pipeline
Turn a workflow into a testable control decision.
Move through the pipeline to see where Loom looks for evidence. Each stage narrows the question without hiding the trade-off.
Find missing controls, likely redundancies, and gaps in the available evidence.
Control diagnosis
Find missing controls, likely redundancies, and gaps in the available evidence.
03 / Working method
One bounded change at a time.
Loom tests one proposed control change against the original workflow. The comparison keeps task completion and protected constraints in view.
Understand
Map the task, tools, authority, protected constraints, and the evidence each action returns.
Diagnose
Find missing controls, likely redundancies, evidence gaps, and the failure mechanism worth testing.
Evaluate
Compare one bounded change with the original workflow on task success and assurance outcomes.
04 / Current evidence
A control improved one assurance metric and damaged task completion.
Hard approval enforcement reduced the observed unapproved-mutation rate. Task success fell by 11.1 percentage points against approval guidance, which failed the experiment’s predefined criteria.
The result: change the enforcement design before release.
Read the method and limitation05 / Where Loom fits
The decision before a harness change.
Teams already have systems for building agents, inspecting runs, measuring behavior, and enforcing rules.
Agent frameworks
Build and run the agent.
Observability
Record what happened.
Evaluations
Measure how the agent performed.
Policy engines
Enforce a rule the team already knows.
Which control is justified by this workflow’s evidence, and how will the team know the change helped?
06 / Technical thesis
How much harness does this workflow need?
The Minimal Effective Harness is the smallest supported combination of capabilities, controls, observations, and workflow structure for this agent and this task.
Evidence — What can the system prove?
07 / Who we want to learn from
Teams releasing agents that take consequential actions.
We want to understand how teams release agents that update records, send messages, change code, or trigger financial and operational actions.
Customer discovery
Bring us one concrete workflow.
Tell us where the release decision became difficult and what evidence the team had at the time.
Talk to us about your agent workflow