Early-stage research

Assurance for AI agents doing real work.

Loom examines the workflow around an agent—what it can do, what its tools can prove, and where its authority ends. Teams use that evidence to choose controls and test a change before release.

Focus
Consequential tool use
Method
Bounded comparisons
Status
Controlled research

Selected What can the team observe?

01 / Release pressure

Every added control changes the workflow.

Agents can update records, issue refunds, write code, contact users, and operate internal systems. Teams surround them with layers of instructions, permissions, approvals, evaluators, and retries.

−

Too little control

Unsafe or unreliable execution.

+

Too much control

Friction rises. Task success falls.

For any given workflow, the job is to decide which controls are necessary, what evidence supports them, and whether the agent can still finish the task.

02 / The full pipeline

Turn a workflow into a testable control decision.

Move through the pipeline to see where Loom looks for evidence. Each stage narrows the question without hiding the trade-off.

Find missing controls, likely redundancies, and gaps in the available evidence.

Stage 04 of 7

Control diagnosis

Find missing controls, likely redundancies, and gaps in the available evidence.

03 / Working method

One bounded change at a time.

Loom tests one proposed control change against the original workflow. The comparison keeps task completion and protected constraints in view.

01

Understand

Map the task, tools, authority, protected constraints, and the evidence each action returns.

02

Diagnose

Find missing controls, likely redundancies, evidence gaps, and the failure mechanism worth testing.

03

Evaluate

Compare one bounded change with the original workflow on task success and assurance outcomes.

Controlled public-benchmark researchNot customer or production results

04 / Current evidence

A control improved one assurance metric and damaged task completion.

Hard approval enforcement reduced the observed unapproved-mutation rate. Task success fell by 11.1 percentage points against approval guidance, which failed the experiment’s predefined criteria.

The result: change the enforcement design before release.

Read the method and limitation
Same benchmarkOnly the approval rule changed.
Task completionHigher is better−11.1 pp
Actions without approvalLower is better−7.54 pp
When approval was suggested, task success was 75.6 percent and 9.76 percent of mutations were unapproved. When approval was required, task success was 64.4 percent and 2.22 percent of mutations were unapproved.

05 / Where Loom fits

The decision before a harness change.

Teams already have systems for building agents, inspecting runs, measuring behavior, and enforcing rules.

01

Agent frameworks

Build and run the agent.

02

Observability

Record what happened.

03

Evaluations

Measure how the agent performed.

04

Policy engines

Enforce a rule the team already knows.

Loom

Which control is justified by this workflow’s evidence, and how will the team know the change helped?

06 / Technical thesis

How much harness does this workflow need?

The Minimal Effective Harness is the smallest supported combination of capabilities, controls, observations, and workflow structure for this agent and this task.

Evidence testIs it required?Can we observe it?Can we test it?
Minimal Effective HarnessSmallest supported set of controls and capabilities for this workflow.

Evidence — What can the system prove?

Outcome, capabilities, authority, evidence, and workflow pass through the required, observable, and testable criteria to define the Minimal Effective Harness.

07 / Who we want to learn from

Teams releasing agents that take consequential actions.

We want to understand how teams release agents that update records, send messages, change code, or trigger financial and operational actions.

Customer discovery

Bring us one concrete workflow.

Tell us where the release decision became difficult and what evidence the team had at the time.

Talk to us about your agent workflow