Research register · v0.1

Evidence before prescription.

Each Loom study begins with a concrete control question. The method, finding, and boundary stay together so a benchmark result is not mistaken for production proof.

Completed findings
3
Active study
1
Research direction
1
01QuestionWhat should change?
02MethodWhat stays fixed?
03ResultWhat did we observe?
04BoundaryWhat remains unknown?
Evidence boundaryControlled public-benchmark research

No customer or production result is claimed. Active work and research directions are visibly separated from completed findings.

How a study moves

The limitation travels with the result.

A control is only useful if the team can explain why it belongs in this workflow and what changed when it was introduced.

  1. 01

    Observe the workflow

    Name the outcome, authority, tools, and evidence that already exists.

  2. 02

    Form one hypothesis

    Identify a missing, redundant, or harmful control worth testing.

  3. 03

    Hold the task fixed

    Compare one bounded change against the original workflow.

  4. 04

    Keep the boundary

    Publish the result with its limitation and evidence source attached.

Current register

Five questions under study.

Architecture active

Harness architecture research

QuestionWhich parts of an agent system determine its effective capabilities, evidence, and authority?

Method
Source-grounded architecture review across tool-using agent systems, organized around observable workflow boundaries.
Current state
Ongoing. Findings will be added here only after the evidence and claim boundaries are reviewed.
Boundary
This is an active research thread, not a completed comparative study.

Verification completed

Evidence-aware verification

QuestionWhen does a mutation require a separate follow-up verification step?

Method
Inspected the pinned benchmark tool contracts and retrospectively reclassified verification obligations using authoritative mutation evidence.
Finding
All 13 relevant mutation types already returned authoritative updated state or a deterministic success receipt. The blanket follow-up requirement was classified as redundant for this tool surface.
Boundary
A bounded result from one benchmark ecosystem. It does not show that follow-up verification is unnecessary in other environments.

Control discovery completed

Policy opportunity discovery

QuestionCan task, tool, and policy evidence identify workflow-specific control hypotheses?

Method
A deterministic, evidence-linked scan of 60 pinned public-benchmark tasks across airline and retail workflows.
Finding
The scan produced 585 task-specific opportunities across 14 policy types. It separated required, optional, redundant, unreachable, and human-decision cases without automatically changing enforcement.
Boundary
Opportunity frequency is not customer prevalence, and a candidate control is not evidence that it improves agent behavior.

Approval assurance completed

Approval enforcement experiment

QuestionDoes authoritative approval enforcement reduce unapproved actions without unacceptable workflow degradation?

Method
A preregistered, controlled 135-run comparison on nine pinned public-benchmark tasks, holding models, tools, evaluator, and task state fixed across arms.
Finding
Enforcement reduced the observed unapproved-mutation rate from 9.76% to 2.22%, while task success fell from 75.6% to 64.4%. The trade-off failed the experiment’s predefined non-inferiority criterion.
Boundary
One benchmark, two customer-support domains, nine tasks, and a specific enforcement-and-repair design. This is not customer or production evidence.

Research direction proposed

Loom Decision Quality

QuestionHow should the quality of a ship, change, or stop recommendation be evaluated?

Method
Define falsifiable decision criteria that preserve task outcomes and protected constraints while exposing uncertainty and evidence gaps.
Current state
Direction only. No completed experimental result is claimed yet.
Boundary
The evaluation contract and benchmark design remain to be frozen.

From benchmark to workflow

Bring us one difficult release decision.

Share the workflow, the control under debate, and the evidence your team trusted.

Talk to us about your agent workflow