Research index · v0.1
Evidence before
prescription.
Loom studies how controls around tool-using AI agents affect both protected constraints and the workflow outcome. Each entry separates the question, method, result, and limitation.
Evidence boundaryControlled public-benchmark research — not customer or production results. Ongoing work and research directions are labeled separately from completed results.
Research items
01Architecture · ongoing
Harness architecture research
- Question
- Which parts of an agent system determine its effective capabilities, evidence, and authority?
- Method
- Source-grounded architecture review across tool-using agent systems, organized around observable workflow boundaries.
- Result
- Ongoing. Findings will be added here only after the evidence and claim boundaries are reviewed.
- Limitation
- This is an active research thread, not a completed comparative study.
02Verification · result
Evidence-aware verification
- Question
- When does a mutation require a separate follow-up verification step?
- Method
- Inspected the pinned benchmark tool contracts and retrospectively reclassified verification obligations using authoritative mutation evidence.
- Result
- All 13 relevant mutation types already returned authoritative updated state or a deterministic success receipt. The blanket follow-up requirement was classified as redundant for this tool surface.
- Limitation
- A bounded result from one benchmark ecosystem. It does not show that follow-up verification is unnecessary in other environments.
- Question
- Can task, tool, and policy evidence identify workflow-specific control hypotheses?
- Method
- A deterministic, evidence-linked scan of 60 pinned public-benchmark tasks across airline and retail workflows.
- Result
- The scan produced 585 task-specific opportunities across 14 policy types. It separated required, optional, redundant, unreachable, and human-decision cases without automatically changing enforcement.
- Limitation
- Opportunity frequency is not customer prevalence, and a candidate control is not evidence that it improves agent behavior.
- Question
- Does authoritative approval enforcement reduce unapproved actions without unacceptable workflow degradation?
- Method
- A preregistered, controlled 135-run comparison on nine pinned public-benchmark tasks, holding models, tools, evaluator, and task state fixed across arms.
- Result
- Enforcement reduced the observed unapproved-mutation rate from 9.76% to 2.22%, while task success fell from 75.6% to 64.4%. The trade-off failed the experiment’s predefined non-inferiority criterion.
- Limitation
- One benchmark, two customer-support domains, nine tasks, and a specific enforcement-and-repair design. This is not customer or production evidence.
- Question
- How should the quality of a ship, change, or stop recommendation be evaluated?
- Method
- Define falsifiable decision criteria that preserve task outcomes and protected constraints while exposing uncertainty and evidence gaps.
- Result
- Direction only. No completed experimental result is claimed yet.
- Limitation
- The evaluation contract and benchmark design remain to be frozen.
Have a workflow that could challenge this research?
We’re looking for concrete release decisions, failures, and control trade-offs—not polished success stories.
Talk to us about your agent workflow