CAPABILITY · OSWORLD TASK SUCCESS
The technical objection is gone. Agents now succeed on real computer tasks two out of three times, up from roughly one in ten just two years ago.
Stanford HAI, 2026 AI Index Report
Self watches what your agents actually do, joins it to human corrections and real business outcomes, and enforces how much autonomy each skill has earned.
WHY NOW
CAPABILITY · OSWORLD TASK SUCCESS
The technical objection is gone. Agents now succeed on real computer tasks two out of three times, up from roughly one in ten just two years ago.
Stanford HAI, 2026 AI Index Report
DEPLOYMENT · ENTERPRISE APPS WITH AGENTS
Agents are forecast to arrive inside the software you already run within the year, whether or not anyone signs off on governing them.
Gartner forecast, August 2025
THE GAP · GOVERNANCE FORECAST
Gartner's own diagnosis: enterprises treat governance as binary — locked down or fully trusted — and that is the root cause of failure.
Gartner, May 2026
THE STATE OF THE ART
Enterprises running agents that move money, change records, or resolve exceptions almost always pick one of these. The third column is the whole product.
Throughput capped
A person checks every action. Volume stops at whatever the review team can absorb — and once the queue outruns them, approvals get faster, not better. The record still looks clean, because nobody is really reading it.
Blind after day one
One evaluation, then standing trust. Nothing in that assessment survives a model release, a changed tool, a rewritten prompt, or an input the agent never saw during testing — and nothing tells you when it stopped being true.
Continuous · skill-specific · revocable
Oversight is set per skill, against evidence, and moves in both directions. A model release puts the affected skills back under full review until they re-earn what they had. Trust accrues slowly. It is revoked at the gate, immediately.
THE METHOD
Self runs the same loop for every skill, and doesn't wait for it to run on its own. Every number on every screen is one click from the sessions that produced it.
01 · INPUTS
What the agent did — including scenarios Self injects itself, to see what it hasn't naturally encountered yet.
What a person approved, edited, escalated, or reversed.
What turned out to be correct, costly, or harmful downstream.
How much autonomy your organization permits this skill to earn.
02 · BEHAVIORAL LEDGER
Append-only. Reference-only storage for sensitive payloads — Self keeps record identifiers, not your invoices.
03 · EVALUATION
04 · GATE
AUTONOMY TIERS
Trust is calculated per skill, per agent version, per environment — never as one score for an agent. Select a tier to see what it means operationally and what Self has to see before it will let a skill operate there unattended.
TIER 4
TIER 3
TIER 2
TIER 1
TIER 0
THE EVIDENCE
This is: one decision, the evidence that produced it, and what changed at the gate because of it. It is what your auditor reads.
Accounts Payable Agent · Production · 60-day window
60 days · production only
Engagement-adjusted, not raw approval rate
Actions joined to a final business result
Critical failures 0
None unresolved
Material errors 0.3% · 6 of 1,840
All corrected before payment
Calibration 2 confidently wrong
Both self-flagged, corrected same day
Drift None detected
Inputs, tool use, escalation rate, calibration
Environment changes None since last review
Model, prompt, tool, workflow, policy
PROPOSED OPERATING BOUNDARY
New-vendor invoices and any invoice over $2,000 still route to a human.
REQUIRED AUDIT SAMPLE
15% of Tier 3 actions
REEVALUATE
In 30 days, or immediately on the next model or prompt change.
TRIGGERING SESSIONS
1,840 sessions, each traceable to the ledger row that produced it.
TRUST OVER TIME
Every action class starts under full review. Authority widens as evidence accumulates — and contracts immediately the moment it shouldn't have.
Every action starts under full review, and a person sees each one before it happens.
THE PILOT