AlkhemyAgent Custody

The Agent Failure Diagnostic · free

Your agent said it was done. It was not done.

Every agent failure lives in one of four layers, and almost every fix gets aimed at the wrong one. This diagnostic maps seven symptoms to the layer that actually owns each, and the wrong instinct each one triggers.

The diagnostic unlocks on this page the moment you submit. No spam, no sequence, one email if the cohort opens a new run.

Or score yourself first

Every number below is a real, logged incident. Every one of these systems reported success while it was broken.

16,781 to 1

Failures to successes over 70 days. A working fallback answered every query, so nothing ever looked broken.

18 hours

A dispatch pipeline dark behind a dashboard that stayed green the entire time.

177,777

Password attempts against three servers in one week, behind a default nobody had checked.

The diagnostic

Seven symptoms, the layer that owns each, the instinct to resist

Harness, loop, and graph engineering are taught everywhere now, mostly free. Custody, who is accountable for the output between the agent claiming success and it reaching production, is the layer nobody teaches, and it is the one that decides whether you can trust the result.

Use it the day something fails: find the symptom, read the layer, and notice the fix you were about to reach for.

Where you are

Seven questions about last week, then your level

Answer for what you actually did, not what you meant to do. No email, nothing stored, about ninety seconds.

0 of 7
Question 1 of 7. Last week, what did a typical request to your agent contain?

The three signals scored here are Anthropic's, from research on ~400,000 interactive sessions from ~235,000 people between October 2025 and April 2026: how precisely the user frames their directions, what they ask Claude to verify, and whether the user tends to correct Claude or Claude tends to correct the user. Read the research.

The five rung names are ours, and this is not their classifier. Theirs read real sessions. This reads seven answers you gave about your own week, so treat it as a mirror rather than a measurement. Anthropic state their own limit: We cannot measure real-world outcomes, like whether code written in a session is actually used or discarded thereafter, or whether it produces an economically valuable artifact.

One finding worth the whole study if you are not an engineer: Every one of the ten largest occupations in our dataset lands within seven points of software engineers in terms of their success. This is method, not credential.

This table is the last ten minutes of Agent Custody, a four hour live cohort that builds the layer underneath it: written law, permissions that fail closed, orchestration that cannot collide, and verification that disbelieves the summary.

Request a seat

Opens an email with the subject filled in. Ten seats per run.