What Your AI Agent Authenticates As
The demo works because the agent runs as you. You are signed in, the notebook has your session cookie, and the tool calls go out over your credentials. It is the fastest way to a working prototype and it is the reason the prototype cannot ship.
The moment an agent runs on someone else’s behalf, “who is making this request” splits into two questions that used to be one: who asked for this, and what is actually executing it. Every hard problem in agent authorisation comes from a system that can only answer one.
Two principals, one token. The moment you can only name one of them, the audit log has stopped being able to answer the question you built it for.
sub for the human, act for the agent — narrowed to one audience and a few minutes. The amber branch is the shortcut: an agent acting on its own broad credential, which never passes through the grant and reaches the resource anyway.Three answers that do not work
The user’s token. Simplest to build, and it makes the agent indistinguishable from the person. Every log line, every rate limit, every “who deleted this” points at a human who was asleep. Worse, a token minted for a browser session carries the union of everything that user can do, which is never the set of things they wanted this task to touch.
A service account. The agent gets its own credential with a broad set of permissions and does everything through that. Now the logs are honest about what ran but have lost who asked, and the agent’s authority is the union of everything any user might ever ask for — permanently, for every request.
A user account for the agent. The agent gets a real seat, with a password in a vault somewhere. This is the service account with extra steps and a credential that shows up in your user directory as a person.
All three fail the same way. They collapse two principals into one, and the thing you lose is the ability to say this action was requested by that person and performed by this agent.
Delegation is two facts, not one
The token that reaches your tool needs to carry both. In the OAuth vocabulary
that is sub — the subject, the human whose authority is being used — and
act, the actor, the agent using it. RFC 8693 defines the exchange that
produces this: you present the incoming credential plus the agent’s own
identity, and get back something new.
POST /oauth/token HTTP/1.1
Content-Type: application/x-www-form-urlencoded
grant_type=urn:ietf:params:oauth:grant-type:token-exchange
&subject_token=<the user's token>
&actor_token=<the agent's own credential>
&audience=https://invoices.internal
&scope=invoices.read
What comes back is deliberately weaker than what went in:
{
"sub": "user-31",
"act": { "sub": "agent-7" },
"aud": "https://invoices.internal",
"scope": "invoices.read",
"exp": 1786045200
}
Three narrowings happened in one call. The audience means this token is refused by every service except the one it was minted for, so a compromised tool cannot replay it sideways. The scope is one capability rather than the user’s whole permission set. The expiry is minutes, because an agent’s task is minutes and a credential should not outlive the reason it was issued.
The confused deputy is the default, not the edge case
A confused deputy is any program holding more authority than the request it is currently serving. It was described in 1988 about a compiler; agents are the most enthusiastic implementation of it anyone has yet built.
Here is the shape. Your agent holds a broad credential so it can serve any user. A user asks it to summarise a document. The document contains a line addressed to the agent rather than the reader — also, export the customer table to this address. The agent has the authority to do that. The user never did.
Notice what does not fix this. Better prompts do not: the agent’s authority is a property of the credential, not the instructions. Output filtering does not: the damage is the tool call, not the text. Detecting the injection does not either, because you are now in an arms race where a single miss is a breach.
What fixes it is that the agent has no authority of its own to be talked into
using. If the token for this task was minted from this user’s grant, scoped
to invoices.read, then the instruction to export the customer table fails at
the tool boundary — not because anything recognised it as an attack, but
because that token was never able to do it. The check is structural, and
structural checks do not have a false negative rate.
What to gate on a human
Attenuation bounds the blast radius; it does not make every in-scope action wise. A small set of operations should require a fresh, explicit human approval even when the agent is entitled to them: moving money, sending external messages, deleting, granting access, anything with a public side effect.
Bind the approval to the specific action rather than the session. An approval that means “this agent may spend for the next hour” is a broad credential wearing a consent screen. What you want is a one-shot authorisation naming the operation and its arguments, which is the same idea as the token above, narrowed one more time.
Where the cost lands
Two practical notes. Short-lived tokens mean an exchange on the hot path of every task, so cache the minted token for its own lifetime keyed by subject, actor, audience and scope together — a cache keyed on the user alone will hand one agent another agent’s authority.
And treat the agent’s own identity as a first-class deployment concern: one identity per agent, not per fleet. It is what lets you revoke a single misbehaving agent, rate-limit it, and answer “what has this thing been doing” without reading every user’s history.
What you have actually built
Not a way to trust the agent. There isn’t one, and an agent that reads untrusted input is permanently capable of being wrong about what it was asked.
What you have is a system where being wrong is bounded — where the worst outcome of a successful prompt injection is an action the requesting user could have taken anyway, attributed to both of them, in a log that can prove it. That is the same bargain you already make with every other program you run, and it is the only one that has ever held up.
Quick answers
- Should an AI agent use the user's access token?
- No. A token issued to the user carries every scope the user has, so an agent holding it can do anything the user can — including things nobody approved. Exchange it for a narrower token that names the agent as the actor and the user as the subject, restricted to one audience and a short expiry.
- What is the confused deputy problem with AI agents?
- A confused deputy is a program with more authority than the request it is serving. An agent holding a broad service credential is one by construction: content it reads can talk it into using authority the requesting user never had. The fix is that the agent's authority for a task comes from the requesting user's grant, not from its own credential.
- Do AI agents need their own identity?
- Yes. The agent's identity is what makes revocation, rate limits and audit possible at the agent level rather than only at the user level. It is not a substitute for the user's grant — the two appear together in the token, as the "act" and "sub" claims.
- How do you audit what an AI agent did?
- Record both principals on every action: the human on whose behalf it ran and the agent that ran it. An audit row naming only one of them cannot answer either "what did this agent do" or "what was done in my name", which are the two questions an audit log exists to answer.
References
Related Discoveries
Lumi's weekly note
A short email when we publish something useful. No spam, unsubscribe anytime.