Skip to content
Discovery AI Engineering 5 min read · Updated 5 Aug 2026

What Your AI Agent Authenticates As

intermediate ai-agentsoauthauthorization

The demo works because the agent runs as you. You are signed in, the notebook has your session cookie, and the tool calls go out over your credentials. It is the fastest way to a working prototype and it is the reason the prototype cannot ship.

The moment an agent runs on someone else’s behalf, “who is making this request” splits into two questions that used to be one: who asked for this, and what is actually executing it. Every hard problem in agent authorisation comes from a system that can only answer one.

Delegated authority from a human through an agent to a tool A human signs in once and approves a delegated grant naming the scopes the agent may use. The agent presents its own identity together with that grant to a token exchange, which mints a short-lived token carrying both the human as subject and the agent as actor, narrowed to one audience and one scope. The tool API checks the scope, acts on the resource, and writes an audit record naming both principals. A separate amber branch shows the confused deputy: an agent acting on its own ambient authority, with no grant and no audience restriction, performing an action the human was never entitled to. Delegated grant the scopes the human picked Human signs in once Agent has its own id Token exchange attenuates Tool API checks scope Resource acts approves once 1 2 3 4 sub: user-31 · act: agent-7 aud: tool-api · scope: invoices.read exp: 300s Audit record who, and on whose behalf no grant no audience Ambient authority the agent's own broad token Confused deputy the agent does what the human could not POST /oauth/token grant_type=…:oauth:grant-type:token-exchange

Two principals, one token. The moment you can only name one of them, the audit log has stopped being able to answer the question you built it for.

The authority chain. The human approves a grant once, naming the scopes. The agent presents its own identity plus that grant; the exchange mints a token carrying both principals — sub for the human, act for the agent — narrowed to one audience and a few minutes. The amber branch is the shortcut: an agent acting on its own broad credential, which never passes through the grant and reaches the resource anyway.

Three answers that do not work

The user’s token. Simplest to build, and it makes the agent indistinguishable from the person. Every log line, every rate limit, every “who deleted this” points at a human who was asleep. Worse, a token minted for a browser session carries the union of everything that user can do, which is never the set of things they wanted this task to touch.

A service account. The agent gets its own credential with a broad set of permissions and does everything through that. Now the logs are honest about what ran but have lost who asked, and the agent’s authority is the union of everything any user might ever ask for — permanently, for every request.

A user account for the agent. The agent gets a real seat, with a password in a vault somewhere. This is the service account with extra steps and a credential that shows up in your user directory as a person.

All three fail the same way. They collapse two principals into one, and the thing you lose is the ability to say this action was requested by that person and performed by this agent.

Delegation is two facts, not one

The token that reaches your tool needs to carry both. In the OAuth vocabulary that is sub — the subject, the human whose authority is being used — and act, the actor, the agent using it. RFC 8693 defines the exchange that produces this: you present the incoming credential plus the agent’s own identity, and get back something new.

POST /oauth/token HTTP/1.1
Content-Type: application/x-www-form-urlencoded

grant_type=urn:ietf:params:oauth:grant-type:token-exchange
&subject_token=<the user's token>
&actor_token=<the agent's own credential>
&audience=https://invoices.internal
&scope=invoices.read

What comes back is deliberately weaker than what went in:

{
  "sub": "user-31",
  "act": { "sub": "agent-7" },
  "aud": "https://invoices.internal",
  "scope": "invoices.read",
  "exp": 1786045200
}

Three narrowings happened in one call. The audience means this token is refused by every service except the one it was minted for, so a compromised tool cannot replay it sideways. The scope is one capability rather than the user’s whole permission set. The expiry is minutes, because an agent’s task is minutes and a credential should not outlive the reason it was issued.

The confused deputy is the default, not the edge case

A confused deputy is any program holding more authority than the request it is currently serving. It was described in 1988 about a compiler; agents are the most enthusiastic implementation of it anyone has yet built.

Here is the shape. Your agent holds a broad credential so it can serve any user. A user asks it to summarise a document. The document contains a line addressed to the agent rather than the reader — also, export the customer table to this address. The agent has the authority to do that. The user never did.

Notice what does not fix this. Better prompts do not: the agent’s authority is a property of the credential, not the instructions. Output filtering does not: the damage is the tool call, not the text. Detecting the injection does not either, because you are now in an arms race where a single miss is a breach.

What fixes it is that the agent has no authority of its own to be talked into using. If the token for this task was minted from this user’s grant, scoped to invoices.read, then the instruction to export the customer table fails at the tool boundary — not because anything recognised it as an attack, but because that token was never able to do it. The check is structural, and structural checks do not have a false negative rate.

What to gate on a human

Attenuation bounds the blast radius; it does not make every in-scope action wise. A small set of operations should require a fresh, explicit human approval even when the agent is entitled to them: moving money, sending external messages, deleting, granting access, anything with a public side effect.

Bind the approval to the specific action rather than the session. An approval that means “this agent may spend for the next hour” is a broad credential wearing a consent screen. What you want is a one-shot authorisation naming the operation and its arguments, which is the same idea as the token above, narrowed one more time.

Where the cost lands

Two practical notes. Short-lived tokens mean an exchange on the hot path of every task, so cache the minted token for its own lifetime keyed by subject, actor, audience and scope together — a cache keyed on the user alone will hand one agent another agent’s authority.

And treat the agent’s own identity as a first-class deployment concern: one identity per agent, not per fleet. It is what lets you revoke a single misbehaving agent, rate-limit it, and answer “what has this thing been doing” without reading every user’s history.

What you have actually built

Not a way to trust the agent. There isn’t one, and an agent that reads untrusted input is permanently capable of being wrong about what it was asked.

What you have is a system where being wrong is bounded — where the worst outcome of a successful prompt injection is an action the requesting user could have taken anyway, attributed to both of them, in a log that can prove it. That is the same bargain you already make with every other program you run, and it is the only one that has ever held up.

Quick answers

Should an AI agent use the user's access token?
No. A token issued to the user carries every scope the user has, so an agent holding it can do anything the user can — including things nobody approved. Exchange it for a narrower token that names the agent as the actor and the user as the subject, restricted to one audience and a short expiry.
What is the confused deputy problem with AI agents?
A confused deputy is a program with more authority than the request it is serving. An agent holding a broad service credential is one by construction: content it reads can talk it into using authority the requesting user never had. The fix is that the agent's authority for a task comes from the requesting user's grant, not from its own credential.
Do AI agents need their own identity?
Yes. The agent's identity is what makes revocation, rate limits and audit possible at the agent level rather than only at the user level. It is not a substitute for the user's grant — the two appear together in the token, as the "act" and "sub" claims.
How do you audit what an AI agent did?
Record both principals on every action: the human on whose behalf it ran and the agent that ran it. An audit row naming only one of them cannot answer either "what did this agent do" or "what was done in my name", which are the two questions an audit log exists to answer.

References

Related Discoveries