Field Notes: A Gateway policy can't see the user if your agent calls with its own token
A re:Post question asked how to stop an AgentCore agent acting beyond the user who asked. I built a demo to find out. A Cedar policy does nothing if the agent calls with its own token. Two gotchas, and the real limit: a stale token claim.
On this page
Series Navigation
- 1. Field Notes: Three things I learned diagnosing a production Bedrock workload
- 2. Field Notes: Turning prompt caching on for a production Bedrock workload
- 3. Every dashboard was green while the agent burned six figures a year
- 4. Field Notes: The AgentCore Memory write that returns success and reads back empty
- 5. The LLM is not a security boundary
- 6. Your LLM security diagram defends the wrong layer
- 7. Field Notes: A Gateway policy can't see the user if your agent calls with its own token ← you are here
Someone on AWS re:Post asked a question I had been meaning to test. Their ops agent, built on Strands, opens pull requests against a GitOps repository. The agent’s own credential can open a PR for anyone who talks to it, including people who are not allowed to push to that repo. Prompt instructions do not help, because the call runs with the agent’s credential whatever the model decides. How do you stop it? (the thread, 24 September 2026, no accepted answer when I read it.)
The only reply so far said Policy in Amazon Bedrock AgentCore should handle it. That is half right, and the half that is wrong is the interesting part. So I built a small repo to find out, and this is what it showed. To be clear about what this is: a built and tested demo against a mock target, not something that happened in production.
The short version:
- A Cedar policy on the Gateway does nothing if the agent calls with its own token. The Gateway sees the agent, not the user.
- Forward the asking user’s token and the same policy blocks the same request, before your tool is ever invoked.
- Policy in AgentCore decides on the token’s claims and the tool arguments only. It cannot ask a downstream system whether the person still has access right now.
I have written the “enforce it in your own code” version of this before, in The LLM is not a security boundary. This is the managed-service version, and where it stops.
The setup
One tool, open_pr, behind two AgentCore Gateways. Both use the same mock Lambda that records the call to DynamoDB and never touches GitHub. One Gateway has no policy. The other has a Cedar policy engine in ENFORCE mode. Cognito issues the tokens.
flowchart LR
U["Bob<br/>not on the platform team"] -->|asks| A[Agent]
A -->|"1. agent's own token"| G1["Open Gateway<br/>no policy"]
A -->|"2. agent's own token"| G2["Guarded Gateway<br/>Cedar policy"]
A -->|"3. forwards Bob's token"| G2
G1 --> T[("open_pr target")]
G2 -->|"permit, else default deny"| T
The “agent” is a deterministic script standing in for whatever the model chooses to call. That is deliberate. Enforcement must not depend on what the model decides, so I took the model out of the loop.
Two repos are in play, both made up: acme/docs (anyone signed in may open a PR) and acme/platform-infra (platform team only). Any other repo has no permit, so it is denied by default. Policy in AgentCore is default-deny, and a matching forbid beats any permit (Cedar semantics in the AgentCore docs, read 30 September 2026).
Cases 1 and 2: the agent’s own credential
| Case | Result |
|---|---|
Open Gateway, agent’s own token, Bob asks for platform-infra | PR opened |
Guarded Gateway, agent’s own token, Bob asks for platform-infra | PR opened |
The second row is the point. I attached a working policy and it changed nothing, because the policy evaluates the principal on the token, and the token was the agent’s. Bob’s name only travelled as a requested_by argument, which the policy has no reason to trust and mine did not read.
If you have put a policy engine on a Gateway and your agent calls with its own service credential, you have added a check that always sees the same caller.
Case 3: forward the asking user’s token
The fix is one change in the agent: send the asking user’s access token to the Gateway instead of your own.
headers = {"Authorization": f"Bearer {users_access_token}"}
The Gateway builds the Cedar request from that token. The JWT sub becomes the principal, the other claims become tags on it, and the tool arguments arrive as context.input (authorization flow, read 30 September 2026). My policy for the protected repo:
permit(
principal is AgentCore::OAuthUser,
action == AgentCore::Action::"github___open_pr",
resource == AgentCore::Gateway::"<guarded gateway arn>"
) when {
context.input.repo == "acme/platform-infra" &&
principal.hasTag("cognito:groups") &&
principal.getTag("cognito:groups") like "*\"platform\"*"
};
| Case | Result |
|---|---|
Bob’s token, platform-infra | Denied, target never invoked |
Alice’s token (platform team), platform-infra | Allowed |
Alice’s token, payments | Denied (no permit) |
Bob’s token, docs | Allowed |
The denial comes back as JSON-RPC error -32002, “Tool call not allowed due to policy enforcement”, and the DynamoDB ledger shows zero new rows. The Lambda was never called. The check happened before your code.
Two things that bit me
1. A bare like on a group name matches more than you meant. My first policy said like "*platform*". It worked on the first run, which is the dangerous kind of working. The cognito:groups claim reaches Cedar as a string that looks like ["platform","oncall"], quotes and brackets included. I tested a user in a group called platform-readonly, and they were allowed onto platform-infra. The fix is to match the quoted element, like "*\"platform\"*". I then checked that platform-readonly is denied and that a user in platform plus oncall is allowed. One other published AgentCore RBAC example (3 Steps to RBAC for AI Agents on AgentCore) uses a bare like on the scope claim. That is the same class of pattern. I did not test scopes, so I am not claiming it breaks.
2. The Gateway’s role needs its permissions before the Gateway exists. With CDK, the role’s permissions are a separate IAM policy resource. CloudFormation created the guarded Gateway first, and the CloudFormation create failed with “Access denied while calling GetPolicyEngine” (the message I saw; the docs describe an InternalServerException when you attach an engine to an existing Gateway, so your symptom may differ). The role needs GetPolicyEngine, AuthorizeAction and PartiallyAuthorizeActions (permissions docs, read 30 September 2026). Give the Gateway an explicit dependency on the role’s policy.
One more, from the docs rather than from a failure: creating a Cedar policy validates it against the Gateway it names, so the role that creates policies also needs InvokeGateway on that Gateway (same page). I deployed as an admin and never hit it, so I have not seen the failure myself.
The limit: a stale claim
Policy in AgentCore knows what the token says. Nothing else.
I added Bob to the platform team, took his token, then removed him from the group. His old token still carried the platform claim, and the guarded Gateway allowed the PR. A fresh token for Bob was denied. Between those two calls the only thing that changed was how old the token was. In the demo the access token lasts an hour, and that is the window.
What closes that gap is not in the Gateway. Short token lifetimes narrow it. Revocation narrows it further, though I did not test Cognito’s revocation options here. For anything high risk, the tool should check the downstream system itself at the moment of the action, which is what the asker in the re:Post thread built as an open-source check-before-acting library. For a low-risk action, an hour of stale entitlement may be an acceptable trade. For “merge to the infrastructure repo” it is not, and that is a judgement for your threat model, not a setting.
What I did not test
- The agent is a script. I did not run an LLM, because the point is that enforcement does not depend on it.
- AgentCore documents the caller’s identity propagating across a Gateway to Runtime to Gateway chain for temporal policy sessions, inside one account and Region (policy sessions and identity propagation). I did not build that path.
- Where the agent gets the user’s token. I forwarded a raw Cognito access token. The same docs page mentions an on-behalf-of (OBO) flow for a shared application credential with per-user
subclaims. I did not test AgentCore Identity token exchange or OBO. - Temporal (session) policies, and any production-scale behaviour.
- The Cognito service user stands in for the agent’s credential. Cognito machine-to-machine tokens carry no group claim, so a real deployment would model the agent identity differently.
Other people have written up forwarding the user’s JWT to Policy in AgentCore, and AWS has its own material on it. What I could not find written up was the side-by-side that shows the policy is inert when the agent uses its own credential, the group-matching collision, or the stale-claim case. I read the ones I could reach and could not open one AWS Builder Center article, so treat that as “I could not find it”, not “nobody has”.
Run it yourself
The repo is rajmurugan01/agentcore-acts-as-the-user. make deploy, make demo, make destroy. It runs every case above and prints PASS or FAIL for each; the last run was 10 of 10. The demo users and password are throwaway.
On cost: AWS lists Gateway at $5 per million invocations and Policy at $25 per million authorisations (AgentCore pricing, read 30 September 2026). A full demo run makes ten tool calls, so my estimate for the AgentCore side is well under a cent. I did not price the small Cognito, DynamoDB and Lambda charges around it. That is an estimate from list prices, not a measured bill.
The question I would like to hear answers to: for the actions your agents take, where is the line between “a token that is up to an hour stale is fine” and “the tool must ask the downstream system every time”?
Related reading
Introducing the AI Architect Roadmap
Eight rungs mapping what shipping production Agentic AI on AWS actually takes, marked honestly as covered, strong, flagship or gap rather than filled in to look finished.
Not a Python tutorial: the patterns that bite in production agents
The senior-lens Python patterns that actually bite in a production agent loop: why boto3 blocking the event loop is the first one, and four more that follow from it.
What a Year 10 study system taught me about production AI failure modes
A personal Bedrock-adjacent build that went through three iterations and an architecture pivot. Five lessons that map directly to production AWS AI work.
Newsletter
A new AWS and AI engineering write-up every Tuesday, directly to your inbox.
Comments
- Loading comments…