The AI Agent With Too Many Permissions
Ask most teams what their AI agent can do, and you'll get an answer about what it's supposed to do. Ask what it's actually authorized to do — every tool, every scope, every credential it can reach — and the answer is usually "we'd have to check." That gap is the finding. Here's why agents end up over-privileged, how to test for it, and what a practical least-privilege model looks like.
The permission an agent uses vs. the permission it has
A support agent that answers billing questions needs read access to invoices. In most deployments we've assessed, it also has write access to the customer database, because that's the same service account the engineering team uses for three other internal tools, and splitting it out felt like unnecessary work at the time. The agent has never once used that write access. It has always had it.
That's the core problem, and it isn't unique to AI agents — it's the same over-provisioning pattern that's plagued IAM for decades, standing access granted for convenience and never revisited. What's different with agents is the trigger for exploitation. A human employee with excess access is a slow-burning insider-threat or credential-theft risk. An AI agent with excess access is one crafted prompt, one poisoned document, one malicious tool description away from using that access right now — because the agent doesn't need to intend the misuse. It only needs to be convinced.
Why agents accumulate permissions by default
- Provisioned once, for every future task. Developers scope access to "everything this agent might plausibly need," not "everything this specific task needs," because re-requesting access per task is friction nobody wants to build.
- No approval workflow, no offboarding trigger. A human's access request goes through a manager; their departure triggers deprovisioning. An agent's access is a config value one engineer can widen in a pull request, and there's rarely an equivalent event that prompts anyone to narrow it back down.
- Shared credentials across tools. The same API key or service account often backs multiple integrations because it's already there and already working. Scoping one agent's access means touching everything else built on that credential.
- Agency creeps forward, not back. Every iteration adds one more tool, one more scope, one more "just in case" permission to unblock a new feature. Nobody schedules a pass to remove the ones the agent stopped needing.
- Testing that never looks for it. A conventional pentest checks whether the agent's endpoints are secure. It rarely asks "if this agent were manipulated, what's the worst plausible action it could take with what it's already allowed to do?" — which is precisely the question excessive-agency testing exists to answer.
Where this shows up in the OWASP taxonomy
Excessive Agency is a named category in the OWASP Top 10 for Agentic Applications (ASI Top 10) — an agent granted more autonomy, tool access or permission scope than its task requires, such that a successful manipulation causes damage far beyond the intended use case. It sits alongside, and compounds, Identity and Privilege Abuse (an agent acting under an identity with broader rights than the acting agent should have) and Agent Goal Hijack (an attacker redirecting what the agent is trying to do). Goal hijack is the trigger. Excessive agency is what turns a redirected agent into a serious incident instead of a contained one.
The relationship matters for prioritization: you can't fully eliminate goal-hijack risk — prompt injection defenses are still an active research problem — but you can substantially cap the damage a successful hijack does, by making sure the agent never had the excess permission to begin with. Least privilege is a blast-radius control, not a prevention control, and it's the one that's fully in your hands regardless of how good your injection defenses get.
How we test for it
- Build the actual permission inventory. Every tool, API scope, database role, file-system path and credential the agent's runtime identity can reach — not what the system prompt claims it's limited to. This is usually the first surprise: the discovered surface is wider than what anyone believed was granted.
- Map each permission to a real task. For every item in that inventory, identify the specific, legitimate task that requires it. Anything without a clear answer is a candidate for removal before any adversarial testing even starts.
- Attempt direct misuse. Ask the agent, straightforwardly, to perform an out-of-scope action using access it holds but shouldn't exercise for this task. No injection, no trickery — just testing whether the authorization boundary exists at all or only exists in the system prompt's wording.
- Attempt manipulated misuse. Chain the same excess-permission actions behind prompt injection, poisoned tool descriptions, or content the agent reads from an untrusted source — testing whether a real attacker could reach the same outcome indirectly.
- Test chained authorization, not just single calls. Individually-authorized tool calls can combine into an unauthorized outcome — read a private document, then send its contents somewhere external, each call legitimate on its own. This is where sequence-level testing catches what a permissions audit alone misses.
- Verify approval gates actually gate. Where a human-in-the-loop step exists for high-risk actions, confirm it can't be bypassed, auto-approved, or presented with misleading context that makes a reviewer rubber-stamp something they wouldn't approve if they understood it.
A practical least-privilege model
| Control | What it does | Stops |
|---|---|---|
| Task-scoped tool grants | Grant tool access per task or session, not standing access to the agent's full identity | An unrelated task exploiting access it never needed |
| Ephemeral, narrowly-scoped credentials | Short-lived tokens minted for one operation, scoped to one resource, expiring fast | A leaked or misused credential remaining useful after the task ends |
| Read/write separation by default | Read access is the default grant; write and delete require explicit justification | An agent that only ever reads data being able to silently modify or destroy it |
| Human approval for consequential actions | Irreversible, high-value or externally-visible actions pause for explicit human sign-off, with the actual intended action shown clearly | A hijacked goal executing a damaging action with no checkpoint |
| Per-agent identity, not shared service accounts | Each agent gets its own credential rather than reusing one shared across tools | One agent's compromise inheriting every other integration's access |
| Scheduled privilege review | A recurring pass that re-justifies every standing grant against current task requirements | Agency creep — permissions that made sense once and were never revisited |
| Action-level audit logging | Every tool call logged with the triggering input, not just the outcome | An excessive-agency incident going undetected because nothing was recorded |
"You can't fully stop an agent from being manipulated. You can absolutely stop a manipulated agent from mattering."
Least agency, not just least privilege
Traditional least privilege asks "what access does this identity need?" For agents, that framing is necessary but incomplete — the sharper question is least agency: what's the minimum autonomy this agent needs to act without a human, for this task, right now? An agent can hold perfectly scoped permissions and still be over-privileged in effect, if it's allowed to chain several of those permissions together autonomously into an outcome no single grant would have permitted on its own. Scoping access and scoping autonomy are two different controls, and a mature agent security program needs both.
Where this fits in a broader engagement
Permission and privilege testing is one part of a full AI/LLM penetration test, alongside prompt injection resistance, data leakage and the rest of the Agentic AI Security surface. For agents built on the Model Context Protocol specifically, the same excess-permission question applies at the server and tool-definition layer — see our MCP server pentesting methodology for how tool-level authorization gets tested there. ComplyArmor maps every finding to the OWASP ASI Top 10, with human verification before anything reaches your report — see our full methodology.
Frequently asked questions
What is least privilege for AI agents?
It's the principle that an AI agent should hold only the tools, data access and credentials strictly required for the task it's performing right now, for only as long as that task takes — not standing, always-on access to everything it might conceivably need across every task it could ever run.
Why do AI agents end up over-privileged?
Agents are usually provisioned once, early, by a developer trying to unblock every future task rather than scope the current one — and unlike a human employee, an agent's "access request" is a single config change with no manager approval step, no offboarding trigger, and often no review afterward. Broad access is the path of least resistance, and it's rarely revisited once things work.
What is excessive agency in the OWASP sense?
Excessive Agency is a named risk category in the OWASP Top 10 for Agentic Applications — an agent granted more autonomy, tool access or permission scope than its task requires, so a successful manipulation (prompt injection, goal hijack) can cause damage far beyond what the intended use case should allow.
How do you test whether an AI agent is over-privileged?
Enumerate every tool, scope and credential the agent can reach, map each one against what its actual task set requires, and then attempt to use the excess access adversarially — via direct requests and via injected instructions the agent might follow. If an out-of-task action succeeds without a human in the loop, that's a finding, independent of whether any injection was needed to trigger it.
Does human-in-the-loop approval fully solve excessive agency?
No, but it substantially reduces blast radius for the highest-risk actions. It only works if the approval step shows the human what the agent actually intends to do in a form they'll genuinely read, and if it's applied to consequential actions specifically rather than everything, which trains reviewers to click approve without looking.
If you can't currently list every permission your production agents hold without going and checking, that's the finding — before any adversarial testing even starts.