PUBLIC-EVIDENCE AI SECURITY CASE STUDY

Entra Agent ID Administrator Scope Gap: Agent Roles Reaching Service Principals

Silverfort reported that Microsoft Entra's Agent ID Administrator role — documented as scoped to agent-related objects — could add an owner to non-agent service principals in the same tenant. Once owner, the researcher could add credentials to the target principal and then authenticate as it; where that principal held privileged directory roles or high-impact Microsoft Graph permissions, the result was privilege escalation. Silverfort also noted that the Entra UI did not present the role as privileged even though the documentation described it that way, which can reduce the review applied to its assignment. Microsoft fixed the issue across clouds by 2026-04-09, blocking Agent ID Administrator from managing owners on non-agent service principals. The assessment below scores only the pre-fix configuration.

Agentic / MCP SystemTier 1Indicative ERS 3.8 (1.2–6.3)Evidence source date 2026-04-29

Silverfort reported that Microsoft Entra's Agent ID Administrator role — documented as scoped to agent-related objects — could add an owner to non-agent service principals in the same tenant. Once owner, the researcher could add credentials to the target principal and then authenticate as it; where that principal held privileged directory roles or high-impact Microsoft Graph permissions, the result was privilege escalation. Silverfort also noted that the Entra UI did not present the role as privileged even though the documentation described it that way, which can reduce the review applied to its assignment. Microsoft fixed the issue across clouds by 2026-04-09, blocking Agent ID Administrator from managing owners on non-agent service principals. The assessment below scores only the pre-fix configuration.

ASSESSED SYSTEM

The Microsoft Entra Agent ID agent-identity control plane as it stood before Microsoft's fix completed across clouds on 2026-04-09: specifically the Agent ID Administrator directory role, the authorization enforcement applied to its actions over Entra application and service-principal objects, and the credential-issuance and ownership operations reachable through Microsoft Graph. This is the identity, credential and authorization substrate on which AI agent identities in a tenant are built, which is the surface AITBM's Cn-5 (Agent Identity Integrity) is defined over.

OUT OF SCOPE

The post-fix Entra configuration (Microsoft blocked the role from managing owners on non-agent service principals across clouds by 2026-04-09); the rest of Entra ID outside the agent-identity role and the service-principal objects it could reach; the AI agents and models that authenticate using these identities, their prompts, tools and behaviour; and any specific customer tenant's role assignments. This assessment scores a point-in-time platform configuration reported by an external researcher, not Microsoft's identity platform generally.

Architecture: Agentic / MCP System (decision tree Q2) — Q1 = NO: the brief evidences no two-or-more agents coordinating, delegating, sharing memory or exchanging messages — Entra Agent ID issues identities, it does not orchestrate agents. Q2 = YES: Entra Agent ID exists to give AI agents (blueprints, agent identities, agent users) first-class workload identities that authenticate and invoke Microsoft Graph and connected tools without explicit human approval for each call, and the assessed component is the permission and credential boundary of that agentic platform. The brief's own framing — 'agent identities still sit on familiar Entra application and service-principal objects' — is the evidence that the agentic identity plane and the general service-principal plane share primitives. Evidence artefacts available: Silverfort's described role behaviour, Microsoft's role documentation as reported, and the fix scope. Tier 1: Tier 1 Critical: enterprise identity infrastructure is infrastructure control in the spec's sense — takeover of a service principal holding privileged directory roles or high-impact Graph permissions yields immediate, at-scale unauthorised access across every workload that authenticates to the tenant; a Tier 3 reading (organisational boundary, administrator users only) is defensible on the user population but understates the consequence test the tier definition actually applies.

Documented attack or failure path

  1. A principal holds the Entra Agent ID Administrator role — an administrator-assignable directory role that the Entra UI did not label as privileged.
  2. Using that role, add an owner to a non-agent service principal: an object class outside the role's documented resource boundary, but not outside its enforced one.
  3. As owner of the target service principal, add a new client secret or certificate credential to it.
  4. Authenticate as the target service principal using the attacker-added credential.
  5. Inherit that principal's directory roles and Microsoft Graph permissions; where those are privileged, an agent-management role has become a tenant privilege-escalation path.

Observed controls and bounded outcomes

Positive credit is given only where the record directly demonstrates a control operating. Recommended or merely presumed controls receive no positive scoring credit.

  • Microsoft shipped a fix across all clouds by 2026-04-09 that blocks Agent ID Administrator from managing owners on non-agent service principals, so the documented boundary is now the enforced boundary.
  • The path was not reachable by an unauthenticated or unprivileged actor: it required an existing assignment of an administrative directory role, so the failure is escalation between administrative scopes rather than an open door.
  • The response guidance recommends correlating role assignment, owner change, credential creation, and sign-in. That guidance does not prove each event was recorded with the necessary identity fields in the assessed configuration, so it receives no positive scoring credit.
  • Microsoft's written documentation did classify the role as privileged; the mismatch Silverfort reported was with the Entra UI's risk labelling, not with the guidance itself.

Layer 1 — Intrinsic Vulnerability Profile

Each sub-metric is placed on its five-level rubric by the evidence quoted beside it. Missing applicable evidence remains unknown. The displayed midpoint and interval are scenario values, not inferred control performance.

Sub-metricScoreRubric basisEvidence
Robustness (Ro) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Fairness (Fa) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Transparency (Tr) — scenario interval 0.05–0.85 (midpoint 0.45), Tier 1 MVT 0.55 indeterminate
Tr-4Model Lineage Disclosure0.25w 0.20Placed at 0.25 rather than 0.50 under the conflicting-evidence rule. Documentation of the effective permission surface existed for the major component (0.50 language: documented but not linked to evaluation evidence), but it materially misstated the enforced boundary, which is worse than a stale-update gap; the lower of the two plausible anchors is taken. Scored against the disclosed authorisation/configuration surface of the agent-identity plane, not model lineage, which is out of scope here.'The role said "agent objects," but enforcement reached broader service principals. Agent ID Administrator was documented as scoped to agent-related objects, yet Silverfort showed it could add an owner to non-agent service principals.' Additionally 'the role was documented as privileged, but the Entra UI did not show it that way.' The brief's closing recommendation is to build a pre-adoption test harness exercising each new agent role against every inherited directory object type — i.e. no such validation evidence was linked to the published scope.source: brief
Privacy (Pr) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Containment (Cn) — scenario interval 0.10–0.69 (midpoint 0.40), Tier 1 MVT 0.65 indeterminate
Cn-1Scope Enforcement0.25w 0.120.25 anchor: basic boundaries exist and are real — the role is scoped RBAC, not unrestricted directory administration — but enforcement was inconsistent, and straightforward parameter manipulation (pointing the same operation at a different object subtype) exceeded the intended scope. Not 0.00, because scope boundaries did exist and constrained most operations.The documented resource boundary of the Agent ID Administrator role was not the enforced one: a direct Graph owner-assignment call against a non-agent service principal succeeded. The brief states the core failure as 'an authorization policy that did not enforce the documented resource boundary' and specifies the correct behaviour as 'Agent ID Administrator may update owners on agent-backed service principals, but must fail closed on non-agent service principals.'source: brief
Cn-2Escalation Prevention0.25w 0.160.25 anchor: basic role checks existed (the actor needed the Agent ID Administrator assignment) but were vulnerable to role confusion — the role's identity as an 'agent' role did not match the object classes it could act on — and the two-step owner-then-credential chain proceeded without an approval gate or execution-time re-check.'Ownership became a takeover primitive. After becoming owner, the researcher could add credentials and authenticate as the target principal. If that principal held privileged directory roles or high-impact Graph permissions, the result became privilege escalation.' The brief notes both sensitive steps — become owner, then add credentials — completed without re-authorisation at execution time, which is why it maps Continuous Authorization Verification (anti-TOCTOU) to the case.source: brief
Cn-5Agent Identity Integrity0.25w 0.14Disputed between the 0.50 anchor (token-based identity with scoped permissions and partial verification, no cryptographic binding to workload or session — a fair description of service-principal identity in general) and 0.25. The lower value is taken per the conflicting-evidence rule: on the demonstrated path, identity spoofing succeeded outright because possession of a credential the attacker could mint was sufficient to be the principal, with no workload attestation and no agent-to-agent verification. disputedAgent identities are built on Entra application and service-principal objects. An actor who could add itself as owner could then mint a new secret or certificate on the target principal and authenticate as it — a full identity assumption, with no attestation binding the credential to an actual workload. The brief's remediation set is credential eviction at the authoritative issuer plus revocation of delegated grants, because 'Removing the attacker-added secret alone is insufficient if the principal can mint replacement access.'source: researcher

Unknown, not N/A: 19 applicable sub-metrics lack admissible public evidence. The lower, midpoint, and upper scenarios evaluate each at 0.00, 0.50, and 1.00 respectively; no weight is redistributed merely because evidence is missing.

Layer 2 — Operational Risk Posture

DimensionScoreJustification
AaAutonomy Amplification0.75w 0.35Supervised autonomy. Two facets: the agent identities this plane issues authenticate and call Microsoft Graph on their own without per-call human approval — that is the purpose of Entra Agent ID — and the control plane itself applies high-consequence changes (owner assignment, credential creation) immediately on a valid API call, with human oversight only asynchronous through audit review and access reviews. Not 1.00: each administrative change is still initiated by a human-held principal, and there is no evidence of the plane autonomously deciding to grant authority.
AsAttack Surface Exposure0.50w 0.25Internet-reachable with authentication and validation. The Microsoft Graph control plane is globally reachable and multi-tenant, but the demonstrated path required an authenticated principal already holding a specific administrative role, and the brief evidences no untrusted-content ingestion, external agent communication or MCP tool integration into the assessed component. More exposed than an internal-only plane (0.25), well short of the maximum-exposure anchor.
CpCascade Potential1.00w 0.25No System Dependency Graph is published for the assessed configuration, so the spec's worst-case default applies — but the reconstruction independently triggers the 1.00 anchor on its merits. The observed path is ungated and terminates at a credential- and permission-issuing node (P4): owning a service principal permits minting credentials for it, and the target class explicitly includes principals holding privileged directory roles and high-impact Graph permissions. That is 'an ungated path reaches a P3/P4 node', with privilege amplification from an agent-scoped administrative origin to arbitrary service-principal authority.GDCP: corroborated by the observed path elevated
RfRemediation Feasibility0.00w 0.15Deterministic fix, already shipped. Microsoft closed the path across clouds by 2026-04-09 with an authorization-policy change blocking the role from managing owners on non-agent service principals — a server-side code/policy correction requiring no retraining and no customer-side model change. Tenant-side follow-up (inventorying role assignments, auditing recent owner and credential changes) is operational hygiene rather than remediation difficulty.

Nelevated = 1 (dimensions strictly above 0.75) → CRM = 1.00.

Layer 3 — Public-evidence confidence diagnostic

ComponentScoreBasis
Pc — Public provenance evidence0.25Minimal provenance in the evidence available. The vendor, platform, role name, affected object classes, exploitation steps and fix date are documented at a high level by the researcher and the vendor's fix notice. There is no AIBOM, no agent or tool inventory, no identity-policy artefact, no per-tenant configuration record, and no cryptographic or independently reviewed evidence tied to a specific assessed deployment.
Ec — Public evaluation coverage0.17coverage 0.17 (4 of 23 applicable sub-metrics) × independence 1.00 × fidelity 1.00. No Full, Standard, or Lite pathway is claimed for a retrospective article.
Tf — Public-evidence freshness0.01Evidence dated 2026-04-23; age 112 days on the workpaper reference date. Components: T_containment 0.01 · T_calendar 0.08 · C_event 0.35 · C_monitor 0.65 · C_evidence 0.85. Binding term: T_containment. dt_days = 112, measured from Silverfort's primary disclosure dated 2026-04-23 to the 2026-08-13 evidence reference date; the AIDEFEND brief republished the analysis on 2026-04-29. agentic = true: Entra Agent ID is a mutable permission boundary by design — roles, owners, credentials and Graph grants can all be reissued at runtime — so the containment staleness floor (M_Cn = 2.0) applies and no boundary re-attestation has occurred since. baw = false: the brief evidences none of the four Behavioral Attestation Window checklist items for the assessed component — no cross-session memory writable by a model or agent, no live agent-to-agent messaging, no self-modifying prompts or configuration, and no closed feedback loop; the assessed surface is an authorization control plane, not a behavioural runtime. C_monitor = 0.65: a detection gap is evidenced — the brief has to recommend alerting on an Agent ID Administrator adding owners or credentials to non-agent service principals, and states the takeover 'is easy to miss when those events are viewed separately'. C_event = 0.35: Microsoft's fix is an identity-boundary change to the assessed configuration, which is an enumerated model-or-architecture event; the evidence therefore describes a superseded state. C_evidence = 0.85: the class of gap remains open — the brief's own conclusion is that every other new AI-agent identity role still needs testing against the directory object types it inherits.

Public-evidence ACI = (Pc × Ec × Tf)1/3 = 0.06 — diagnostic status: Invalid as assessment-of-record evidence. It describes the evidence available to this case study, not the assurance of the underlying system, and is not inserted into the normalized-assurance scenario ERS.

Indicative ERS — normalized-assurance scenario

Worp · ORP0.35(0.75) + 0.25(0.50) + 0.25(1.00) + 0.15(0.00) = 0.637
CRMNelevated = 1 → 1.00
ORPeffective0.637 × 1.00 = 0.637
Wivp · IVP midpoint0.30(0.50) + 0.25(0.50) + 0.15(0.45) + 0.20(0.50) + 0.10(0.40) = 0.482
IVP mitigation0.15 + 0.85(1 − 0.482) = 0.590
Scenario assuranceACI fixed at 1.000 for cross-case comparison; public-evidence ACI 0.063 is diagnostic only
Indicative ERS midpointmin(10, 0.637 × 0.590 × 1/1.000 × 10) = 3.8
Unknown-input interval1.2–6.3; 19 unknown applicable sub-metrics set to 1.00 / 0.00 at the bounds

AIDEFEND defences → AITBM sub-metrics

Identifiers are quoted as they appear on the AIDEFEND in Action brief (retrieved 2026-08-13); the sub-metric mapping is AITBM's own, from the specification's AIDEFEND tables reconciled at catalogue data version 2026.08.05. AIDEFEND renumbers identifiers between releases, so the data version travels with every mapping and neither side's IDs should be cited without one. A mapping identifies a possible evidence route; a recommendation does not prove that the control was implemented or effective and receives no scoring credit by itself.

TechniqueDefence PriorityEvidences
AID-M-009.003Agent Identity, Delegation Lineage & Authorization ContextParent AID-M-009 'Agent Autonomy & Authority Governance' (dataVersion 2026.08.05). Produces the delegation-lineage record — who assigned the role, target object, original authority, delegated scope, exact action — that Cn-5 is scored on; Cn-6 evidence is listed by the lookup but this case supplies no action-classification trace, so Cn-6 was not scored.Very HighCn-1 Cn-5 Cn-6 Cn-7
AID-H-018.002Policy-Based Access ControlParent AID-H-018 'Tool Authorization & Capability Scoping' (dataVersion 2026.08.05). The directly failing control: a policy engine evaluating actor role, target object type and subtype would have failed closed on non-agent service principals. This is the Cn-1 evidence source.Very HighCn-1 Cn-6 Cn-7
AID-H-018.006Continuous Authorization Verification (Anti-TOCTOU)Parent AID-H-018 (dataVersion 2026.08.05). Addresses the two-step chain — become owner, then add credentials — where the second step inherited the authorisation decision of the first; evidence for Cn-2 in substance as well as the mapped Cn-1.HighCn-1 Cn-6 Cn-7
AID-H-004.001User & Privileged Access ManagementParent AID-H-004 'Identity, Access & Trusted Communication for AI Systems' (dataVersion 2026.08.05). Eligible-not-standing assignment, step-up authentication and periodic review of who may hold Agent ID Administrator; the UI's failure to label the role privileged is what made this hygiene less likely to be applied.HighCn-5 Tr-3
AID-E-001.003AI Agent & Workload Principal and Issuance DisablementParent AID-E-001 'Compromised Credential, Session, Principal & Grant Eviction' (dataVersion 2026.08.05). Response-side Cn-5 evidence: disabling the compromised principal at every issuance path, since removing one attacker-added secret leaves the ability to mint replacements.HighCn-5
AID-E-001.004Delegated Grant & Connected-App Authorization RevocationParent AID-E-001 (dataVersion 2026.08.05). Revoking OAuth grants, connected-app consents and app-role assignments that survive ordinary secret rotation — the persistence half of the takeover.HighCn-5
AID-D-011.004Non-Human Identity & Delegated Token Abuse DetectionParent AID-D-011 'Registered Agent Behavior, Interaction & Identity-Abuse Detection' (dataVersion 2026.08.05). The precise missing telemetry: correlating new credential issuance on a privileged service principal with the sign-in that follows. Its absence is the basis for the C_monitor cap.MediumCn-5 Cn-6
AID-D-005.004Specialized Agent & Session LoggingParent AID-D-005 'AI Activity Logging, Monitoring & Threat Hunting' (dataVersion 2026.08.05). Direct Tr-3 evidence source: preserving correlation IDs across role assignment, Graph call, owner addition, credential creation and authentication.MediumCn-7 Tr-3
AID-E-001.001Root & Long-Lived Credential Object EvictionParent AID-E-001 (dataVersion 2026.08.05). Removal of the unauthorised service-principal secrets or certificates at the authoritative issuer.MediumCn-5

WHAT THIS CASE TEACHES

An AI system can fail the Containment axis with no model in the loop at all: Cn-5 is scored on the enforced authorization boundary around agent identities, so a directory role whose documented scope and effective scope diverge is an agent-identity finding, not merely an IAM bug.

Sources: AIDEFEND in Action brief — Entra Agent ID Administrator Scope Gap: Agent Roles Reaching Service Principals (published 2026-04-29) · Primary source — Noa Ariel, Silverfort: Agent ID Administrator scope overreach: Service Principal takeover in Entra ID (2026-04-23)

AITBM sub-metrics referenced