Gap analysis

Twelve structural deficiencies in existing AI security assessment methodologies — grouped into four domains — and how AITBM addresses each.

THE BOTTOM LINE

The comparison identifies recurring scoring gaps across artifacts with different purposes. AITBM is the only compared framework with full coverage across all fifteen capabilities — and the only one with five-level rubrics and required test methods per sub-metric. It doesn't replace OWASP AIVSS or AIUC-1; it measures system risk across dimensions, context, and time where they don't.

12
structural gaps identified
11 / 12
fully addressed today
10
frameworks compared
43+
research sources

Four failure modes

1. Methodological foundations

Ambiguous severity language produces measurable assessor divergence — in the largest CVSS consistency study, 68% of assessors re-scored identical vulnerabilities differently (Wunder et al., IEEE S&P 2024) — plus single-score reductionism that discards multi-dimensional signal.

2. Operational reality

Systems are scored in isolation. Deployment context, threat sophistication, and evidence staleness go unmodeled.

3. Scope blindness

Static-model assumptions miss agentic and MCP threats: autonomy risks, agent identity, and tool/function-calling attack surface.

4. Structural impossibilities

Frameworks imply zero risk is achievable. Emergent behavior and cross-layer cascades make that false.

The twelve gaps

Twelve structural deficiencies across four domains, each with a short read on what it means. Eleven are fully addressed today; one is partially addressed — a residual inherent to deterministic scoring, mitigated by the α floor and behavioral attestation decay. The framework comparison below shows how every other framework we’ve mapped fares on the same capabilities.

Gap Impact AITBM solution Coverage
Methodological foundations
Subjectivity in scoringSeverity words like “high”/“medium” admit several readings, so assessors diverge on the same system.68% re-score divergence (CVSS; Wunder et al., IEEE S&P 2024)Five-level quantitative rubricsFull
Single-score reductionismOne 0–10 number hides which dimension actually drives the risk, so controls can’t be aimed.Multi-dimensional signal lossThree-layer IVP / ORP / ACI architectureFull
Incomplete scoring guidanceTiers are defined but with no test method, threshold, or tool to measure them.Implementation paralysisTest methods + tool integration per sub-metricFull
Operational reality
Context-blind scoringIntrinsic model risk is scored while where the system is deployed is ignored.Critical vs. low-risk indistinguishableORP layer + Compound Risk Multiplier (CRM)Full
Temporal evidence decayPoint-in-time assessments quietly go stale as models drift and threats evolve.Stale assessments, false confidenceACI decay + re-assessment triggersFull
Operational vs. intrinsic conflationRuntime controls are credited as if they reduced intrinsic risk, hiding control failure.Zero-risk fallacyIVP + ORP separation; α = 0.15 floorFull
Scope blindness
Agentic autonomy risksStatic-model frameworks don’t score autonomous, tool-using, coordinating agents.Agentic threat classes unaddressedContainment axis (Cn-1 … Cn-7)Full
Identity security gapsHuman IAM doesn’t fit ephemeral, delegating agents; agent identity goes unscored.Agent impersonation; no IAM for agentsCn-5 Agent Identity IntegrityFull
Tool / function-calling securityThe MCP and tool-calling surface — poisoning, CVEs — goes unassessed.Tool poisoning; MCP CVEsRo-4 + Cn-1/Cn-2 + Tr-3Full
Structural impossibilities
Residual risk floorFrameworks imply controls can reach zero risk; emergent and semantic risk make that false.False belief in zero riskα = 0.15 irreducible riskFull
Emergent-behavior unpredictabilityBehaviour not predictable from component tests — collusion, drift, memory poisoning.Non-deterministic risk underestimatedα = 0.15 floor + behavioral attestation decay (BAW: T_behavior, M_Em = 3.0) + C_behavior monitoring caps + §7 BBD drift protocolPartial
Cross-layer cascading failuresVulnerabilities amplify across the agent stack; independent layer scoring misses it.Amplification not modeledGraph-derived Cp (GDCP): verified dependency graph, layer reachability, privilege amplification, fault-injection blast radius; worst indicator governs; Cn-6 chain rule retainedFull
11 / 12 fully addressed 1 / 12 partially addressed (inherent to deterministic scoring; mitigated by the α floor + behavioral attestation decay)

Evidence from 2025–2026 research

The analysis draws on five research segments and 43+ sources. The agentic/MCP segment in particular surfaced the threat data that motivates AITBM's Containment axis.

72.8%

tool-poisoning attack success rate reported by the MCPTox benchmark (arXiv 2508.14925).

40.55%

of 7,973 live remote MCP servers exposed tools without authentication (Zhou et al., arXiv 2605.22333). Separately, Censys measured 12,520 internet-accessible services across 8,758 IP addresses.

82%

of 2,614 analyzed MCP servers were prone to path-traversal flaws (Endor Labs, 2026).

CVSS 9.3

EchoLeak (CVE-2025-32711) — zero-click prompt-injection in Microsoft 365 Copilot (Microsoft score; NVD base 7.5).

CVSS 9.6

RCE in the mcp-remote connector (CVE-2025-6514; ~437k downloads) — JFrog, 2025.

SPIFFE / OIDC-A

agent-identity standards mapped directly onto the Cn-5 rubric levels.

Figures are drawn from published 2025–2026 security research cited in the AITBM gap analysis. See the Resources page for the full documentation.

Current comparison basis

Framework comparison

How AITBM compares with every framework we’ve mapped, across the same capabilities the twelve gaps describe. Most are a different kind of artefact — vulnerability catalogs (OWASP LLM Top 10), control checklists (OWASP AISVS), threat taxonomies (MITRE ATLAS), governance frameworks (NIST AI RMF, ISO 42001), regulation (EU AI Act), or certification (AIUC-1) — so a “No” on a scoring capability is a scope difference, not a criticism. AITBM adds a quantitative layer the others are not designed to provide, and it ingests their outputs as scoring evidence. Re-validated against OWASP AIVSS v0.8 (March 2026) and extended with AIUC-1 (launched 2025).

Capability CVSS 4.0 LLM Top 10 AISVS ATLAS AIVSS v0.8 RAISE NIST RMF ISO 42001 EU AI Act AIUC-1 AITBM
AI-native (non-deterministic) No Partial Partial Yes Partial Yes Yes Yes Yes Yes Yes
Multi-dimensional profile No No No No No Yes Partial No No No Yes
Deterministic weights N/A N/A N/A N/A Yes Partial N/A N/A N/A N/A Yes
Epistemic confidence scoring No No No No No No No No No No Yes
Behavioral drift monitoring No No No No No No No No No Partial Yes
Supply chain integration (AIBOM) No Partial Partial Partial Partial No Partial Partial No Partial Yes
Tiered SME pathway N/A No Partial No No No Partial No Partial Partial Yes
Stateful/agentic risk modeling No Partial Partial Partial Yes No No No No Partial Yes
MVT enforcement per dimension No No No No No No No No No Partial Yes
Compound operational risk No No No No No No Partial No No No Yes
Residual deployment risk floor No No No No Partial No Partial Partial No Partial Yes
Graduated MVT severity No No No No No No No No Partial No Yes
Jurisdictional fairness No No No No No No Partial No Partial Partial Yes
Architecture classification tree No No No No Partial No No No Partial Partial Yes
Inter-rater reliability targets No No No No No No No No No No Yes

Verdicts for CVSS, AIVSS v0.8, RAISE, AIUC-1 and AITBM mirror the current framework specification; the catalog, governance and regulatory frameworks are rated on the same capabilities for completeness. RAISE is an academic responsible-AI evaluation framework (arXiv 2510.18559) — not a security/vulnerability scorer — included here only for breadth. Scroll horizontally to see all columns.

Scoring systems

CVSS 4.0 and OWASP AIVSS v0.8 produce a single security-severity number, collapsed to one scalar with no standing evidence-freshness or confidence layer (CVSS 4.0's Environmental group does capture deployment context, but folds it into the same number rather than preserving it as a separate signal). AIVSS v0.8 fixes its factor weights and adds a 0.67 multiplicative mitigation floor, but keeps a single 0–10 score on three-point factor anchors without operational test methods, and scopes itself to security — by design it does not cover fairness, privacy or transparency. RAISE is a different kind of tool: an academic responsible-AI evaluator (Explainability, Fairness, Robustness, Sustainability) that also rolls up to one score, shown here for breadth rather than as a security scorer.

CVSS 4.0 does have a time-varying dimension — its Threat group (renamed from v3.x Temporal; a single Exploit Maturity metric). But that scores a known vulnerability's exploitation state — is there public exploit code, is it being attacked — set manually from threat intel, with no decay function. A CVSS score never ages on its own. That is a different concept from evidence-freshness decay: AITBM's ACI discounts assessment confidence as evidence goes stale and the system drifts.

Catalogs & taxonomies

OWASP LLM Top 10, OWASP AISVS and MITRE ATLAS enumerate threats and controls; by design they don't output a quantitative score, profile, or temporal model. AITBM maps their elements to evidence for applicable sub-metrics (e.g. LLM01 → Ro-1 + Cn-3); the score then comes from the assessed deployment and its evidence, not from the threat or checklist item alone.

Governance, regulation & certification

NIST AI RMF, ISO 42001, the EU AI Act and AIUC-1 govern, regulate and certify — the EU AI Act imposes binding, risk-tiered obligations; the others classify risk qualitatively or issue pass/fail certificates — but none outputs a quantitative, time-aware score. AIUC-1's agent-identity coverage is thin: the AIVSS–AIUC-1 crosswalk maps about two controls (E016, F001) to agent-identity impersonation, and they are policy-and-disclosure oriented rather than a graduated cryptographic-identity rubric — the depth Cn-5 adds. AITBM can use their attestations and test reports as scoring evidence. AIUC-1's insurance mechanism transfers risk, while AITBM's α = 0.15 floor quantifies a separate assumption of non-zero residual risk; one does not validate the other.

Verdict: across all fifteen capabilities AITBM is the only one of the compared artefacts designed to address every one in a single model, and it is distinctive in specifying a five-level rubric and a required test method for each sub-metric. The others were never intended to produce a quantitative, multi-dimensional, time-aware risk score; AITBM doesn’t replace them, it quantifies their outputs across dimensions, context, and time.