IVP
Intrinsic Vulnerability ProfileAsks: how strong is the system itself?
23 checks across 5 security areas — Robustness, Fairness, Transparency, Privacy, Containment — each scored at one of five normalized anchors from 0.00 to 1.00.
Open framework · Community-built · Free
AITBM — the AI Trust Benchmarking and Maturity Framework — evaluates deployed LLM, RAG, agentic, MCP, and model-pipeline systems with fixed rubrics, deployment context, and evidence confidence. You get a single 0–10 risk score while the safety and security detail behind it remains visible.
A repeatable way to measure how risky an AI system is — built to narrow the disagreement between assessors that vague severity ratings leave wide open.
One 0–10 Effective Risk Score, plus the five-dimension profile that explains it.
Security assessors, compliance & risk teams, and engineers shipping LLM and agentic systems.
Nothing. Open, community-built, and MIT-licensed — no paid tooling required.
Start from your question. Each guide connects the scoring model to complete rubrics, framework crosswalks, and public-evidence incident references.
Measure robustness, fairness, transparency, privacy, containment, deployment risk, and confidence.
A repeatable methodology for LLM, RAG, agentic, MCP, and model-pipeline deployments.
Select evidence-producing tests across all 23 fixed scoring rubrics.
Connect OWASP, MITRE, NIST, ISO, EU, CSA, CVSS, and AIDEFEND to measured evidence.
Browse prompt injection, agents, MCP, RAG, supply chain, data exposure, coding agents, and model security.
AITBM looks at a system through three independent lenses, then composes them into a single ERS. Each lens answers a different question — and a strong result in one cannot hide a weak result in another.
Asks: how strong is the system itself?
23 checks across 5 security areas — Robustness, Fairness, Transparency, Privacy, Containment — each scored at one of five normalized anchors from 0.00 to 1.00.
Asks: how risky is where and how it's used?
Autonomy, exposure, blast radius, and how hard it is to fix — combined into a Compound Risk Multiplier (CRM).
Asks: how much can we trust what we know?
Discounts the score as evidence ages or thins out — and flags when it's time to re-assess.
ERS = α + (1 − α) × f(IVP, ORP, ACI) where α = 0.15
One number, 0–10. The residual-risk floor (α = 0.15) is deliberate: even perfect controls leave an irreducible 15% of risk. AI risk can never be zeroed out — so AITBM never pretends it can.
SEVERITY SCALE
A higher ERS means higher residual risk — and triggers a deeper assessment tier.
The comparison of CVSS adaptations, OWASP AIVSS, and the OWASP Top 10 for LLMs identifies recurring scoring gaps across artifacts with different scopes. AITBM was designed to address them.
In the largest published CVSS consistency study, 68% of assessors changed at least one metric when re-scoring identical vulnerabilities (Wunder et al., IEEE S&P 2024).
AITBM → fixed 0–4 rubrics
A medical-diagnosis model and a recipe chatbot with identical flaws score identically.
AITBM → ORP deployment layer
Static-model assumptions miss agentic and MCP threats — tool poisoning, rogue agents, identity spoofing.
AITBM → Containment axis + Cn-5
Frameworks imply enough controls eliminate risk. Emergent behavior makes that impossible.
AITBM → α = 0.15 risk floor + behavioral attestation decay
New to AITBM? Here's the shortest route to what you need.
Understand the scoring model, see it applied to real incidents with the evidence behind every sub-metric, then score a system of your own.
See how AITBM maps to ISO 42001, NIST AI RMF, and the EU AI Act — and where it goes further.
Pilot an assessment, validate inter-assessor consistency, or open an issue or pull request.
Finbot is a financial-advisory agent that combines RAG with tool calling — AITBM's canonical test case. Because its rubric and weights are fixed rather than left to judgment, independent assessors converge far more closely than vague severity ratings allow — that consistency is the point.
See controls take Finbot from 9.7 to 3.2 →AITBM maps against OWASP, MITRE ATLAS, NIST AI RMF, ISO 42001/42005, the EU AI Act, and the AIDEFEND defensive taxonomy — turning controls into measurable evidence rather than checklists.