Open framework · Community-built · Free

Benchmark AI safety and security across the system you deploy.

AITBM — the AI Trust Benchmarking and Maturity Framework — evaluates deployed LLM, RAG, agentic, MCP, and model-pipeline systems with fixed rubrics, deployment context, and evidence confidence. You get a single 0–10 risk score while the safety and security detail behind it remains visible.

0–10
single risk score (ERS)
3
assessment layers
23
checks across 5 areas
5-level
evidence-bound rubrics
AITBM IN 60 SECONDS

The short version

What it is

A repeatable way to measure how risky an AI system is — built to narrow the disagreement between assessors that vague severity ratings leave wide open.

What you get

One 0–10 Effective Risk Score, plus the five-dimension profile that explains it.

Who it's for

Security assessors, compliance & risk teams, and engineers shipping LLM and agentic systems.

What it costs

Nothing. Open, community-built, and MIT-licensed — no paid tooling required.

ASSESSMENT & REFERENCE GUIDES

Find the benchmark, method, framework, or case you need

Start from your question. Each guide connects the scoring model to complete rubrics, framework crosswalks, and public-evidence incident references.

HOW IT WORKS

Three lenses, combined into one score

AITBM looks at a system through three independent lenses, then composes them into a single ERS. Each lens answers a different question — and a strong result in one cannot hide a weak result in another.

LAYER 1 · THE SYSTEM

IVP

Intrinsic Vulnerability Profile

Asks: how strong is the system itself?

23 checks across 5 security areas — Robustness, Fairness, Transparency, Privacy, Containment — each scored at one of five normalized anchors from 0.00 to 1.00.

LAYER 2 · THE DEPLOYMENT

ORP

Operational Risk Posture

Asks: how risky is where and how it's used?

Autonomy, exposure, blast radius, and how hard it is to fix — combined into a Compound Risk Multiplier (CRM).

LAYER 3 · THE EVIDENCE

ACI

Assurance Confidence Index

Asks: how much can we trust what we know?

Discounts the score as evidence ages or thins out — and flags when it's time to re-assess.

ERS

Effective Risk Score
ERS = α + (1 − α) × f(IVP, ORP, ACI)      where α = 0.15

One number, 0–10. The residual-risk floor (α = 0.15) is deliberate: even perfect controls leave an irreducible 15% of risk. AI risk can never be zeroed out — so AITBM never pretends it can.

SEVERITY SCALE

0 Low5 Moderate10 Critical

A higher ERS means higher residual risk — and triggers a deeper assessment tier.

WHY IT EXISTS

What today's frameworks miss

The comparison of CVSS adaptations, OWASP AIVSS, and the OWASP Top 10 for LLMs identifies recurring scoring gaps across artifacts with different scopes. AITBM was designed to address them.

Guesswork

In the largest published CVSS consistency study, 68% of assessors changed at least one metric when re-scoring identical vulnerabilities (Wunder et al., IEEE S&P 2024).

AITBM → fixed 0–4 rubrics

Context blindness

A medical-diagnosis model and a recipe chatbot with identical flaws score identically.

AITBM → ORP deployment layer

Scope blindness

Static-model assumptions miss agentic and MCP threats — tool poisoning, rogue agents, identity spoofing.

AITBM → Containment axis + Cn-5

Zero-risk fallacy

Frameworks imply enough controls eliminate risk. Emergent behavior makes that impossible.

AITBM → α = 0.15 risk floor + behavioral attestation decay

See the full 12-gap analysis and head-to-head comparison →
START HERE

Pick the path that fits you

New to AITBM? Here's the shortest route to what you need.

2

I own compliance & risk

See how AITBM maps to ISO 42001, NIST AI RMF, and the EU AI Act — and where it goes further.

3

I want to contribute

Pilot an assessment, validate inter-assessor consistency, or open an issue or pull request.

WORKED EXAMPLE

Proven on a real scenario: "Finbot"

Finbot is a financial-advisory agent that combines RAG with tool calling — AITBM's canonical test case. Because its rubric and weights are fixed rather than left to judgment, independent assessors converge far more closely than vague severity ratings allow — that consistency is the point.

See controls take Finbot from 9.7 to 3.2 →
10.0
ERS
1.35
CRM
Critical
MVT severity

Aligned to the standards you already use

AITBM maps against OWASP, MITRE ATLAS, NIST AI RMF, ISO 42001/42005, the EU AI Act, and the AIDEFEND defensive taxonomy — turning controls into measurable evidence rather than checklists.

Explore the framework Resources & standards