Use these records as incident references, test-design inputs, and examples of evidence-to-rubric reasoning. They are retrospective scenarios, not current vendor ratings or assessments of record.
2026-08-12 · Agentic / MCP System
MCP onboarding cannot be a one-time trust decision: clients must bind descriptor and prompt semantics, detect drift, and re-authorize every consequential action against current user intent.
Indicative ERS 6.4 · 3.2–9.6
2026-08-02 · Standalone LLM / Generative AI
A familiar model name is not an identity: deployment admission must bind approved bytes, signer, source ownership, and loader policy before first load and every reload.
Indicative ERS 4.0 · 1.5–6.4
2026-06-01 · Agentic / MCP System
Ro-4 poisoning resistance is not only about training data and RAG corpora: an agent's own instruction and configuration files are an ingestion channel, and where they are read as authoritative guidance with no provenance, signature or hidden-character check, Ro-4 sits at the 0.00 anchor regardless of how well the training pipeline is protected.
Indicative ERS 6.7 · 3.8–9.7
2026-05-08 · Agentic / MCP System
The adversary never touched the model — it optimised the documentation the model reads — so the scoring weight lands on Ro-4 ingestion integrity and Cn-1/Cn-6 execution gating rather than on jailbreak resistance, and the case shows why AITBM scores the pipeline configuration rather than the assistant that proposed the change.
Indicative ERS 6.1 · 3.8–8.4
2026-05-03 · Agentic / MCP System
Containment can score 0.00 on an AI product whose model behaved perfectly: the assessed boundary here is the extension host and the credential store, so a client-side trust-boundary failure lands squarely on Cn-1 and Cn-5 — and demonstrates that a 'no model involvement' incident is still an AI security finding, not an exemption from scoring.
Indicative ERS 4.0 · 1.8–6.3
2026-04-25 · Agentic / MCP System
The trust boundary that fails here is neither the model nor the tool but the transport intermediary between them, and AITBM localises it precisely — Cn-5 = 0.25 for an unverifiable response origin under API-key-only identity and Cn-6 = 0.00 for ARCR = 0 under auto-approve — a failure that no model-layer sub-metric and no prompt-injection test would have surfaced.
Indicative ERS 7.8 · 4.8–10.0
2026-04-18 · Tool-Calling LLM / Connected GenAI
AITBM scores a software supply-chain compromise as an AI-system finding without distorting either: the AI-specific signal lands on Ro-4 (the dependency ingestion path is a poisoning channel), Cn-5 (static credentials with no workload binding) and Tr-4 (no AIBOM to bound the blast radius), while the thinness of published deployment detail is carried honestly by a base coverage of 4 of 22 rather than by inventing scores — the epistemic penalty appears in ACI, which is where the framework intends it.
Indicative ERS 4.2 · 1.6–6.8
2026-04-18 · Agentic / MCP System
A containment control can hold and the deployment still be compromised: the agent sandbox blocked in-runtime execution, which is why Cn-1 is scored up to 0.50 rather than down at a failure anchor, while the compromise travelled the one path the sandbox never covered — an instruction relayed through the human. AITBM captures this only because Cn-1's 0.50 anchor names delegated workflows explicitly, and because the assessment is required to score the control that held alongside the ones that did not.
Indicative ERS 5.9 · 2.7–9.2