Use these records as incident references, test-design inputs, and examples of evidence-to-rubric reasoning. They are retrospective scenarios, not current vendor ratings or assessments of record.
2026-08-05 · Standalone LLM / Generative AI
An internal AI message bus is still an untrusted code boundary: peer identity, segmentation, and safe serialization must hold before bytes reach a deserializer.
Indicative ERS 3.5 · 1.4–5.5
2026-08-02 · Agentic / MCP System
Model weights are only one executable release component; chat templates must be digest-bound, provenance-checked, and security-regression-tested with the same rigor as the weights.
Indicative ERS 4.8 · 2.0–7.7
2026-08-02 · Standalone LLM / Generative AI
A familiar model name is not an identity: deployment admission must bind approved bytes, signer, source ownership, and loader policy before first load and every reload.
Indicative ERS 4.0 · 1.5–6.4
2026-07-06 · Agentic / MCP System
JADEPUFFER is the case that forces the assessor to keep the attacker out of the assessed system: the agentic behaviour that compressed the kill chain to 31-second self-repair belonged to the offence, so it drives no ORP or BAW score here — what AITBM scores is a Tier 1 workflow host with Cn-1, Cn-2 and Cn-5 all at 0.00 and a corroborated Cp = 1.00 traced from an unauthenticated origin to a P4 backdoor-administrator terminal.
Indicative ERS 6.9 · 3.0–10.0
2026-07-02 · Standalone LLM / Generative AI
Architecture class is scored on the assessed boundary, not on the attacker's behaviour: the exposed backend is a Standalone LLM with Aa = 0.25 and baw/agentic both false, yet three Containment sub-metrics sit at 0.00 and Cp is corroborated at 1.00 — an unauthenticated endpoint can be maximally exposed and trivially remediable (Rf = 0.00) at the same time, and the thin two-axis evidence base correctly forces the Lite Ec cap.
Indicative ERS 4.3 · 1.9–6.8
2026-06-27 · Multi-Agent / MCP System
Detection capability and evidence freshness are scored in different layers, and this case separates them cleanly: the monitor's blindness is a point-in-time IVP finding at Ro-1 and Cn-3, while the same sensor loss independently caps ACI Temporal Freshness through C_monitor and Band 0 C_behavior — and because the report measured visibility rather than consequence, Cn-1 and Cn-6 are correctly omitted instead of guessed.
Indicative ERS 5.0 · 2.0–8.0
2026-05-26 · RAG / Retrieval-Augmented System
Remediation Feasibility sits at its bottom anchor — a patchable ordering bug with a CVE and a fixed release — while Cascade Potential still forces 1.00, which is exactly the layer separation AITBM is built for: how easily a defect is fixed and how far it reaches are scored independently, so a cleanly patchable flaw never gets to look harmless.
Indicative ERS 4.8 · 2.7–7.0
2026-05-15 · Agentic / MCP System
Nothing in this case reaches the model — there is no prompt, no injection, and no Robustness-1 signal at all — so AITBM's not-applicable redistribution rule carries the whole assessment on Containment, identity, and posture evidence, which is the correct answer for an incident where, as the brief puts it, security was lost before the model saw any prompt.
Indicative ERS 7.2 · 3.4–10.0
2026-04-24 · Traditional ML / Classifier
An artifact-supply-chain incident with no agent, no prompt and no model output is still fully scorable — the evidence lands on Ro-4, Tr-3/Tr-4 and Cn-1/Cn-2 — and because the assessed configuration was observed in February 2024, Temporal Freshness rather than the attack itself is what collapses assurance confidence, which is the temporal layer behaving exactly as designed.
Indicative ERS 4.8 · 2.3–7.4
2026-04-16 · Research
This brief reports a controlled research result about a training-time mechanism, not an incident against a deployed AI system. The teacher and student models were created by the researchers to demonstrate the effect; there is no victim deployment, no operator, no production configuration, and no attack against a running system. Consequently the three AITBM layers have no referent: the IVP would have to be scored against a research artefact rather than an assessed configuration, and every ORP dimension - autonomy, attack surface, cascade potential, remediation feasibility - would have to be invented, since a distillation pipeline demonstrated in a laboratory has no deployment context to score. The protocol's representative-configuration exception does not rescue it either: the brief studies a class of training pipeline, not a class of deployment, and it supplies no configuration facts (autonomy, exposure, downstream reach) from which a representative deployment could be reconstructed. Scoring it would manufacture numbers the evidence cannot support. The brief is nonetheless directly useful to AITBM as a scoping input, recorded in key_finding below.
Research note · no ERS