All 64 current AIDEFEND in Action analyses are represented: 56 have
evidence-bounded AITBM scenarios with uncertainty intervals, while 8 retain
traceable AIDEFEND mappings without an ERS.
These are not assessments of record. Search by incident, architecture, AITBM sub-metric, or
AIDEFEND technique, or browse the categories below.
64source analyses
5.2median indicative ERS
2Critical MVT confirmed
43Compound Risk Alerts
Indicative ERS midpoints span 3.5–8.0:
9 High, 23 Moderate, and 24 Low-Moderate.
Unknown controls remain unknown; no case is forced to Critical because evidence is absent.
INCIDENT REFERENCE COLLECTIONS
Browse the library by threat and system type
Use focused collections to find comparable incidents,
reusable evaluation questions, related AITBM sub-metrics, and framework routes.
SCORING METHODHow these scores were produced
Read the assumptions, uncertainty rules, and evidence limits behind every case.
Each case is a retrospective public-evidence scenario for the affected configuration, built from the
AIDEFEND in Action
analysis plus cited vendor and researcher material. It is a website companion, not an AITBM assessment
of record and not a statement about any vendor's current product.
Unknown is not NOT APPLICABLE
Architecture alone determines applicability. Missing public
evidence remains unknown. The result is recomputed with unknown applicable inputs at 1.00, 0.50,
and 0.00 to show the best-security bound, maximum-entropy midpoint, and worst-security bound.
Only observed controls earn credit
A demonstrated allowlist rejection, authorization denial, or
integrity verification can raise a sub-metric. A recommended defence, presumed log, or product
capability does not prove operation in the assessed configuration and earns no positive credit.
Cascade Potential is graph-derived, and usually defaults
Cp is not read off a prose ladder: it comes from a verified System
Dependency Graph, and spec 3.2.1 forces Cp = 1.00 whenever no such graph exists. No
organisation publishes one after an incident, so Cp sits at 1.00 throughout. Each case says whether
that was corroborated by an observed ungated path to a privileged node, or is merely the
worst-case default — 19 of 56 rest on the default alone.
Assurance is normalized for comparison
The comparative scenario fixes ACI at 1.00 so article completeness
does not overwhelm the system-risk signal. The public-evidence ACI is still calculated and shown as a
diagnostic, but it is not represented as actual system assurance or inserted into the scenario ERS.
How to read the midpoint and interval
The midpoint assigns every unknown applicable sub-metric 0.50, the expected value of a maximum-entropy
prior over the five rubric anchors. This is a deterministic comparison convention, not a finding that
an undisclosed control achieved 0.50. The interval is the primary uncertainty statement.
The midpoint supports transparent comparison between the case studies; it must always be quoted with
its interval and the label Indicative ERS (normalized-assurance
scenario). A first-party assessment replaces the interval with measured inputs and uses its
actual ACI in the normative ERS formula.
Temporal evidence is diagnostic, not normalized to seven days
Every workpaper preserves its documented evidence age and event caps. Those values calculate a
public-evidence ACI diagnostic on the stated reference date. The generator no longer overwrites every
case with a seven-day age, because a constant would be a scenario assumption rather than measured
evidence freshness.
The Behavioral Attestation Window uses the specification's non-double-counting rule:
lambdabehavior = lambdatier × Mthreat ×
max(MTDI, MEm). The public-evidence ACI does not determine the comparative
midpoint; it tells readers how much confidence the article package itself can support.
READ THIS BEFORE QUOTING A NUMBER
These are case-study scenarios, not vendor ratings. Each result
describes one documented configuration or representative boundary. Remediation and current products
are outside the score unless a separate reassessment is performed.
All 56 scenarios use the current 23-sub-metric basis as of
2026-08-13. Cn-7 remains UNKNOWN unless the public evidence establishes resource-budget
and execution-loop containment; UNKNOWN stays in the interval and is never converted into inferred credit.
Quote the midpoint with its interval. Public evidence cannot
establish all 23 current sub-metrics, a verified dependency graph, or an assessment-of-record ACI. The
midpoint is useful only with its best-to-worst unknown-input bounds and normalized-assurance label.
Judgement was applied, and it is shown. Rubric placement is an
assessor call. Every placement on this page carries the evidence behind it precisely so a reader can
disagree with a specific number instead of the whole result.
The arithmetic and unknown-input rule are fixed. Scores are computed by
the framework's own formulas, and the engine is checked against the Finbot worked example
(ERS = 10.0, CRM = 1.35, Critical MVT) on every build. If that anchor ever fails to
reproduce, this page refuses to build.
CASE DISCOVERY
Find the evidence you need
Search directly, browse by category, or scan the compact
directory. Every result opens the complete evidence trail and calculation.
MCP onboarding cannot be a one-time trust decision: clients must bind descriptor and prompt semantics, detect drift, and re-authorize every consequential action against current user intent.
Opaque state is not safe state: reasoning artifacts need confidentiality, principal/session binding, bounded retention, and an output gate before they enter logs or client-visible trajectories.
Observability metadata is routing authority when it can select an exporter; trace context must be schema-restricted, destination-pinned, and minimized before export.
Tool-calling LLM
Indicative ERS4.1range 1.5–6.6
MVT evidenceMVT indeterminate
2026-08-05 · Standalone LLM / Generative AI · Tier 1
An internal AI message bus is still an untrusted code boundary: peer identity, segmentation, and safe serialization must hold before bytes reach a deserializer.
Serialized conversation history is active authority when deserialization triggers I/O; replay paths require the same destination policy as an explicit network tool call.
Long-term memory is persistent control input: every write needs provenance and promotion policy, and every recall must be re-authorized against the current task before it can influence tools.
Model weights are only one executable release component; chat templates must be digest-bound, provenance-checked, and security-regression-tested with the same rigor as the weights.
Agentic / MCP
Indicative ERS4.8range 2.0–7.7
MVT evidenceMVT indeterminate
2026-08-02 · Standalone LLM / Generative AI · Tier 1
A familiar model name is not an identity: deployment admission must bind approved bytes, signer, source ownership, and loader policy before first load and every reload.
Persistent memory changes prompt injection from a one-chat event into delayed control flow; trusted intent must govern both memory promotion and every later rehydration.
A sandbox is only as strong as its least-governed sibling tool: the Bash boundary held, but the in-process Read path crossed the same CI-secret boundary without equivalent enforcement.
A platform can ship an approval gate and still score Cn-6 = 0.00 when that gate is writable by the very workflow it governs — AITBM scores the exercised configuration, not the nominal capability — and…
A case can have eight or nine genuinely effective controls — an SSRF allowlist that never failed, 84 straight denials at the credential store, digest verification that kept the supply chain clean — and…
A human-approval gate only counts for as much of the chain as it actually governs: Notion's page-edit approval was real enough to raise Aa to 0.50 and hold Cn-6 above zero, yet the render hop that fired…
This is the reference case for cp_basis = default: the agents are maximally exposed and their navigation action is entirely ungated, yet the measured terminal impact is a benign HTTP request, so Cp = 1.00…
AgentFlayer is the batch's clearest case of a control that half-held changing two scores at once: because a URL reputation check existed and had to be bypassed, Cn-3 sits at the rubric's bypass anchor…
This is the cleanest Cn-6 = 0.00 in the set and shows why AITBM separates Cn-6 from Cn-1: the agent never exceeded its granted CRM permissions, so a permission-only assessment would find nothing wrong, yet…
This case shows the Behavioral Attestation Window being correctly withheld: the agent is unambiguously Agentic and scores at the ORP ceiling, but none of the four BAW checklist items is evidenced for a…
LangGraph is the cleanest available demonstration that agent memory is an Ro-4 poisoning surface rather than a storage detail: an attacker-planted checkpoint row promoted straight into trusted agent state…
Two AITBM behaviours meet in this case: As sitting exactly on 0.75 keeps N_elevated at one and CRM at 1.00 even though the cascade path is corroborated all the way to a credential-issuing node, and…
A loopback-bound control plane is not a lower-exposure deployment: because a browsing agent renders untrusted content on the same host, As reaches the 1.00 anchor and the Cp path is corroborated end-to-end…
This is the batch's only case with no patch to point at, and it lands where AITBM says it should: Rf = 0.75 rather than 0.00 because every remedy is an external containment layer, while Cn-6 = 0.00 for a…
PerplexedBrowser is the case where the assessed system's own security controls were largely irrelevant to the outcome: the vault held, the model filter partly held, and the system still failed — because…
When a client treats repository content as configuration, the poisoning surface is the tool manifest rather than the model — Ro-4 and Cn-5 both collapse to 0.00 on the same evidence, and a shipped,…
DifyTap is the case that separates audit-trail completeness from audit-trail safety: Tr-3 field coverage was strong enough for the 0.75 band and still capped at 0.50, because the 0.75 criterion requires the…
JADEPUFFER is the case that forces the assessor to keep the attacker out of the assessed system: the agentic behaviour that compressed the kill chain to 31-second self-repair belonged to the offence, so it…
An AI gateway with no agents, no memory and no autonomy still classifies as Connected GenAI and still forces Cp = 1.00 on its merits — because the corroborating path runs through a credential-issuing (P4)…
Tool-calling LLM
Indicative ERS4.8range 2.0–7.5
MVT evidenceMVT indeterminate
2026-07-02 · Standalone LLM / Generative AI · Tier 3
Architecture class is scored on the assessed boundary, not on the attacker's behaviour: the exposed backend is a Standalone LLM with Aa = 0.25 and baw/agentic both false, yet three Containment sub-metrics…
Detection capability and evidence freshness are scored in different layers, and this case separates them cleanly: the monitor's blindness is a point-in-time IVP finding at Ro-1 and Cn-3, while the same…
This is the clean illustration of cp_basis = 'default' versus 'corroborated': all four stack layers were reachable and data did leave the tenant, but every terminal node in the observed path was…
A tool integration that is read-only by design can still drive an agent to code execution: the boundary that failed is the MCP output boundary, not the model, which is why Ro-1 stays at 0.25 while five…
A confirmation gate that exists but is not bound to a canonical action summary earns Cn-6 = 0.25, not credit for human-in-the-loop control — and because the gate cannot be claimed at CBR >= 0.95, the same…
Cascade Potential can legitimately be a default rather than a corroborated 1.00: the injected content traverses all four stack layers, but the terminal node is an external fetch rather than a write-external…
Ro-4 poisoning resistance is not only about training data and RAG corpora: an agent's own instruction and configuration files are an ingestion channel, and where they are read as authoritative guidance with…
Authentication is not authorisation: a cryptographically valid, correctly authenticated peer session still scores Cn-5 low, because Cn-5 measures whether identity is bound to instruction provenance and tool…
Agentic / MCP
Indicative ERS8.0range 4.3–10.0
MVT evidenceMVT floor: at least Major
2026-05-26 · RAG / Retrieval-Augmented System · Tier 3
Remediation Feasibility sits at its bottom anchor — a patchable ordering bug with a CVE and a fixed release — while Cascade Potential still forces 1.00, which is exactly the layer separation AITBM is built…
Nothing in this case reaches the model — there is no prompt, no injection, and no Robustness-1 signal at all — so AITBM's not-applicable redistribution rule carries the whole assessment on Containment,…
The adversary never touched the model — it optimised the documentation the model reads — so the scoring weight lands on Ro-4 ingestion integrity and Cn-1/Cn-6 execution gating rather than on jailbreak…
The exploited input never reached the model, which is exactly why Ro-1 must be scored over the agent's whole task-setup surface rather than its prompt: an agent's adversarial-input resistance is only as…
Containment can score 0.00 on an AI product whose model behaved perfectly: the assessed boundary here is the extension host and the credential store, so a client-side trust-boundary failure lands squarely…
There is no adversary in this case at all, and AITBM still scores it as a Containment collapse — Cn-1, Cn-2 and Cn-6 are driven by what the agent was technically able to do, not by whether anyone attacked…
An AI system can fail the Containment axis with no model in the loop at all: Cn-5 is scored on the enforced authorization boundary around agent identities, so a directory role whose documented scope and…
Tool abuseAgentic / MCP
Indicative ERS3.8range 1.2–6.3
MVT evidenceMVT indeterminate
2026-04-29 · RAG / Retrieval-Augmented System · Tier 2
An incident with no established root cause is still scorable — the enforced boundary and the missing release gate are directly observable from the outcome — and the unresolved mechanism belongs in ACI (thin…
A population study can be scored honestly as a representative configuration, but only if the assurance layer carries the cost: Pc = 0.10 and a class-level Cp classified as default rather than corroborated…
Authentication succeeded and the provider's write boundary held, so the failure has to be located precisely rather than described as 'broken identity': AITBM puts it in Cn-5 = 0.40 (a valid managed token…
Controls that held move the numbers as much as the ones that failed: a measured 76% attack success rate coexists with Cn-1 and Cn-2 at 0.50 (the OS app sandbox bounded the blast radius) and Rf at 0.25 (a…
The trust boundary that fails here is neither the model nor the tool but the transport intermediary between them, and AITBM localises it precisely — Cn-5 = 0.25 for an unverifiable response origin under…
An artifact-supply-chain incident with no agent, no prompt and no model output is still fully scorable — the evidence lands on Ro-4, Tr-3/Tr-4 and Cn-1/Cn-2 — and because the assessed configuration was…
Model pipeline
Indicative ERS4.8range 2.3–7.4
MVT evidenceMVT indeterminate
2026-04-23 · RAG / Retrieval-Augmented System · Tier 2
This is the batch's clearest demonstration that a control which held still gets scored, and scored up: CSP blocked the direct egress domain and the XPIA classifier forced the attacker to craft plain prose,…
The scoring turns on an evidentiary rule rather than a judgement call: the 0.50 Ro-1 anchor describes this case qualitatively — common attacks resisted, multi-step tool-mediated attacks still effective —…
AITBM scores a software supply-chain compromise as an AI-system finding without distorting either: the AI-specific signal lands on Ro-4 (the dependency ingestion path is a poisoning channel), Cn-5 (static…
A containment control can hold and the deployment still be compromised: the agent sandbox blocked in-runtime execution, which is why Cn-1 is scored up to 0.50 rather than down at a failure anchor, while the…
Cn-5 is the case's hinge and shows why a mechanism-only reading of the rubric is wrong: the gateway had real authentication (pairing, device tokens, a password), which is the 0.25 anchor, but the measured…
Identity, not the model, is the AI attack surface here: every low sub-metric sits in Containment (Cn-1, Cn-2, Cn-4, Cn-5, Cn-6 all at or near the floor) while Robustness barely features - and the case is…
Data exposureTool-calling LLM
Indicative ERS5.6range 3.2–7.9
MVT evidenceCritical MVT confirmed
2026-04-17 · RAG / Retrieval-Augmented System · Tier 3
The sub-metric that should have stopped this is Ro-4, not a web-application control: once the prompt and retrieval tables were writable, the AI platform's integrity depended on chunk and configuration…
ForcedLeak shows why Cp = 1.00 here is corroborated rather than defaulted: a sink gate that a researcher's payload was observed crossing cannot be claimed at CBR >= 0.95, so the path counts as ungated to a…
A case with no exploit still moves the AITBM score: a change to an agent's authority boundary trips the C_event <= 0.35 cap and, with mutable behavioural state, the M_Em = 3.0 behavioural staleness floor -…
Agentic / MCP
Indicative ERS6.1range 2.0–10.0
MVT evidenceMVT indeterminate
No matching case studiesTry a broader term or clear the active category.
Showing 1–10 of 56
Indicative ERS is a normalized-assurance comparison scenario
and must be read with its best-to-worst interval. Public ACI describes the article evidence package;
MVT reports only a severity floor proved across the uncertainty bounds.
Featured case: the Hugging Face intrusion
The July 2026 Hugging Face case is the fullest public record in
this set — a first-party technical timeline, a corroborating account from OpenAI, and an independent
analysis — which makes it the best demonstration of how AITBM consumes evidence. Its dedicated page
contains the complete calculation, source trail, and uncertainty analysis.
FEATURED CASE
Frontier Lab Agent Intrusion into Hugging Face: Technical Reconstruction and Defensive Priorities
A case can have eight or nine genuinely effective controls — an SSRF allowlist that never failed, 84 straight denials at the credential store, digest verification that kept the supply chain clean — and still sit at the ceiling of the ORP layer, because AITBM scores boundaries independently rather than crediting an incident for the boundaries that happened not to be on the attacker's path: one ungated escalation route from an entry-exposed workload to credential-issuing nodes forces Cp to 1.00 on its own merits, and no number of denials elsewhere reduces it.
8 of the 64 published analyses describe a
technique, a research result, or a survey of many deployments rather than one assessed system. AITBM
scores a deployment, not a threat, so forcing a score onto these would manufacture precision that the
evidence cannot support. They are listed here for completeness.
Subliminal Learning: Behavioural Traits Leak Through Semantically Unrelated Distillation Data(2026-04-16)ResearchThis brief reports a controlled research result about a training-time mechanism, not an incident against a deployed AI system. The teacher and student models were created by the researchers to demonstrate the effect; there is no victim deployment, no operator, no production configuration, and no attack against a running system. Consequently the three AITBM layers have no referent: the IVP would have to be scored against a research artefact rather than an assessed configuration, and every ORP dimension - autonomy, attack surface, cascade potential, remediation feasibility - would have to be invented, since a distillation pipeline demonstrated in a laboratory has no deployment context to score. The protocol's representative-configuration exception does not rescue it either: the brief studies a class of training pipeline, not a class of deployment, and it supplies no configuration facts (autonomy, exposure, downstream reach) from which a representative deployment could be reconstructed. Scoring it would manufacture numbers the evidence cannot support. The brief is nonetheless directly useful to AITBM as a scoping input, recorded in key_finding below.
Real Attackers Don’t Compute Gradients: Operational Threat Modeling for ML Security(2026-04-26)ResearchMethodology brief with no assessed deployment. The source is a research paper (Apruzzese, Anderson, Dambra, Freeman, Pierazzi, Roundy) arguing that adversarial-ML evaluation over-weights gradient-style model attacks relative to the cheaper system-level bypasses real attackers use. Its illustrative material — Facebook's abuse-fighting funnel, an unnamed commercial phishing detector, the MLSEC competition — is second-hand and generic: no named system in a stated configuration, no incident, no measured control outcome, and no evidence of which defences were present or absent at a point in time. Under protocol section 7 this is a methodology/population study, and the representative-configuration exception does not apply because the brief does not describe any single deployment in enough detail to reconstruct one. Scoring it would require inventing sub-metric evidence.
AID-D-005.001 AI System Log Generation & Collection → Cn-7, Tr-3
AID-H-001 Adversarial Robustness Training → Ro-1
System-Level Agent Defenses: Why Indirect Prompt Injection Needs Plan and Policy Boundaries(2026-05-18)ResearchThere is no assessed deployment. The brief summarises an academic position paper (arXiv 2603.30016, 'Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection') that argues for plan, policy, approval, execution, and feedback boundaries in general-purpose agents. It reports no incident, no victim system, no observed configuration, and no measurement of any deployment — every statement is an architectural proposal or a critique of benchmark methodology. Scoring it would require inventing a system's IVP, ORP, and ACI inputs from prescriptive text, which the protocol forbids. The brief is retained for its AIDEFEND mapping and for what its benchmark critique implies about AITBM's own evidence-quality machinery.
Sysdig marimo: When an LLM Agent Drives Post-Exploitation After RCE(2026-06-06)IncidentThere is no assessed AI system on the victim side. The compromised asset is a marimo notebook runtime — a browser-based interactive Python execution environment used for AI and data-science work — reached through CVE-2026-39987 on its /terminal/ws WebSocket. It contains no model, no retrieval layer, no tool-calling interface and no agent: nothing in the deployed victim system produces or consumes model output. The only AI agency in this incident belonged to the attacker's post-exploitation tooling, and protocol section 1 places the attacker out of scope. Consequently the AITBM architecture decision tree (spec 4.4) has no class that fits without misrepresentation: Q1 through Q5 are all NO, and the Q6 fall-through to Traditional ML / Classifier would be factually wrong for a code-execution surface with no model. Every IVP axis that gives AITBM its distinctive signal — Robustness, Fairness, Transparency's explainability and calibration sub-metrics, Privacy's model-behaviour sub-metrics — is unassessable, and what remains (network scope, credential hygiene, segmentation, telemetry) is conventional infrastructure security that AITBM does not exist to score. Forcing a score here would produce an ERS driven almost entirely by ORP with a Containment-only IVP, which would misrepresent both the incident and the framework. The correct AITBM treatment of this case is as an input to other assessments rather than as an assessment of its own: it is a spec 3.3.3.4 threat-override condition (a new exploited vulnerability with an active incident, M_threat = 1.50, Tf capped at 0.50 for exploited relevance) for any AI workload hosted on or reachable from such runtimes, and an exposure-management finding for AI development surfaces generally.
AIDEFEND evidence routes
AID-H-004.001 User & Privileged Access Management → Cn-5, Tr-3
Shadow AI and Enterprise Data Governance: What 22.4M Prompts Reveal(2026-06-24)ResearchThere is no assessed AI deployment. The brief is a cross-vendor population study — Harmonic's telemetry over 22,458,240 prompts and uploads across 665 tools, plus Cyberhaven, Verizon DBIR and IBM breach-cost figures — measuring how enterprise data flows out through employee use of unmanaged third-party AI accounts. The systems receiving the data (ChatGPT, Gemini, Claude, Copilot, Perplexity and 660 others) are not deployments the assessed organisation configures, and the brief supplies no evidence about any of their intrinsic security properties; it measures the volume and category of data leaving governed paths. AITBM scores a specific system in a specific configuration, and 'shadow AI' is precisely the condition of there being no assessed configuration to score. Forcing an architecture class, tier, IVP or ORP onto a 665-tool population would fabricate a system that does not exist. The findings belong instead as evidence inputs to the assessment of the organisation's *governed* AI paths — chiefly Pr-3 (data minimisation), Cn-1 (scope enforcement at the egress boundary), Tr-3 (audit trail completeness), Cn-5 (account and workload identity) and the ORP Attack Surface dimension.
AIDEFEND evidence routes
AID-M-001.004 AI Service & Embedded SaaS AI Discovery → Cn-5, Fa-3, Tr-4
AID-I-002.002 Secure External AI Service Connectivity → Cn-4
AID-D-005.001 AI System Log Generation & Collection → Cn-7, Tr-3
AID-D-005.002 AI Detection Rule Lifecycle, Delivery & Health → Cn-7, Tr-3
AID-DV-002 Honey Data, Decoy Artifacts & Canary Tokens for AI → Pr-1
hTAG's Browser-Agent Benchmark Shows Why Model Refusal Is Not Authorization(2026-08-05)ExerciseThe benchmark aggregates 20 scenarios across five browser-agent products and reports vendor-level completion counts without a complete prompt-and-trace corpus, exact reproducible builds, or repeated-run counts. It is useful evidence for authorization testing, but not one bounded deployment. A single score would average incompatible products, account states, tools, and transaction effects and would therefore manufacture a system that was never assessed.
Vector Databases Are Not Just Vectors: Orca Exposure Findings and Milvus Authentication Failures(2026-08-05)Validated ResearchThe source combines an Internet-wide exposure study, several unrelated vector-database deployments, credential validation against external SaaS systems, and two different Milvus vulnerabilities with separate preconditions and fixed versions. There is no single assessed database, architecture, dependency graph, or remediation state. Combining them into one score would conflate population prevalence, two software defects, and several organizations' configurations.
AIDEFEND evidence routes
AID-H-003.010 Deployed AI Software Vulnerability Remediation Lifecycle → Ro-4, Tr-4
AID-I-002.001 Internal AI Network Segmentation → Cn-4
AID-H-005.005 Embedding & Vector Store Confidentiality → Pr-1, Pr-4
AID-H-004.002 Service & API Authentication → Cn-5, Tr-3
AID-M-001.005 Public AI Endpoint & Agent-Service Exposure Discovery → Cn-5, Fa-3, Tr-4
IDEsaster Shows How Agent File Writes Can Activate Trusted IDE Features(2026-08-05)Validated ResearchThe source reports a vulnerability class spanning more than ten AI IDE products, over 30 findings, and 24 CVEs. It does not define one product version, workspace trust state, enabled native feature set, agent tool policy, or deployment boundary from which a single IVP/ORP/ACI result can be calculated. Selecting one representative configuration would require an assessor choice not made by the source, so the analysis is preserved with its AIDEFEND routes but carries no ERS.
MVT — Minimum Viability Threshold. It is the per-tier minimum each assessed IVP axis must meet. A breach sets an independent severity floor, so Critical MVT does not automatically mean Critical ERS.