PUBLIC-EVIDENCE TOPIC COLLECTION

AI Model Security Cases: Loading, Inference, and Backdoors

AI model security cases covering malicious model loading, inference backdoors, tokenizer blind spots, exposed inference, namespace reuse, and distillation.

10 matched references · all link to complete source and scoring records

Why this topic matters

Model security extends beyond adversarial examples. Serialization formats, chat templates, tokenizers, serving messages, namespaces, exposed compute, and distillation pipelines can all change behavior or open a code-execution boundary.

Evaluation questions

  • Are model artifacts and executable serialization paths verified before loading?
  • Can templates, tokenizers, adapters, or serving messages alter behavior outside weight review?
  • Are model names, revisions, registries, and lineage cryptographically bound and auditable?
  • Do tests cover extraction, inference abuse, behavioral transfer, and monitoring blind spots?

Framework routes: MITRE ATLAS · OWASP AISVS. External frameworks guide threat, control, and evidence selection; measured deployment evidence determines AITBM scores.

AI model security case library

Use these records as incident references, test-design inputs, and examples of evidence-to-rubric reasoning. They are retrospective scenarios, not current vendor ratings or assessments of record.

2026-07-06 · Agentic / MCP System

JADEPUFFER Shows Agentic Ransomware Moving From AI RCE to Database Extortion

JADEPUFFER is the case that forces the assessor to keep the attacker out of the assessed system: the agentic behaviour that compressed the kill chain to 31-second self-repair belonged to the offence, so it drives no ORP or BAW score here — what AITBM scores is a Tier 1 workflow host with Cn-1, Cn-2 and Cn-5 all at 0.00 and a corroborated Cp = 1.00 traced from an unauthenticated origin to a P4 backdoor-administrator terminal.

Indicative ERS 6.9 · 3.0–10.0

2026-07-02 · Standalone LLM / Generative AI

Exposed LLM Backends Are Becoming Attacker AI Compute

Architecture class is scored on the assessed boundary, not on the attacker's behaviour: the exposed backend is a Standalone LLM with Aa = 0.25 and baw/agentic both false, yet three Containment sub-metrics sit at 0.00 and Cp is corroborated at 1.00 — an unauthenticated endpoint can be maximally exposed and trivially remediable (Rf = 0.00) at the same time, and the thin two-axis evidence base correctly forces the Lite Ec cap.

Indicative ERS 4.3 · 1.9–6.8

2026-06-27 · Multi-Agent / MCP System

Lingua Ex Machina: When the AI Monitor Cannot See What the Executor Sees

Detection capability and evidence freshness are scored in different layers, and this case separates them cleanly: the monitor's blindness is a point-in-time IVP finding at Ro-1 and Cn-3, while the same sensor loss independently caps ACI Temporal Freshness through C_monitor and Band 0 C_behavior — and because the report measured visibility rather than consequence, Cn-1 and Cn-6 are correctly omitted instead of guessed.

Indicative ERS 5.0 · 2.0–8.0

2026-05-26 · RAG / Retrieval-Augmented System

ChromaToast: ChromaDB Pre-Auth RCE Through Malicious Hugging Face Model Loading

Remediation Feasibility sits at its bottom anchor — a patchable ordering bug with a CVE and a fixed release — while Cascade Potential still forces 1.00, which is exactly the layer separation AITBM is built for: how easily a defect is fixed and how far it reaches are scored independently, so a cleanly patchable flaw never gets to look harmless.

Indicative ERS 4.8 · 2.7–7.0

2026-05-15 · Agentic / MCP System

AI App Misconfigurations: Public Agent Endpoints as RCE and Credential-Leak Paths

Nothing in this case reaches the model — there is no prompt, no injection, and no Robustness-1 signal at all — so AITBM's not-applicable redistribution rule carries the whole assessment on Containment, identity, and posture evidence, which is the correct answer for an incident where, as the brief puts it, security was lost before the model saw any prompt.

Indicative ERS 7.2 · 3.4–10.0

2026-04-24 · Traditional ML / Classifier

Malicious Hugging Face Models: When Loading a Model Opens a Backdoor

An artifact-supply-chain incident with no agent, no prompt and no model output is still fully scorable — the evidence lands on Ro-4, Tr-3/Tr-4 and Cn-1/Cn-2 — and because the assessed configuration was observed in February 2024, Temporal Freshness rather than the attack itself is what collapses assurance confidence, which is the temporal layer behaving exactly as designed.

Indicative ERS 4.8 · 2.3–7.4

2026-04-16 · Research

Subliminal Learning: Behavioural Traits Leak Through Semantically Unrelated Distillation Data

This brief reports a controlled research result about a training-time mechanism, not an incident against a deployed AI system. The teacher and student models were created by the researchers to demonstrate the effect; there is no victim deployment, no operator, no production configuration, and no attack against a running system. Consequently the three AITBM layers have no referent: the IVP would have to be scored against a research artefact rather than an assessed configuration, and every ORP dimension - autonomy, attack surface, cascade potential, remediation feasibility - would have to be invented, since a distillation pipeline demonstrated in a laboratory has no deployment context to score. The protocol's representative-configuration exception does not rescue it either: the brief studies a class of training pipeline, not a class of deployment, and it supplies no configuration facts (autonomy, exposure, downstream reach) from which a representative deployment could be reconstructed. Scoring it would manufacture numbers the evidence cannot support. The brief is nonetheless directly useful to AITBM as a scoping input, recorded in key_finding below.

Research note · no ERS

Use the evidence, not just the incident name

Each record distinguishes observed controls, missing evidence, architecture-based exclusions, operational risk, confidence limits, and the uncertainty interval. When using a case as a reference, compare the documented path to your own system boundary and rerun the applicable test methods rather than copying its indicative ERS.