PUBLIC-EVIDENCE AI SECURITY CASE STUDY

Poisoned GGUF Chat Templates Backdoor Inference Without Changing Model Weights

Researchers from Pillar Security and Fujitsu showed that a GGUF artifact can carry a modified Jinja chat template while its model weights remain unchanged. When a trigger appears, the inference engine renders attacker instructions into system-level context before the model runs; controlled tests degraded model behavior and hijacked BrowserUse and OpenHands tool use. Defenders need to treat chat templates as executable release artifacts, verify their exact bytes and provenance, regression-test every conversion, and enforce tool and data-flow boundaries.

Agentic / MCP SystemTier 1Indicative ERS 4.8 (2.0–7.7)Evidence source date 2026-08-02

Researchers from Pillar Security and Fujitsu showed that a GGUF artifact can carry a modified Jinja chat template while its model weights remain unchanged. When a trigger appears, the inference engine renders attacker instructions into system-level context before the model runs; controlled tests degraded model behavior and hijacked BrowserUse and OpenHands tool use. Defenders need to treat chat templates as executable release artifacts, verify their exact bytes and provenance, regression-test every conversion, and enforce tool and data-flow boundaries.

ASSESSED SYSTEM

A representative BrowserUse or OpenHands agent using a GGUF model artifact whose weights are legitimate but whose bundled Jinja chat template contains trigger-conditioned attacker instructions.

OUT OF SCOPE

Clean digest-pinned GGUF artifacts, deployments without tools or sensitive data, and real-world victims; the reported demonstrations used synthetic data and researcher-controlled infrastructure.

Architecture: Agentic / MCP System (decision tree Q4) — The controlled demonstrations combined poisoned inference context with browser or coding tools that produced external actions. Tier 1: Tier 1 because a release artifact can silently redirect payment, personal-data, credential, or code-writing actions at inference time.

Documented attack or failure path

  1. The attacker changes the template, not the weights. A legitimate open-weight model is repackaged with malicious Jinja logic in its GGUF chat template and redistributed as a plausible model artifact.
  2. The inference engine activates the backdoor. On every request, the engine renders the bundled template before tokenization. A matching phrase or field lets the template insert attacker instructions into system-level context.
  3. The behavior remained dormant until triggered. Across 18 models, seven families, and four inference engines, the paper measured severe triggered accuracy loss and reliable attacker-link emission while preserving normal behavior outside the trigger.
  4. Agent tools turned model manipulation into system impact. Controlled BrowserUse and OpenHands demonstrations redirected synthetic payment and personal data, inserted attacker code, and exposed test credentials.
  5. This was controlled research. The BrowserUse demonstration used a mock storefront, synthetic data, and a loopback receiver; the OpenHands demonstration used researcher-controlled infrastructure and test credentials. The sources do not report malicious exploitation against real users.

Observed controls and bounded outcomes

Positive credit is given only where the record directly demonstrates a control operating. Recommended or merely presumed controls receive no positive scoring credit.

  • The backdoor remained dormant outside the trigger, preserving ordinary behavior; this bounded outcome also explains why ordinary regression alone was insufficient.
  • The demonstrations used mock or researcher-controlled infrastructure and synthetic data, so no victim loss is inferred.

Layer 1 — Intrinsic Vulnerability Profile

Each sub-metric is placed on its five-level rubric by the evidence quoted beside it. Missing applicable evidence remains unknown. The displayed midpoint and interval are scenario values, not inferred control performance.

Sub-metricScoreRubric basisEvidence
Robustness (Ro) — scenario interval 0.00–0.75 (midpoint 0.38), Tier 1 MVT 0.60 indeterminate
Ro-4Poisoning Attack Resistance0.00w 0.25The 0.00 anchor: artifact poisoning was accepted and activated without an effective complete-artifact integrity or conversion-regression gate.A legitimate open-weight model was repackaged with malicious Jinja logic; inference engines rendered attacker instructions into system-level context without changing the weights.source: primary/brief
Fairness (Fa) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Transparency (Tr) — scenario interval 0.05–0.85 (midpoint 0.45), Tier 1 MVT 0.55 indeterminate
Tr-4Model Lineage Disclosure0.25w 0.20The 0.25 anchor: partial model provenance existed, but the full GGUF bytes, template, conversion settings, and loader identity were not bound together.An unchanged weight hash could coexist with a modified bundled template, so weight-only lineage did not identify the executable release artifact actually loaded.source: primary/brief
Privacy (Pr) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Containment (Cn) — scenario interval 0.07–0.78 (midpoint 0.42), Tier 1 MVT 0.65 indeterminate
Cn-1Scope Enforcement0.25w 0.12The 0.25 anchor: tool surfaces existed within controlled environments, but the model-originated data and action scope was not deterministically constrained at dispatch.Triggered BrowserUse and OpenHands tests redirected synthetic payment and personal data, inserted attacker code, and exposed test credentials through ordinary agent tools.source: primary/brief
Cn-3Output Filtering Robustness0.25w 0.18The 0.25 anchor: basic model behavior remained normal off-trigger, but adaptive unsafe outputs and tool-laundered effects were not reliably blocked.The trigger-conditioned template injected system-level instructions that survived into tool use and external output without an independent sink or release gate.source: primary/brief

Unknown, not N/A: 19 applicable sub-metrics lack admissible public evidence. The lower, midpoint, and upper scenarios evaluate each at 0.00, 0.50, and 1.00 respectively; no weight is redistributed merely because evidence is missing.

Layer 2 — Operational Risk Posture

DimensionScoreJustification
AaAutonomy Amplification0.75w 0.35Placed from the demonstrated decision and action authority of the assessed boundary; public evidence supports this bounded level but not a broader claim about current product defaults.
AsAttack Surface Exposure0.75w 0.25Placed from who can supply the initiating content or protocol message and from the trust status of that source in the documented configuration.
CpCascade Potential1.00w 0.25No verified System Dependency Graph with DGC at least 0.90 is public for this boundary, so the specification's worst-case graph default governs rather than an assessor-estimated blast radius.GDCP: worst-case default — no verified dependency graph published elevated
RfRemediation Feasibility0.50w 0.15Placed from the documented remediation class: deterministic package/configuration change where available, otherwise provider, architecture, or multi-layer changes. It does not assert fleet-wide closure.

Nelevated = 1 (dimensions strictly above 0.75) → CRM = 1.00.

Layer 3 — Public-evidence confidence diagnostic

ComponentScoreBasis
Pc — Public provenance evidence0.35The public record identifies the affected product or representative configuration, attack path, and principal control boundaries, but does not provide a complete asset manifest, verified dependency graph, configuration export, or assessment evidence manifest.
Ec — Public evaluation coverage0.17coverage 0.17 (4 of 23 applicable sub-metrics) × independence 1.00 × fidelity 0.95. No Full, Standard, or Lite pathway is claimed for a retrospective article.
Tf — Public-evidence freshness0.00Evidence dated 2026-02-04; age 190 days on the workpaper reference date. Components: T_containment 0.00 · T_calendar 0.01 · C_monitor 0.65 · C_event 0.65 · C_evidence 0.85. Binding term: T_containment. Evidence age is measured from 2026-02-04 to the 2026-08-13 workpaper reference date. Public sources do not provide a passing containment or behavioral re-attestation receipt; event, monitoring, and unresolved-evidence caps remain diagnostic only.

Public-evidence ACI = (Pc × Ec × Tf)1/3 = 0.02 — diagnostic status: Invalid as assessment-of-record evidence. It describes the evidence available to this case study, not the assurance of the underlying system, and is not inserted into the normalized-assurance scenario ERS.

Indicative ERS — normalized-assurance scenario

Worp · ORP0.35(0.75) + 0.25(0.75) + 0.25(1.00) + 0.15(0.50) = 0.775
CRMNelevated = 1 → 1.00
ORPeffective0.775 × 1.00 = 0.775
Wivp · IVP midpoint0.30(0.38) + 0.25(0.50) + 0.15(0.45) + 0.20(0.50) + 0.10(0.42) = 0.448
IVP mitigation0.15 + 0.85(1 − 0.448) = 0.620
Scenario assuranceACI fixed at 1.000 for cross-case comparison; public-evidence ACI 0.021 is diagnostic only
Indicative ERS midpointmin(10, 0.775 × 0.620 × 1/1.000 × 10) = 4.8
Unknown-input interval2.0–7.7; 19 unknown applicable sub-metrics set to 1.00 / 0.00 at the bounds

AIDEFEND defences → AITBM sub-metrics

Identifiers are quoted as they appear on the AIDEFEND in Action brief (retrieved 2026-08-13); the sub-metric mapping is AITBM's own, from the specification's AIDEFEND tables reconciled at catalogue data version 2026.08.05. AIDEFEND renumbers identifiers between releases, so the data version travels with every mapping and neither side's IDs should be cited without one. A mapping identifies a possible evidence route; a recommendation does not prove that the control was implemented or effective and receives no scoring credit by itself.

TechniqueDefence PriorityEvidences
AID-H-007.006Post-Training Optimization & Format-Conversion Safety RegressionAIDEFEND dataVersion 2026.08.05. Treat every GGUF conversion or repackaging step as a security-relevant transformation. Bind the candidate to approved source bytes and conversion settings, then compare both artifacts on the same signed suite for trigger-conditioned behavior, instruction-hierarchy changes, leakage, and tool-policy violations. Block promotion when lineage is missing or any security category regresses. This relationship is an evidence route only and supplies no positive score credit without observed control operation.Very HighRo-4
AID-H-003.002CI/CD Release Gating, Model Artifact Signing & Secure DistributionAIDEFEND dataVersion 2026.08.05. Production should load only reviewed, signed, digest-pinned GGUF bytes from an internal mirror. The promotion gate should verify the artifact source, complete digest, bundled template, loader policy, regression evidence, and approval record before deployment or hot reload. A public model name or unchanged weight hash is not sufficient admission evidence. This relationship is an evidence route only and supplies no positive score credit without observed control operation.Very HighRo-4 Tr-4
AID-H-003.006Model SBOM & Provenance AttestationAIDEFEND dataVersion 2026.08.05. Bind the complete GGUF digest, model format, source, loader commit, tokenizer, and configuration evidence into a signed model SBOM and attestation. Verify the signature, predicate, signer identity, and actual artifact digest before admission and every load. This proves that the runtime received the approved bytes; it does not determine whether the bundled Jinja logic is semantically safe. This relationship is an evidence route only and supplies no positive score credit without observed control operation.HighRo-4 Tr-4
AID-H-018.005Value-Level Capability Metadata & Data Flow Sink EnforcementAIDEFEND dataVersion 2026.08.05. For BrowserUse or another live action that actually passes through the governed dispatcher, require a server-side content identifier for payment fields, personal data, credentials, or repository content. Deny labelled content at attacker-controlled network, write, or file sinks, and keep agent egress default-deny. This control does not govern an OpenHands-generated script after the code leaves that dispatcher. This relationship is an evidence route only and supplies no positive score credit without observed control operation.MediumCn-1 Cn-6 Cn-7

WHAT THIS CASE TEACHES

Model weights are only one executable release component; chat templates must be digest-bound, provenance-checked, and security-regression-tested with the same rigor as the weights.

Sources: AIDEFEND in Action — Poisoned GGUF Chat Templates Backdoor Inference Without Changing Model Weights · Primary source — Inference-Time Backdoors via Chat Templates: From LLM Supply Chains to Agentic System Compromise · Supporting primary source cited by AIDEFEND

AITBM sub-metrics referenced