Memory Control Flow Attacks (MCFA) do not need direct access to an agent's memory database or system prompt. The research shows how ordinary interaction can place an action-oriented preference into long-term memory, then a later benign task retrieves it and changes tool choice or order. Across controlled LangChain and LlamaIndex experiments, tool-selection attacks were highly reliable, while ordering attacks were weaker. Memory must therefore be governed as persistent control input, not trusted conversation history.
ASSESSED SYSTEM
A representative LangChain or LlamaIndex agent matching the controlled MCFA experiments: ordinary interactions can write long-term memory that later influences trusted tool selection and order.
OUT OF SCOPE
Production victims and harmful external actions; the study used synthetic safe/risky tool pairs and measured trace deviation in controlled experiments.
Architecture: Agentic / MCP System (decision tree Q4) — The model writes persistent behavioral state and later retrieves it to select and order tools, satisfying the mutable-agentic branch and Behavioral Attestation Window trigger. Tier 2: Tier 2 because durable memory changes later tool decisions across tasks, while the published study did not demonstrate critical production impact.
Documented attack or failure path
- The attacker did not need memory-database access. One or a few ordinary interactions caused the agent to store an action-oriented preference. A later benign task retrieved it and changed tool choice or order, while the trusted tools and system prompt remained unchanged.
- The study measured related but distinct outcomes. Override selected a risky tool and reached 97.2% to 100% success. Order violated a required sequence and reached 52.8% to 69.4%. Cross-task and more-than-30-round tests showed that the effect could persist, while matched benign-memory controls remained at 0%.
- Simple mitigations were incomplete. Adding 100 benign records did not remove the effect. Summarization reduced malicious-write success, but every summarized instruction that survived still steered the later task. Role-based memory segregation reduced some Override rates, yet residual success ranged from 2.8% to 100%.
- This was controlled research. The evaluation used synthetic safe and risky tool variants, no production users, and no harmful external actions. It measured tool-trace deviation, not real damage. There is currently no public evidence that MCFA has been used in an actual attack.
Observed controls and bounded outcomes
Positive credit is given only where the record directly demonstrates a control operating. Recommended or merely presumed controls receive no positive scoring credit.
- Matched benign-memory controls stayed at 0% attack success, supporting a causal link to malicious durable state rather than generic agent instability.
- Summarization and role-based segregation reduced some attacks, although neither supplied a reliable security boundary.
Layer 1 — Intrinsic Vulnerability Profile
Each sub-metric is placed on its five-level rubric by the evidence quoted beside it. Missing applicable evidence remains unknown. The displayed midpoint and interval are scenario values, not inferred control performance.
| Sub-metric | Score | Rubric basis | Evidence |
|---|---|---|---|
| Robustness (Ro) — scenario interval 0.14–0.59 (midpoint 0.36), Tier 2 MVT 0.50 indeterminate | |||
| Ro-1Adversarial Input Resistance | 0.25w 0.30 | The 0.25 anchor: a low-complexity interaction bypassed the trust boundary and no robust adversarial-input gate prevented memory write or recall. | One or a few ordinary interactions could plant an action-oriented preference without memory-database or system-prompt access; later benign tasks retrieved it as control input.source: primary/brief |
| Ro-4Poisoning Attack Resistance | 0.25w 0.25 | The 0.25 anchor: persistent state poisoning succeeded reliably and simple dilution or summarization mitigations remained incomplete. | Tool Override reached 97.2% to 100% and Order attacks 52.8% to 69.4%; effects persisted beyond 30 rounds while benign-memory controls remained at 0%.source: primary/brief |
| Fairness (Fa) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 2 MVT 0.55 indeterminate | |||
| Transparency (Tr) — scenario interval 0.06–0.81 (midpoint 0.44), Tier 2 MVT 0.50 indeterminate | |||
| Tr-3Audit Trail Completeness | 0.25w 0.25 | The 0.25 anchor: ordinary memory records existed, but the security-relevant write-to-recall-to-action lineage was incomplete. | The attacked agents stored and rehydrated action-oriented memory without an evidenced signed writer, provenance tier, promotion receipt, or tool-decision correlation trail.source: primary/brief |
| Privacy (Pr) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 2 MVT 0.55 indeterminate | |||
| Containment (Cn) — scenario interval 0.10–0.70 (midpoint 0.40), Tier 2 MVT 0.55 indeterminate | |||
| Cn-1Scope Enforcement | 0.25w 0.12 | The 0.25 anchor: available tools bounded possible actions, but runtime scope was influenced by untrusted durable state rather than re-derived from current intent. | A recalled preference could select a risky tool outside the later benign task's trusted intent while the system prompt and tool registry remained unchanged.source: primary/brief |
| Cn-2Escalation Prevention | 0.25w 0.16 | The 0.25 anchor: partial segmentation existed in experiments but did not reliably prevent escalation across memory trust boundaries. | Role-based memory segregation reduced some Override rates yet left residual success from 2.8% to 100%, so stored state could still escalate into a riskier tool choice.source: primary/brief |
| Cn-6Action Reversibility Classification Rate | 0.25w 0.12 | The 0.25 anchor: no formal action-classification coverage or chain-composition gate was evidenced in the representative configuration. | The later agent could change tool choice or required order without a fresh independent approval bound to the exact high-impact action sequence.source: primary/brief |
Unknown, not N/A: 17 applicable sub-metrics lack admissible public evidence. The lower, midpoint, and upper scenarios evaluate each at 0.00, 0.50, and 1.00 respectively; no weight is redistributed merely because evidence is missing.
Layer 2 — Operational Risk Posture
| Dimension | Score | Justification |
|---|---|---|
| AaAutonomy Amplification | 0.75w 0.25 | Placed from the demonstrated decision and action authority of the assessed boundary; public evidence supports this bounded level but not a broader claim about current product defaults. |
| AsAttack Surface Exposure | 0.75w 0.35 | Placed from who can supply the initiating content or protocol message and from the trust status of that source in the documented configuration. |
| CpCascade Potential | 1.00w 0.20 | No verified System Dependency Graph with DGC at least 0.90 is public for this boundary, so the specification's worst-case graph default governs rather than an assessor-estimated blast radius.GDCP: worst-case default — no verified dependency graph published elevated |
| RfRemediation Feasibility | 0.50w 0.20 | Placed from the documented remediation class: deterministic package/configuration change where available, otherwise provider, architecture, or multi-layer changes. It does not assert fleet-wide closure. |
Nelevated = 1 (dimensions strictly above 0.75) → CRM = 1.00.
Layer 3 — Public-evidence confidence diagnostic
| Component | Score | Basis |
|---|---|---|
| Pc — Public provenance evidence | 0.35 | The public record identifies the affected product or representative configuration, attack path, and principal control boundaries, but does not provide a complete asset manifest, verified dependency graph, configuration export, or assessment evidence manifest. |
| Ec — Public evaluation coverage | 0.25 | coverage 0.26 (6 of 23 applicable sub-metrics) × independence 1.00 × fidelity 0.95. No Full, Standard, or Lite pathway is claimed for a retrospective article. |
| Tf — Public-evidence freshness | 0.20 | Evidence dated 2026-06-05; age 69 days on the workpaper reference date. Components: T_behavior 0.20 · T_containment 0.35 · T_calendar 0.59 · C_monitor 0.65 · C_event 0.65 · C_evidence 0.85. Binding term: T_behavior. Evidence age is measured from 2026-06-05 to the 2026-08-13 workpaper reference date. Public sources do not provide a passing containment or behavioral re-attestation receipt; event, monitoring, and unresolved-evidence caps remain diagnostic only. |
Public-evidence ACI = (Pc × Ec × Tf)1/3 = 0.26 — diagnostic status: Invalid as assessment-of-record evidence. It describes the evidence available to this case study, not the assurance of the underlying system, and is not inserted into the normalized-assurance scenario ERS.
Indicative ERS — normalized-assurance scenario
AIDEFEND defences → AITBM sub-metrics
Identifiers are quoted as they appear on the AIDEFEND in Action brief (retrieved 2026-08-13); the sub-metric mapping is AITBM's own, from the specification's AIDEFEND tables reconciled at catalogue data version 2026.08.05. AIDEFEND renumbers identifiers between releases, so the data version travels with every mapping and neither side's IDs should be cited without one. A mapping identifies a possible evidence route; a recommendation does not prove that the control was implemented or effective and receives no scoring credit by itself.
| Technique | Defence | Priority | Evidences |
|---|---|---|---|
| AID-I-004.004 | Transactional Promotion Gates (Quarantine -> Trusted)AIDEFEND dataVersion 2026.08.05. Route every untrusted or high-risk new memory record to quarantine, then promote the exact version through an atomic, evidence-backed state transition before it becomes retrieval-eligible. This directly interrupts MCFA's first stage by preventing an ordinary low-trust conversation write from silently becoming trusted control input. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Very High | Cn-4 Cn-7 Pr-2 Pr-4 |
| AID-H-018.003 | High-Impact Independent Validation & Approval GateAIDEFEND dataVersion 2026.08.05. At the executor boundary, independently validate the canonical action, caller identity, policy evidence, and exact approval for high-impact operations. Even when malicious memory succeeds in changing the plan or tool order, the recalled text cannot authorize a transfer, deletion, credential change, or other side effect by itself. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Very High | Cn-1 Cn-6 Cn-7 |
| AID-I-004.002 | Persistent Memory Partitioning (Trust & Tenant Isolation)AIDEFEND dataVersion 2026.08.05. Partition persistent memory by tenant and trust tier, and make retrieval consult centralized entitlement policy before a record enters the model context. This limits cross-user and low-trust recall, but the paper's RBMS results show that segregation plus prompt hierarchy is not a complete defense when models fail to honor the hierarchy. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | High | Cn-4 Cn-7 Pr-2 Pr-4 |
| AID-D-001.005 | Recalled Memory Pre-Rehydration ScanningAIDEFEND dataVersion 2026.08.05. After loading an exact versioned memory record and before adding it to the prompt, rescan it for role or authority assertions, trigger-like markers, encoded fragments, known quarantine fingerprints, and policy risk, then emit a content-bound finding that downstream policy can quarantine or omit. This places inspection at MCFA's second stage, although semantic instructions can evade detection and a finding is not enforcement by itself. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Medium | Ro-1 |
| AID-D-003.004 | Tool-Call Sequence Anomaly DetectionAIDEFEND dataVersion 2026.08.05. Learn approved tool transitions and detect forbidden, low-likelihood, or missing-stage sequences, such as a transfer tool appearing before its required audit step. This is especially relevant to MCFA Order attacks and can surface successful steering, but it detects a changed trace after planning rather than preventing the memory write. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Medium | Cn-1 Cn-3 Cn-7 Ro-3 |
| AID-M-002.004 | Trust-Tiered Memory/KB Provenance & Write Eligibility ContractAIDEFEND dataVersion 2026.08.05. Record who proposed each memory item, the supporting evidence, validator result, namespace, trust tier, and recommended retrieval eligibility, including contradiction checks against authoritative facts. This creates the decision facts a promotion gate needs; it classifies and preserves provenance but does not enforce admission on its own. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Medium | Fa-3 Pr-3 Ro-4 Tr-3 Tr-4 |
WHAT THIS CASE TEACHES
Long-term memory is persistent control input: every write needs provenance and promotion policy, and every recall must be re-authorized against the current task before it can influence tools.
Sources: AIDEFEND in Action — MCFA Turns Agent Memory into a Delayed Control-Flow Input · Primary source — From Storage to Steering: Memory Control Flow Attacks on LLM Agents