Radware showed how hidden instructions in email or documents read through ChatGPT connectors could turn a task into data exfiltration. Fixed URLs encoded stolen text one character per request, Memory preserved instructions across chats, and a contact-harvesting branch spread the payload; Radware says OpenAI fixed it on December 16, 2025. Workspace administrators can disable unnecessary connectors or Memory and narrow or revoke grants, while platform implementers own raw-content isolation, Memory gates, and default-deny egress; value-level sink checks cover only explicit transfers retaining server-side content IDs.
ASSESSED SYSTEM
The controlled ChatGPT connector configuration reported by Radware, including connected email or document retrieval, cross-chat Memory, URL fetching, and the attacker's external collection service.
OUT OF SCOPE
Current ChatGPT behavior after the reported 2025-12-16 fix, connector-free workspaces, and any tenant whose grants, Memory, and egress posture were not part of the research.
Architecture: Agentic / MCP System (decision tree Q4) — The system retrieved external content, persisted model-selected Memory across chats, and could initiate network or messaging effects, which is mutable agentic behavior. Tier 1: Tier 1 because connector data, durable Memory, contact discovery, and outbound requests combine into a persistent sensitive-data and propagation boundary.
Documented attack or failure path
- A normal connector task delivers the instruction. The attacker hides a prompt in an email or document. When a later user request causes ChatGPT to retrieve that source, the model can treat the hidden text as authority.
- Static URLs carry the stolen value. The prompt supplies a fixed URL alphabet. The model selects one existing URL per character, and the resulting request sequence encodes sensitive text without constructing or modifying a URL.
- Memory turns one retrieval into persistence. A malicious source can instruct ChatGPT to save a recurring trigger and collected sensitive information in Memory, allowing the behavior to reappear in later chats.
- Contact theft enables propagation. One branch extracts recent email addresses. The attacker's server, not ChatGPT itself, then sends the malicious email to those recipients.
- This was controlled research. Radware says OpenAI fixed the reported issue on December 16. OpenAI later described source-sink analysis and Safe Url defenses, but public material does not establish that every persistence and propagation variant was eliminated.
Observed controls and bounded outcomes
Positive credit is given only where the record directly demonstrates a control operating. Recommended or merely presumed controls receive no positive scoring credit.
- OpenAI reported that the demonstrated issue was fixed on 2025-12-16; the case remains historical and does not score current product behavior.
- Connector and Memory features are administratively controllable, which gives defenders a concrete exposure-reduction and grant-revocation path.
Layer 1 — Intrinsic Vulnerability Profile
Each sub-metric is placed on its five-level rubric by the evidence quoted beside it. Missing applicable evidence remains unknown. The displayed midpoint and interval are scenario values, not inferred control performance.
| Sub-metric | Score | Rubric basis | Evidence |
|---|---|---|---|
| Robustness (Ro) — scenario interval 0.14–0.59 (midpoint 0.36), Tier 1 MVT 0.60 confirmed breach | |||
| Ro-1Adversarial Input Resistance | 0.25w 0.30 | The 0.25 anchor: the controlled attack bypassed the instruction boundary with ordinary retrieved content; no measured broad defense rate supports a higher placement. | Hidden instructions in an email or document were treated as authority during a later connector task and redirected the workflow into data exfiltration and persistence.source: primary/brief |
| Ro-4Poisoning Attack Resistance | 0.25w 0.25 | The 0.25 anchor: durable behavioral state accepted externally influenced content without a quarantined promotion gate or demonstrated provenance enforcement. | The malicious source could cause an action-oriented instruction to be written into cross-chat Memory and retrieved during later benign conversations.source: primary/brief |
| Fairness (Fa) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate | |||
| Transparency (Tr) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.55 indeterminate | |||
| Privacy (Pr) — scenario interval 0.09–0.74 (midpoint 0.41), Tier 1 MVT 0.60 indeterminate | |||
| Pr-3Data Minimization Compliance | 0.25w 0.35 | The 0.25 anchor: some tenant grant boundaries existed, but the demonstrated task exposed more durable state and outbound authority than the immediate read required. | Connector scope exposed email or document content and recent contact addresses to a workflow that could also select external URLs and retain data in Memory.source: primary/brief |
| Containment (Cn) — scenario interval 0.10–0.69 (midpoint 0.40), Tier 1 MVT 0.65 indeterminate | |||
| Cn-1Scope Enforcement | 0.25w 0.12 | The 0.25 anchor: connector grants bounded available services, but effective action scope was not re-derived from trusted intent at each step. | An inbox-reading task could acquire Memory-write, URL-open, and contact-harvesting effects from retrieved text rather than from the authenticated user's visible request.source: primary/brief |
| Cn-3Output Filtering Robustness | 0.25w 0.18 | The 0.25 anchor: URL protections existed but did not enforce information-flow policy over which approved URL was selected. | A fixed alphabet of prebuilt URLs encoded stolen text one request per character, bypassing defenses focused on model-constructed or modified destinations.source: primary/brief |
| Cn-6Action Reversibility Classification Rate | 0.25w 0.12 | The 0.25 anchor: user-level connector authorization existed, but no evidenced action taxonomy or single-use approval governed these consequential steps. | Memory promotion and outbound requests could occur without a content-bound independent approval tied to the exact retained instruction or external effect.source: primary/brief |
Unknown, not N/A: 17 applicable sub-metrics lack admissible public evidence. The lower, midpoint, and upper scenarios evaluate each at 0.00, 0.50, and 1.00 respectively; no weight is redistributed merely because evidence is missing.
Layer 2 — Operational Risk Posture
| Dimension | Score | Justification |
|---|---|---|
| AaAutonomy Amplification | 0.75w 0.35 | Placed from the demonstrated decision and action authority of the assessed boundary; public evidence supports this bounded level but not a broader claim about current product defaults. |
| AsAttack Surface Exposure | 1.00w 0.25 | Placed from who can supply the initiating content or protocol message and from the trust status of that source in the documented configuration. elevated |
| CpCascade Potential | 1.00w 0.25 | No verified System Dependency Graph with DGC at least 0.90 is public for this boundary, so the specification's worst-case graph default governs rather than an assessor-estimated blast radius.GDCP: worst-case default — no verified dependency graph published elevated |
| RfRemediation Feasibility | 0.50w 0.15 | Placed from the documented remediation class: deterministic package/configuration change where available, otherwise provider, architecture, or multi-layer changes. It does not assert fleet-wide closure. |
Nelevated = 2 (dimensions strictly above 0.75) → CRM = 1.15.
Compound Risk Alert. Two or more dimensions are simultaneously elevated (spec 3.2.2); architectural decomposition is recommended before deployment.
Layer 3 — Public-evidence confidence diagnostic
| Component | Score | Basis |
|---|---|---|
| Pc — Public provenance evidence | 0.35 | The public record identifies the affected product or representative configuration, attack path, and principal control boundaries, but does not provide a complete asset manifest, verified dependency graph, configuration export, or assessment evidence manifest. |
| Ec — Public evaluation coverage | 0.25 | coverage 0.26 (6 of 23 applicable sub-metrics) × independence 1.00 × fidelity 0.95. No Full, Standard, or Lite pathway is claimed for a retrospective article. |
| Tf — Public-evidence freshness | 0.00 | Evidence dated 2026-01-08; age 217 days on the workpaper reference date. Components: T_containment 0.00 · T_behavior 0.00 · T_calendar 0.01 · C_monitor 0.65 · C_event 0.65 · C_evidence 0.85. Binding term: T_behavior. Evidence age is measured from 2026-01-08 to the 2026-08-13 workpaper reference date. Public sources do not provide a passing containment or behavioral re-attestation receipt; event, monitoring, and unresolved-evidence caps remain diagnostic only. |
Public-evidence ACI = (Pc × Ec × Tf)1/3 = 0.00 — diagnostic status: Invalid as assessment-of-record evidence. It describes the evidence available to this case study, not the assurance of the underlying system, and is not inserted into the normalized-assurance scenario ERS.
Indicative ERS — normalized-assurance scenario
AIDEFEND defences → AITBM sub-metrics
Identifiers are quoted as they appear on the AIDEFEND in Action brief (retrieved 2026-08-13); the sub-metric mapping is AITBM's own, from the specification's AIDEFEND tables reconciled at catalogue data version 2026.08.05. AIDEFEND renumbers identifiers between releases, so the data version travels with every mapping and neither side's IDs should be cited without one. A mapping identifies a possible evidence route; a recommendation does not prove that the control was implemented or effective and receives no scoring credit by itself.
| Technique | Defence | Priority | Evidences |
|---|---|---|---|
| AID-H-017.007 | Dual-LLM Isolation PatternAIDEFEND dataVersion 2026.08.05. Send raw email, documents, and connector content only to a quarantined model with no privileged tools. A validating broker may release only allowlisted domain fields and opaque record identifiers in a signed envelope; a generic free-text summary is not a security boundary. The privileged model never receives the original attacker text, and every proposal still goes through the executor's authorization checks. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Very High | Cn-5 Cn-7 |
| AID-I-004.004 | Transactional Promotion Gates (Quarantine -> Trusted)AIDEFEND dataVersion 2026.08.05. Externally influenced content must not write directly into trusted Memory. Store proposed memories in quarantine with tenant, source, content digest, and target namespace, then require a content-bound approval and controlled-writer receipt before atomic promotion. Prompt assembly must exclude every pending or rejected state. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Very High | Cn-4 Cn-7 Pr-2 Pr-4 |
| AID-H-018.005 | Value-Level Capability Metadata & Data Flow Sink EnforcementAIDEFEND dataVersion 2026.08.05. Put URL fetches, remote subresources, and outbound mail behind default-deny provider egress with a change-managed destination allowlist. For explicit value transfers that already carry a server-side content identifier, apply the labelled sink policy at dispatch. This blocks the attacker domain in the demonstrated chain, but the current guidance does not automatically track the implicit flow in which a secret character selects one prebuilt URL. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | High | Cn-1 Cn-6 Cn-7 |
| AID-H-018.004 | Intent-Based Dynamic Capability ScopingAIDEFEND dataVersion 2026.08.05. Register connector read, Memory write, web open, and outbound message as distinct tool names. Derive the signed set of allowed names, action budget, and expiry from the authenticated user's visible request, so an inbox-reading task cannot acquire unrelated tools from retrieved text. This guidance does not constrain mailbox resources or result cardinality inside one broadly designed connector operation. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | High | Cn-1 Cn-6 Cn-7 |
| AID-D-001.005 | Recalled Memory Pre-Rehydration ScanningAIDEFEND dataVersion 2026.08.05. Before any saved Memory record re-enters prompt assembly, scan the exact versioned content for authority claims, delayed triggers, encoded fragments, and known quarantine fingerprints. Bind the finding to the tenant, record version, content digest, detector revision, and policy. This detector only recommends a disposition; a downstream Harden or Isolate policy must decide and enforce blocking or quarantine. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Medium | Ro-1 |
| AID-H-006.002 | Text, Markup & Structured Output Sanitization and Release GateAIDEFEND dataVersion 2026.08.05. For the renderer-based branch, hold generated output until remote subresources and URLs have passed sink-specific policy. Remove or proxy unapproved images and previews, disable automatic fetches, and verify with browser-egress tests that Markdown, HTML, CSS, SVG, and reference links cannot silently trigger the fixed-URL requests. This relationship is an evidence route only and supplies no positive score credit without observed control operation. | Medium | Cn-3 Ro-3 |
WHAT THIS CASE TEACHES
Persistent memory changes prompt injection from a one-chat event into delayed control flow; trusted intent must govern both memory promotion and every later rehydration.
Sources: AIDEFEND in Action — ZombieAgent Shows How Connector Prompt Injection Can Persist and Spread · Primary source — ZombieAgent: New ChatGPT Vulnerabilities Let Data Theft Continue (and Spread) · Supporting primary source cited by AIDEFEND