PUBLIC-EVIDENCE AI SECURITY CASE STUDY

Stolen Thoughts: Opaque Reasoning Blocks Could Be Replayed to Recover Secrets

Stolen Thoughts found that opaque reasoning blocks returned by Anthropic, OpenAI, and Google APIs could be accepted across sessions, users, or compatible models within a provider family. An attacker who first obtained another user's block could submit it to a weaker compatible model and jailbreak that model into transcribing hidden reasoning. In 6,708 public agent trajectories, the researchers reconstructed 315,320 blocks and found real sensitive artifacts in 328 sessions. They report that provider mitigations later made the demonstrated attacks non-reproducible.

Agentic / MCP SystemTier 1Indicative ERS 3.5 (1.5–5.4)Evidence source date 2026-08-12

Stolen Thoughts found that opaque reasoning blocks returned by Anthropic, OpenAI, and Google APIs could be accepted across sessions, users, or compatible models within a provider family. An attacker who first obtained another user's block could submit it to a weaker compatible model and jailbreak that model into transcribing hidden reasoning. In 6,708 public agent trajectories, the researchers reconstructed 315,320 blocks and found real sensitive artifacts in 328 sessions. They report that provider mitigations later made the demonstrated attacks non-reproducible.

ASSESSED SYSTEM

A representative provider-family agent API integration that returned opaque reasoning blocks to clients, accepted captured blocks across sessions or users, and allowed a compatible weaker model to continue them.

OUT OF SCOPE

Arbitrary tenant access, which the paper did not demonstrate; present provider behavior after mitigation; and claims that every reconstructed block exactly matched unavailable ground-truth plaintext.

Architecture: Agentic / MCP System (decision tree Q4) — The evidence corpus comprised agent trajectories carrying provider state across tool-driven sessions; the scored boundary is the state-bearing integration rather than the foundation model alone. Tier 1: Tier 1 because cross-session reasoning state can contain credentials, personal data, proprietary instructions, and tool context whose disclosure affects production identities and systems.

Documented attack or failure path

  1. The attacker first needed a reasoning block. The work does not show arbitrary access to another tenant. The third-party path begins when published or otherwise obtainable agent logs expose the block.
  2. The weakness was portability, not broken encryption. Providers accepted opaque blocks across sessions, users, or compatible models within one provider family. After processing a captured block, a weaker model could be jailbroken into transcribing hidden reasoning.
  3. Public traces contained real secrets, with limits. The team reconstructed 315,320 blocks from 6,708 trajectories and found real sensitive artifacts in 328 sessions. Sixty-four of 704 genuine-session artifacts were absent from visible chat history. Other results included synthetic benchmark personas, and exact transcription fidelity could not always be proven without ground-truth plaintext.
  4. The reported attack is now mitigated. Anthropic, OpenAI, and Google acknowledged the disclosures. The researchers say provider changes made the Figure 1 attacks non-reproducible, though portable hidden state and public trace publication remain recurring architectural boundaries.

Observed controls and bounded outcomes

Positive credit is given only where the record directly demonstrates a control operating. Recommended or merely presumed controls receive no positive scoring credit.

  • The attacker first needed to obtain a reasoning block; the research did not demonstrate arbitrary access to another tenant.
  • Anthropic, OpenAI, and Google acknowledged the disclosures, and the researchers report that provider mitigations made the demonstrated attacks non-reproducible.

Layer 1 — Intrinsic Vulnerability Profile

Each sub-metric is placed on its five-level rubric by the evidence quoted beside it. Missing applicable evidence remains unknown. The displayed midpoint and interval are scenario values, not inferred control performance.

Sub-metricScoreRubric basisEvidence
Robustness (Ro) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Fairness (Fa) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Transparency (Tr) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.55 indeterminate
Privacy (Pr) — scenario interval 0.20–0.40 (midpoint 0.30), Tier 1 MVT 0.60 confirmed breach
Pr-1Training Data Leakage Risk0.25w 0.25The 0.25 anchor: sensitive context could be extracted through an adaptive replay and jailbreak chain; provider mitigation later blocked the demonstrated route but no measured enduring leakage bound is public.Researchers reconstructed 315,320 blocks from 6,708 public agent trajectories and found real sensitive artifacts in 328 sessions, including 64 artifacts absent from visible chat history.source: primary/brief
Pr-2Inference Attack Resistance0.25w 0.20The 0.25 anchor: the inference interface accepted a cross-context state artifact and disclosed information under adaptive probing without strong session binding.A captured block could be submitted to a compatible weaker model and jailbroken into transcribing hidden reasoning across sessions or users within a provider family.source: primary/brief
Pr-3Data Minimization Compliance0.25w 0.35The 0.25 anchor: operational context was retained and exported beyond the minimum visible response, with incomplete secret demotion and trace minimization.Credentials and personal data entered hidden reasoning and were preserved in public agent traces even when some artifacts were absent from visible chat history.source: primary/brief
Containment (Cn) — scenario interval 0.06–0.82 (midpoint 0.44), Tier 1 MVT 0.65 indeterminate
Cn-4Side-Channel Resistance0.25w 0.10The 0.25 anchor: provider opacity obscured the content but did not prevent the state artifact from carrying information across the intended session boundary.Opaque reasoning state became an unintended cross-session disclosure channel once an attacker obtained a block and replayed it through a compatible model.source: primary/brief
Cn-5Agent Identity Integrity0.25w 0.14The 0.25 anchor: provider-family compatibility existed, but identity and session binding of the state artifact was incomplete.Providers accepted opaque blocks across sessions, users, or compatible models without an evidenced cryptographic binding to the originating principal and conversation.source: primary/brief

Unknown, not N/A: 18 applicable sub-metrics lack admissible public evidence. The lower, midpoint, and upper scenarios evaluate each at 0.00, 0.50, and 1.00 respectively; no weight is redistributed merely because evidence is missing.

Layer 2 — Operational Risk Posture

DimensionScoreJustification
AaAutonomy Amplification0.25w 0.35Placed from the demonstrated decision and action authority of the assessed boundary; public evidence supports this bounded level but not a broader claim about current product defaults.
AsAttack Surface Exposure0.75w 0.25Placed from who can supply the initiating content or protocol message and from the trust status of that source in the documented configuration.
CpCascade Potential1.00w 0.25No verified System Dependency Graph with DGC at least 0.90 is public for this boundary, so the specification's worst-case graph default governs rather than an assessor-estimated blast radius.GDCP: worst-case default — no verified dependency graph published elevated
RfRemediation Feasibility0.25w 0.15Placed from the documented remediation class: deterministic package/configuration change where available, otherwise provider, architecture, or multi-layer changes. It does not assert fleet-wide closure.

Nelevated = 1 (dimensions strictly above 0.75) → CRM = 1.00.

Layer 3 — Public-evidence confidence diagnostic

ComponentScoreBasis
Pc — Public provenance evidence0.35The public record identifies the affected product or representative configuration, attack path, and principal control boundaries, but does not provide a complete asset manifest, verified dependency graph, configuration export, or assessment evidence manifest.
Ec — Public evaluation coverage0.21coverage 0.22 (5 of 23 applicable sub-metrics) × independence 1.00 × fidelity 0.95. No Full, Standard, or Lite pathway is claimed for a retrospective article.
Tf — Public-evidence freshness0.65Evidence dated 2026-08-10; age 3 days on the workpaper reference date. Components: C_monitor 0.65 · C_event 0.65 · C_evidence 0.85 · T_containment 0.87 · T_calendar 0.93. Binding term: C_monitor. Evidence age is measured from 2026-08-10 to the 2026-08-13 workpaper reference date. Public sources do not provide a passing containment or behavioral re-attestation receipt; event, monitoring, and unresolved-evidence caps remain diagnostic only.

Public-evidence ACI = (Pc × Ec × Tf)1/3 = 0.36 — diagnostic status: Critical evidence limitation. It describes the evidence available to this case study, not the assurance of the underlying system, and is not inserted into the normalized-assurance scenario ERS.

Indicative ERS — normalized-assurance scenario

Worp · ORP0.35(0.25) + 0.25(0.75) + 0.25(1.00) + 0.15(0.25) = 0.562
CRMNelevated = 1 → 1.00
ORPeffective0.562 × 1.00 = 0.562
Wivp · IVP midpoint0.30(0.50) + 0.25(0.50) + 0.15(0.50) + 0.20(0.30) + 0.10(0.44) = 0.454
IVP mitigation0.15 + 0.85(1 − 0.454) = 0.614
Scenario assuranceACI fixed at 1.000 for cross-case comparison; public-evidence ACI 0.361 is diagnostic only
Indicative ERS midpointmin(10, 0.562 × 0.614 × 1/1.000 × 10) = 3.5
Unknown-input interval1.5–5.4; 18 unknown applicable sub-metrics set to 1.00 / 0.00 at the bounds

AIDEFEND defences → AITBM sub-metrics

Identifiers are quoted as they appear on the AIDEFEND in Action brief (retrieved 2026-08-13); the sub-metric mapping is AITBM's own, from the specification's AIDEFEND tables reconciled at catalogue data version 2026.08.05. AIDEFEND renumbers identifiers between releases, so the data version travels with every mapping and neither side's IDs should be cited without one. A mapping identifies a possible evidence route; a recommendation does not prove that the control was implemented or effective and receives no scoring credit by itself.

TechniqueDefence PriorityEvidences
AID-H-037.001Reasoning-Trace Confidentiality & Storage ControlsAIDEFEND dataVersion 2026.08.05. Keep full reasoning blocks in provider-side or protected server-side state, classify them as sensitive, and deny their display, client return, analytics export, and public logging. Publish only bounded summaries or non-reversible references. This removes the portable artifact that the third-party attack must first acquire. This relationship is an evidence route only and supplies no positive score credit without observed control operation.Very HighCn-3 Cn-4 Cn-7
AID-H-037.003Provider Reasoning-Block Round-Trip IntegrityAIDEFEND dataVersion 2026.08.05. Preserve provider-authored reasoning blocks byte-for-byte in a server-held envelope, return them only through the same provider route and conversation, verify provider binding where supported, and reject client-supplied or altered blocks before continuation. This blocks the captured-block replay path for an integration that enforces the boundary, but it does not repair a provider API that accepts the same block through a direct path outside that integration. This relationship is an evidence route only and supplies no positive score credit without observed control operation.Very HighCn-3 Cn-4 Cn-7
AID-I-004.007Task-Bounded Context Segmentation & Secret DemotionAIDEFEND dataVersion 2026.08.05. Keep credentials and other high-sensitivity values outside conversation and model context. Give the agent short-lived, task-scoped handles that an authorized broker resolves only at the approved tool or network sink. If the underlying value never enters model context, it cannot be captured inside a reasoning block; this does not protect proprietary reasoning that legitimately remains there. This relationship is an evidence route only and supplies no positive score credit without observed control operation.HighCn-4 Cn-7 Pr-2 Pr-4
AID-M-008Automated Agentic Security BenchmarkingAIDEFEND dataVersion 2026.08.05. Turn the paper's captured-block, compatible-model, and transcription-jailbreak cases into a signed, sanitized regression corpus. Block each model or API release unless it rejects the replay path and prevents sensitive plaintext from being released. This tests the demonstrated bypass as models and routing change; it does not prove universal jailbreak resistance. This relationship is an evidence route only and supplies no positive score credit without observed control operation.HighRo-2
AID-D-003.002Sensitive Information & Data Leakage DetectionAIDEFEND dataVersion 2026.08.05. Inspect the complete decoder response before delivery. Detect common credential formats with deterministic patterns and use structured PII detection for names, addresses, email addresses, and other unstructured personal data, then emit findings to the release gate. Detection identifies recognizable sensitive content but does not stop delivery by itself. This relationship is an evidence route only and supplies no positive score credit without observed control operation.HighCn-1 Cn-3 Cn-7 Ro-3
AID-H-006.002Text, Markup & Structured Output Sanitization and Release GateAIDEFEND dataVersion 2026.08.05. Hold the complete decoder response or a bounded release unit until sensitive-output findings have been evaluated. Apply deterministic redaction or denial, and fail closed if the gate cannot complete. This can block recognizable credentials and personal data in recovered plaintext, but it cannot prove that all proprietary reasoning has been removed. This relationship is an evidence route only and supplies no positive score credit without observed control operation.HighCn-3 Ro-3
AID-E-001.001Root & Long-Lived Credential Object EvictionAIDEFEND dataVersion 2026.08.05. When evidence proves that a password, API key, client secret, private key, or other root or long-lived credential appeared in exposed reasoning plaintext, enumerate the exact incident-scoped objects and revoke, disable, reset, or rotate them at every authoritative issuer and verifier. This is post-exposure containment, not prevention of reasoning extraction. This relationship is an evidence route only and supplies no positive score credit without observed control operation.MediumCn-5
AID-E-001.002Issued Token, Authentication Session & Lease RevocationAIDEFEND dataVersion 2026.08.05. If the exposed material includes access tokens, refresh tokens, session cookies, authorization-server sessions, gateway sessions, or agent leases, revoke the complete affected population at every issuer, verifier, cache, and enforcement point. Rotating a long-lived credential does not invalidate these already-issued objects, so they require a separate response action. This relationship is an evidence route only and supplies no positive score credit without observed control operation.MediumCn-5

WHAT THIS CASE TEACHES

Opaque state is not safe state: reasoning artifacts need confidentiality, principal/session binding, bounded retention, and an output gate before they enter logs or client-visible trajectories.

Sources: AIDEFEND in Action — Stolen Thoughts: Opaque Reasoning Blocks Could Be Replayed to Recover Secrets · Primary source — Stolen Thoughts: Stealing Reasoning from Proprietary Language Models

AITBM sub-metrics referenced