PUBLIC-EVIDENCE AI SECURITY CASE STUDY

Deadbugz: GitHub PRs Delivered a Delayed, Shape-Shifting Malicious MCP Server

Pillar Security documented Deadbugz, an active campaign that opened 23 GitHub pull requests in about 74 minutes to place a malicious MCP server in popular repositories and MCP directories. The server behaved normally at first, then changed later tool descriptions and prompt content into instructions for credential discovery, covert file access, command execution, and exfiltration. None of the reviewed pull requests was merged, and the public evidence does not show victim execution or confirmed data theft.

Agentic / MCP SystemTier 1Indicative ERS 6.4 (3.2–9.6)Evidence source date 2026-08-12

Pillar Security documented Deadbugz, an active campaign that opened 23 GitHub pull requests in about 74 minutes to place a malicious MCP server in popular repositories and MCP directories. The server behaved normally at first, then changed later tool descriptions and prompt content into instructions for credential discovery, covert file access, command execution, and exfiltration. None of the reviewed pull requests was merged, and the public evidence does not show victim execution or confirmed data theft.

ASSESSED SYSTEM

A representative MCP client that has installed the observed Deadbugz server, refreshes tool descriptors or prompts after the server's delayed behavior change, and exposes filesystem, command, repository, or network tools to its agent.

OUT OF SCOPE

The reviewed GitHub pull requests themselves, none of which was merged; any actual victim compromise, execution, credential access, or theft, none of which is confirmed publicly.

Architecture: Agentic / MCP System (decision tree Q4) — The server dynamically supplies tool and prompt context to an agent that can select privileged tools and external effects. Tier 1: Tier 1 because the representative boundary includes credential discovery, covert file access, command execution, and exfiltration authority.

Documented attack or failure path

  1. The campaign used real pull requests for delivery. The actor submitted remote and local MCP configurations, plus directory listings, to established projects. This made repository review the supply-chain boundary the actor was trying to cross.
  2. The server delayed its malicious behavior. It tracked calls by client IP and kept two tools benign. After three tool calls, a later tools/list or prompts/get response returned poisoned content. No change notification was sent, so the client still had to refresh the list or request the prompt.
  3. The payload relied on client authority. It asked the agent to find credential files, conceal access, modify files, run commands, and exfiltrate results. The server could not read victim files directly; success required a client to ingest the changed content and an agent with sufficient tools, permissions, and compliance.
  4. No victim compromise is confirmed. Pillar found no merged pull request in the reviewed set, and the public record does not prove installation, execution, credential access, or exfiltration.

Observed controls and bounded outcomes

Positive credit is given only where the record directly demonstrates a control operating. Recommended or merely presumed controls receive no positive scoring credit.

  • All 23 reviewed pull requests were rejected or remained unmerged, so repository review held for the observed delivery attempts; that outcome is outside the representative post-install score.
  • The server could not read victim files directly and still required a client to ingest changed content plus an agent with sufficient permissions and compliance.

Layer 1 — Intrinsic Vulnerability Profile

Each sub-metric is placed on its five-level rubric by the evidence quoted beside it. Missing applicable evidence remains unknown. The displayed midpoint and interval are scenario values, not inferred control performance.

Sub-metricScoreRubric basisEvidence
Robustness (Ro) — scenario interval 0.07–0.53 (midpoint 0.30), Tier 1 MVT 0.60 confirmed breach
Ro-1Adversarial Input Resistance0.25w 0.30The 0.25 anchor: ordinary protocol content could act as delayed indirect prompt injection and no measured robust instruction/data separation is evidenced.After benign behavior, changed tool descriptions and prompt content delivered instructions for credential discovery, covert access, command execution, and exfiltration to the client model.source: primary/brief
Ro-4Poisoning Attack Resistance0.00w 0.25The 0.00 anchor: the representative installed artifact can silently poison its own tool and prompt metadata after admission.The campaign attempted to place a server into repositories and directories, then changed its effective behavior after installation without a client-notified trusted update.source: primary/brief
Fairness (Fa) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Transparency (Tr) — scenario interval 0.05–0.85 (midpoint 0.45), Tier 1 MVT 0.55 indeterminate
Tr-4Model Lineage Disclosure0.25w 0.20The 0.25 anchor: a named server and initial manifest existed, but digest-bound runtime descriptor and prompt lineage was incomplete.The server could present benign descriptors during onboarding and later return different content without a change notification, leaving the approved and runtime semantics unbound.source: primary/brief
Privacy (Pr) — scenario interval 0.00–1.00 (midpoint 0.50), Tier 1 MVT 0.60 indeterminate
Containment (Cn) — scenario interval 0.14–0.59 (midpoint 0.36), Tier 1 MVT 0.65 confirmed breach
Cn-1Scope Enforcement0.25w 0.12The 0.25 anchor: installed-tool scope provided a coarse boundary, but current intent and exact targets were not deterministically enforced at dispatch.Success required a client agent with sufficient filesystem, shell, messaging, repository, or network authority; poisoned server text could steer that ambient authority beyond the user's task.source: primary/brief
Cn-2Escalation Prevention0.25w 0.16The 0.25 anchor: no evidenced independent policy prevented untrusted protocol metadata from requesting a higher-impact action class.Delayed content attempted to escalate from ordinary MCP use into credential discovery, concealed file access, commands, and exfiltration.source: primary/brief
Cn-5Agent Identity Integrity0.25w 0.14The 0.25 anchor: basic endpoint identity may exist, but continuous attestation of the server's effective tool contract and state is absent.A known server endpoint could change descriptor and prompt semantics after approval; server identity alone did not bind the trusted content or delegation lineage.source: primary/brief
Cn-6Action Reversibility Classification Rate0.25w 0.12The 0.25 anchor: installation approval is not a formal per-action reversibility classification or worst-hop chain gate.The representative client could execute credential reads, commands, or outbound transfers without an independent, single-use approval bound to each exact irreversible or delegated effect.source: primary/brief

Unknown, not N/A: 16 applicable sub-metrics lack admissible public evidence. The lower, midpoint, and upper scenarios evaluate each at 0.00, 0.50, and 1.00 respectively; no weight is redistributed merely because evidence is missing.

Layer 2 — Operational Risk Posture

DimensionScoreJustification
AaAutonomy Amplification1.00w 0.35Placed from the demonstrated decision and action authority of the assessed boundary; public evidence supports this bounded level but not a broader claim about current product defaults. elevated
AsAttack Surface Exposure0.75w 0.25Placed from who can supply the initiating content or protocol message and from the trust status of that source in the documented configuration.
CpCascade Potential1.00w 0.25No verified System Dependency Graph with DGC at least 0.90 is public for this boundary, so the specification's worst-case graph default governs rather than an assessor-estimated blast radius.GDCP: worst-case default — no verified dependency graph published elevated
RfRemediation Feasibility0.50w 0.15Placed from the documented remediation class: deterministic package/configuration change where available, otherwise provider, architecture, or multi-layer changes. It does not assert fleet-wide closure.

Nelevated = 2 (dimensions strictly above 0.75) → CRM = 1.15.

Compound Risk Alert. Two or more dimensions are simultaneously elevated (spec 3.2.2); architectural decomposition is recommended before deployment.

Layer 3 — Public-evidence confidence diagnostic

ComponentScoreBasis
Pc — Public provenance evidence0.35The public record identifies the affected product or representative configuration, attack path, and principal control boundaries, but does not provide a complete asset manifest, verified dependency graph, configuration export, or assessment evidence manifest.
Ec — Public evaluation coverage0.29coverage 0.30 (7 of 23 applicable sub-metrics) × independence 1.00 × fidelity 0.95. No Full, Standard, or Lite pathway is claimed for a retrospective article.
Tf — Public-evidence freshness0.65Evidence dated 2026-08-12; age 1 days on the workpaper reference date. Components: C_monitor 0.65 · C_event 0.65 · C_evidence 0.85 · T_behavior 0.93 · T_containment 0.95 · T_calendar 0.98. Binding term: C_monitor. Evidence age is measured from 2026-08-12 to the 2026-08-13 workpaper reference date. Public sources do not provide a passing containment or behavioral re-attestation receipt; event, monitoring, and unresolved-evidence caps remain diagnostic only.

Public-evidence ACI = (Pc × Ec × Tf)1/3 = 0.40 — diagnostic status: Critical evidence limitation. It describes the evidence available to this case study, not the assurance of the underlying system, and is not inserted into the normalized-assurance scenario ERS.

Indicative ERS — normalized-assurance scenario

Worp · ORP0.35(1.00) + 0.25(0.75) + 0.25(1.00) + 0.15(0.50) = 0.863
CRMNelevated = 2 → 1.15
ORPeffective0.863 × 1.15 = 0.992
Wivp · IVP midpoint0.30(0.30) + 0.25(0.50) + 0.15(0.45) + 0.20(0.50) + 0.10(0.36) = 0.419
IVP mitigation0.15 + 0.85(1 − 0.419) = 0.644
Scenario assuranceACI fixed at 1.000 for cross-case comparison; public-evidence ACI 0.404 is diagnostic only
Indicative ERS midpointmin(10, 0.992 × 0.644 × 1/1.000 × 10) = 6.4
Unknown-input interval3.2–9.6; 16 unknown applicable sub-metrics set to 1.00 / 0.00 at the bounds

AIDEFEND defences → AITBM sub-metrics

Identifiers are quoted as they appear on the AIDEFEND in Action brief (retrieved 2026-08-13); the sub-metric mapping is AITBM's own, from the specification's AIDEFEND tables reconciled at catalogue data version 2026.08.05. AIDEFEND renumbers identifiers between releases, so the data version travels with every mapping and neither side's IDs should be cited without one. A mapping identifies a possible evidence route; a recommendation does not prove that the control was implemented or effective and receives no scoring credit by itself.

TechniqueDefence PriorityEvidences
AID-H-021.001Client-Side Configuration EnforcementAIDEFEND dataVersion 2026.08.05. Validate every proposed MCP configuration against a strict schema and policy before commit or build acceptance. Reject unapproved remote endpoints, hidden local script paths, and configurations that grant broad client capabilities; require protected branches and code-owner review for production configuration. This blocks the delivery path before a Deadbugz pull request can make the server available to users. This relationship is an evidence route only and supplies no positive score credit without observed control operation.Very HighCn-2 Cn-5
AID-H-024.002MCP Tool Descriptor Hash Binding & Drift DetectionAIDEFEND dataVersion 2026.08.05. Canonicalize every approved MCP tool descriptor, pin its digest during onboarding, and compare each later tools/list response with that manifest before exposing it to the agent. Deadbugz changed descriptor text after benign calls, so a mismatch can fail closed before the injected instructions reach the privileged model. This control applies to tool descriptors, not to prompts/get content. This relationship is an evidence route only and supplies no positive score credit without observed control operation.Very HighCn-5
AID-H-024.004Approved Tool Contract Semantics & Change AdmissionAIDEFEND dataVersion 2026.08.05. Encode policy checks for each MCP tool's approved operation, side effects, privilege scope, and destructive behavior. Bind the exact descriptor and semantic profile to a signed approval, and require a new security review whenever either changes. This catches malicious initial contracts as well as Deadbugz's delayed change in tool meaning. This relationship is an evidence route only and supplies no positive score credit without observed control operation.Very HighCn-5
AID-H-018.004Intent-Based Dynamic Capability ScopingAIDEFEND dataVersion 2026.08.05. At the start of each request or session, intersect the authenticated user's stated task with the exact versioned tool registry, action budget, and grant lifetime, then enforce that signed scope at the dispatcher. A formatting or summarization task should never receive credential-reading, shell, email, or repository-write capabilities, so the poisoned instructions have no matching authority to invoke. This relationship is an evidence route only and supplies no positive score credit without observed control operation.Very HighCn-1 Cn-6 Cn-7
AID-H-017.002Least-Privilege Tool ArchitectureAIDEFEND dataVersion 2026.08.05. Expose only allowlisted, strongly typed, single-purpose tools with bounded paths, destinations, credentials, and side effects. Remove generic shell, broad filesystem, arbitrary HTTP, and equivalent ambient-power tools from the agent. Even if poisoned MCP text reaches the model, it no longer has a general mechanism for reading credential files, running arbitrary commands, or sending data anywhere. This relationship is an evidence route only and supplies no positive score credit without observed control operation.HighCn-5 Cn-7
AID-H-018.003High-Impact Independent Validation & Approval GateAIDEFEND dataVersion 2026.08.05. Before executing code, sending email, changing a repository, or reading a protected credential path, independently validate the exact immutable action, target, parameters, current identity, and signed policy. When policy requires approval, bind it to that same action and consume it once at the executor. Attacker-authored MCP text cannot serve as authorization for the requested effect. This relationship is an evidence route only and supplies no positive score credit without observed control operation.HighCn-1 Cn-6 Cn-7
AID-H-018.005Value-Level Capability Metadata & Data Flow Sink EnforcementAIDEFEND dataVersion 2026.08.05. Attach provenance and sensitivity labels to values read from credential paths, shell history, or Kubernetes configuration, then enforce sink policy at command, email, network, and repository dispatch. This blocks a sensitive value from leaving through a tool even if the model follows the injected instructions. This relationship is an evidence route only and supplies no positive score credit without observed control operation.HighCn-1 Cn-6 Cn-7
AID-E-003.003Confirmed Malicious Code & Persistence EvictionAIDEFEND dataVersion 2026.08.05. If investigation confirms that a Deadbugz configuration, script, package, or startup artifact was installed, remove the exact object from every authoritative configuration source and loaded runtime, preserve evidence, and verify that no runtime can load it again. This contains an installed artifact but does not replace review of actions already performed. This relationship is an evidence route only and supplies no positive score credit without observed control operation.MediumRo-4
AID-E-001.001Root & Long-Lived Credential Object EvictionAIDEFEND dataVersion 2026.08.05. When evidence shows that a password, API key, SSH private key, client secret, or other long-lived credential was read or exposed, enumerate the exact affected objects and revoke, disable, or rotate them at every authoritative issuer and verifier. This is incident-scoped containment after exposure, not a reason to rotate all credentials based only on the presence of a malicious pull request. This relationship is an evidence route only and supplies no positive score credit without observed control operation.MediumCn-5

WHAT THIS CASE TEACHES

MCP onboarding cannot be a one-time trust decision: clients must bind descriptor and prompt semantics, detect drift, and re-authorize every consequential action against current user intent.

Sources: AIDEFEND in Action — Deadbugz: GitHub PRs Delivered a Delayed, Shape-Shifting Malicious MCP Server · Primary source — Deadbugz: Currently Active MCP Supply-Chain Campaign

AITBM sub-metrics referenced