AI SECURITY RESEARCH NOTE · NO ERS

System-Level Agent Defenses: Why Indirect Prompt Injection Needs Plan and Policy Boundaries

This source is retained for threat and defensive-evidence research, but AITBM does not manufacture a deployment score where no assessable system boundary exists.

Why this analysis has no ERS

There is no assessed deployment. The brief summarises an academic position paper (arXiv 2603.30016, 'Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection') that argues for plan, policy, approval, execution, and feedback boundaries in general-purpose agents. It reports no incident, no victim system, no observed configuration, and no measurement of any deployment — every statement is an architectural proposal or a critique of benchmark methodology. Scoring it would require inventing a system's IVP, ORP, and ACI inputs from prescriptive text, which the protocol forbids. The brief is retained for its AIDEFEND mapping and for what its benchmark critique implies about AITBM's own evidence-quality machinery.

Classification: Research · source date 2026-05-18

AIDEFEND evidence routes

AIDEFEND defences → AITBM sub-metrics

Identifiers are quoted as they appear on the AIDEFEND in Action brief (retrieved 2026-08-13); the sub-metric mapping is AITBM's own, from the specification's AIDEFEND tables reconciled at catalogue data version 2026.08.05. AIDEFEND renumbers identifiers between releases, so the data version travels with every mapping and neither side's IDs should be cited without one. A mapping identifies a possible evidence route; a recommendation does not prove that the control was implemented or effective and receives no scoring credit by itself.

TechniqueDefence PriorityEvidences
AID-M-009.002Authority Envelope & Action Risk ClassificationParent AID-M-009 (Agent Autonomy & Authority Governance), catalogue dataVersion 2026.08.05. The brief calls this the conceptual anchor for the paper's plan/policy split; in AITBM terms it is the control whose presence or absence is the primary evidence for Cn-1 and Cn-6 in any agentic assessment.Very HighCn-1 Cn-5 Cn-6 Cn-7
AID-H-018.004Intent-Based Dynamic Capability ScopingParent AID-H-018 (Tool Authorization & Capability Scoping), dataVersion 2026.08.05.Very HighCn-1 Cn-6 Cn-7
AID-H-018.002Policy-Based Access ControlParent AID-H-018, dataVersion 2026.08.05. Externalising the decision into checkable policy logic is what makes Cn-1 measurable by SVSR rather than by assessor opinion.Very HighCn-1 Cn-6 Cn-7
AID-H-018.006Continuous Authorization Verification (Anti-TOCTOU)Parent AID-H-018, dataVersion 2026.08.05. Maps to the paper's requirement to re-check sensitive steps after context, plan, policy, or delegation state changes.HighCn-1 Cn-6 Cn-7
AID-H-017.007Dual-LLM Isolation PatternParent AID-H-017 (Secure Agent Architecture), dataVersion 2026.08.05. The quarantined-reader/privileged-executor split the paper argues for.HighCn-5 Cn-7
AID-H-018.005Value-Level Capability Metadata & Data Flow Sink EnforcementParent AID-H-018, dataVersion 2026.08.05. Operationalises the paper's lattice-style information-flow policies by tagging runtime values with provenance and blocking unsafe movement into external sinks.HighCn-1 Cn-6 Cn-7
AID-H-018.003High-Impact Independent Validation & Approval GateParent AID-H-018, dataVersion 2026.08.05. The paper's plan/policy approver is precisely the Cn-6 pre-execution gate.HighCn-1 Cn-6 Cn-7
AID-M-006.001HITL Checkpoint Design & DocumentationParent AID-M-006 (Human-in-the-Loop Control Design & Readiness), dataVersion 2026.08.05. Triggers, operator roles, default-deny timeouts, and SOPs are what distinguish a Cn-6 score of 0.75 (human authority required for delegated-irreversible actions) from 1.00 (verifiable approval recorded in a tamper-evident trail).MediumCn-2 Cn-6
AID-M-008Automated Agentic Security BenchmarkingdataVersion 2026.08.05. The paper's benchmark critique lands here; in AITBM terms it argues for the Ec fidelity factor and for the Behavioral Attestation Battery's pre-registered canary and adaptive-payload requirements, because a static benchmark run once in a sandbox cannot support a high Evaluation Coverage score.MediumRo-2

Related AITBM rubrics

Sources