Skip to content
Search
paper2026Unreviewed

Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers

Haochuan Wang, Zechen Zhang

Abstract

Multi-agent LLM systems are entering productionprocessing documents, managing workflows, acting on behalf of users-yet their resilience to prompt injection is still evaluated with a single binary: did the attack succeed? This leaves architects without the diagnostic information needed to harden real pipelines. We introduce a kill-chain canary methodology that tracks a cryptographic token through four stages (EXPOSED → PERSISTED → RELAYED → EXECUTED) across 950 runs, five frontier LLMs, six attac

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{wang2026killchain,
  title = {{Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers}},
  author = {Haochuan Wang and Zechen Zhang},
  year = {2026},
  doi = {10.2139/ssrn.6604918},
  url = {https://doi.org/10.2139/ssrn.6604918}
}