Skip to content
Search
paperMay 2026Unreviewed

On the Geometric Limits of Transformer Defenses against Obfuscation Attacks: Latent Embedding Collapse & Performance Robustness Gap

Becky Mashaido, Tapadhir Das

Abstract

Prompt injection attacks pose significant risks to language model safety, yet existing defenses are typically evaluated using classification performance. We show that high detection performance does not imply representational robustness. Specifically, multi-operator obfuscated prompts (combining homoglyphs, zero-width characters, and punctuation or emoji noise) can partially collapse onto the embedding manifold of clean prompts, a phenomenon we term latent embedding collapse. Results indicate th

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{mashaido2026geometric,
  title = {{On the Geometric Limits of Transformer Defenses against Obfuscation Attacks: Latent Embedding Collapse \& Performance Robustness Gap}},
  author = {Becky Mashaido and Tapadhir Das},
  year = {2026},
  month = may,
  eprint = {2605.19159},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.19159}
}