May 2026Unreviewed
On the Geometric Limits of Transformer Defenses against Obfuscation Attacks: Latent Embedding Collapse & Performance Robustness Gap
Becky Mashaido, Tapadhir Das
Abstract
Prompt injection attacks pose significant risks to language model safety, yet existing defenses are typically evaluated using classification performance. We show that high detection performance does not imply representational robustness. Specifically, multi-operator obfuscated prompts (combining homoglyphs, zero-width characters, and punctuation or emoji noise) can partially collapse onto the embedding manifold of clean prompts, a phenomenon we term latent embedding collapse. Results indicate th
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{mashaido2026geometric,
title = {{On the Geometric Limits of Transformer Defenses against Obfuscation Attacks: Latent Embedding Collapse \& Performance Robustness Gap}},
author = {Becky Mashaido and Tapadhir Das},
year = {2026},
month = may,
eprint = {2605.19159},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.19159}
}