July 2026Unreviewed
Geometric Configurations of Perturbed Jailbreak Prompts
Lynn Delcon, Andres Algaba, Vincent Ginis
Abstract
Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probabilit
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{delcon2026geometric,
title = {{Geometric Configurations of Perturbed Jailbreak Prompts}},
author = {Lynn Delcon and Andres Algaba and Vincent Ginis},
year = {2026},
month = jul,
eprint = {2607.20581},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.20581}
}