Skip to content
Search
paperJune 2026Unreviewed

PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations

Muberra Ozmen, Subhabrata Majumdar

Abstract

Multi-turn jailbreak attacks on large language models (LLMs) reveal a mismatch in current guardrails: they operate on individual turns, while attacks unfold as trajectories across conversations. We propose a shift from content to dynamics, modeling conversations as paths in representation space and asking whether adversarial intent is encoded early in their geometry. We introduce PsychoPass, a framework that extracts geometric features from conversation trajectories in embedding space to predict

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{ozmen2026psychopass,
  title = {{PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations}},
  author = {Muberra Ozmen and Subhabrata Majumdar},
  year = {2026},
  month = jun,
  eprint = {2606.03136},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.03136}
}