Skip to content
Search
paperMarch 2025Unreviewed

Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States

Xin Wei Chia, Jonathan Pan

arXiv.org

Abstract

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to adversarial manipulations such as jailbreaking via prompt injection attacks. These attacks bypass safety mechanisms to generate restricted or harmful content. In this study, we investigated the underlying latent subspaces of safe and jailbroken states by extracting hidden activations from a LLM. Inspired by attractor dynamics in neuroscience, we hypothesized that LLM activat

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{chia2025probing,
  title = {{Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States}},
  author = {Xin Wei Chia and Jonathan Pan},
  year = {2025},
  month = mar,
  eprint = {2503.09066},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2503.09066},
  url = {https://www.semanticscholar.org/paper/03f458d740dea240d07abe93cd7e756cb6a20cb8}
}