March 2025Unreviewed
Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
Xin Wei Chia, Jonathan Pan
arXiv.org
Abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to adversarial manipulations such as jailbreaking via prompt injection attacks. These attacks bypass safety mechanisms to generate restricted or harmful content. In this study, we investigated the underlying latent subspaces of safe and jailbroken states by extracting hidden activations from a LLM. Inspired by attractor dynamics in neuroscience, we hypothesized that LLM activat
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{chia2025probing,
title = {{Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States}},
author = {Xin Wei Chia and Jonathan Pan},
year = {2025},
month = mar,
eprint = {2503.09066},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2503.09066},
url = {https://www.semanticscholar.org/paper/03f458d740dea240d07abe93cd7e756cb6a20cb8}
}