Skip to content
Search
paperSeptember 2026Unreviewed

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning

Robin Haselhorst, Lucie Flek, Florian Mai

Abstract

Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect without prior knowledge what to look for. We present activation-matched finetuning, an unsupervised detection method that assumes no knowledge of the trigger or the target behavior. Given a suspect model and a publicly available anchor, we finetune the anchor to re

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{haselhorst2026detecting,
  title = {{Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning}},
  author = {Robin Haselhorst and Lucie Flek and Florian Mai},
  year = {2026},
  month = sep,
  eprint = {2609.00351},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.00351}
}