Skip to content
Search
paperApril 2026Unreviewed

Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models

Nay Myat Min, Long H. Pham, Jun Sun

Abstract

Large language models deployed at runtime can misbehave in ways that clean-data validation cannot anticipate: training-time backdoors lie dormant until triggered, jailbreaks subvert safety alignment, and prompt injections override the deployer's instructions. Existing runtime defenses address these threats one at a time and often assume a clean reference model, trigger knowledge, or editable weights, assumptions that rarely hold for opaque third-party artifacts. We introduce Layerwise Convergenc

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{min2026layerwise,
  title = {{Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models}},
  author = {Nay Myat Min and Long H. Pham and Jun Sun},
  year = {2026},
  month = apr,
  eprint = {2604.24542},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.24542}
}