April 2026Unreviewed
Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
Nay Myat Min, Long H. Pham, Jun Sun
Abstract
Large language models deployed at runtime can misbehave in ways that clean-data validation cannot anticipate: training-time backdoors lie dormant until triggered, jailbreaks subvert safety alignment, and prompt injections override the deployer's instructions. Existing runtime defenses address these threats one at a time and often assume a clean reference model, trigger knowledge, or editable weights, assumptions that rarely hold for opaque third-party artifacts. We introduce Layerwise Convergenc
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{min2026layerwise,
title = {{Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models}},
author = {Nay Myat Min and Long H. Pham and Jun Sun},
year = {2026},
month = apr,
eprint = {2604.24542},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.24542}
}