Skip to content
Search
paperJune 2026Unreviewed

GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

Paulo Ricardo Ferreira Neves, Edson Rodrigues da Cruz Filho, Paulo Henrique Eleuterio Falsetti, João Vitor Pavan, Ian Degaspari, Henrique Vieira Laturrague, Patrick Vieira Laturrague, Guilherme Nielsen Dias, Marccello Wilson Perez Berto, Gustavo Voltani Von Atzingen

Abstract

Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attacks. In addition, benchmark evaluations may be affected by contamination and partial information leakage, compromising performance estimates. This work presents GuardNet, a guardrail system based on an ensemble of shallow neural networks (BiLSTMs) with approximately 47 million parameters. We investigate the hypothesis that robustness in adversarial

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{neves2026guardnet,
  title = {{GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection}},
  author = {Paulo Ricardo Ferreira Neves and Edson Rodrigues da Cruz Filho and Paulo Henrique Eleuterio Falsetti and João Vitor Pavan and Ian Degaspari and Henrique Vieira Laturrague and Patrick Vieira Laturrague and Guilherme Nielsen Dias and Marccello Wilson Perez Berto and Gustavo Voltani Von Atzingen},
  year = {2026},
  month = jun,
  eprint = {2606.05566},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.05566}
}