GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection
Paulo Ricardo Ferreira Neves, Edson Rodrigues da Cruz Filho, Paulo Henrique Eleuterio Falsetti, João Vitor Pavan, Ian Degaspari, Henrique Vieira Laturrague, Patrick Vieira Laturrague, Guilherme Nielsen Dias, Marccello Wilson Perez Berto, Gustavo Voltani Von Atzingen
Abstract
Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attacks. In addition, benchmark evaluations may be affected by contamination and partial information leakage, compromising performance estimates. This work presents GuardNet, a guardrail system based on an ensemble of shallow neural networks (BiLSTMs) with approximately 47 million parameters. We investigate the hypothesis that robustness in adversarial
Categories
Framework mappings
- LLM01Prompt Injection
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{neves2026guardnet,
title = {{GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection}},
author = {Paulo Ricardo Ferreira Neves and Edson Rodrigues da Cruz Filho and Paulo Henrique Eleuterio Falsetti and João Vitor Pavan and Ian Degaspari and Henrique Vieira Laturrague and Patrick Vieira Laturrague and Guilherme Nielsen Dias and Marccello Wilson Perez Berto and Gustavo Voltani Von Atzingen},
year = {2026},
month = jun,
eprint = {2606.05566},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.05566}
}