August 2026Unreviewed
Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings
Istiaque Ahmed, Afia Anjum Borsha, Ranat Das Prangon, Abu-fuad Ahmad, Thi Hong Tran
Abstract
Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. However, they often add a delay of about 250-900 ms to each request. This delay is too high for real-time applications, when the system usually needs to respond in less than 100 ms. Furthermore, routing user prompts through external moderat
Categories
Cite
@misc{ahmed2026reflexguard,
title = {{Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings}},
author = {Istiaque Ahmed and Afia Anjum Borsha and Ranat Das Prangon and Abu-fuad Ahmad and Thi Hong Tran},
year = {2026},
month = aug,
eprint = {2608.17556},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.17556}
}