Skip to content

Guardrails

Constitutional AI, safety layers, and behavioral constraints

Resources
268
Page
2/6

Newest first

Search instead
paper2026Unreviewed

Parallax: Why AI Agents That Think Must Never Act

Joel Fokou

Autonomous AI agents are rapidly transitioning from experimental tools to operational infrastructure, with projections that 80% of enterprise applications will embed AI copilots by the end of 2026. As agents gain the ability to execute real-world actions (reading files, running…

paper2026Unreviewed

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

Tongxi Wu, Jian Zhang, Yang Gao

Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshold mechanism. We challenge this assumption by revealing that safety behavior is governed by an instability region where small…

paper2026Unreviewed

Exploring and Developing a Pre-Model Safeguard with Draft Models

Hongyu Cai, Arjun Arunasalam, Yiming Liang +2

Large Language Model (LLM) alignment remains vulnerable to jailbreak attacks that elicit unsafe responses, motivating pre-model and post-model guards. Pre-model guards audit the safety of prompts before invoking target models. However, relying solely on the prompt often leads to…

paper2026Unreviewed

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs

Wenqi Chen, Ziyan Zhang, Bing Wang +3

While Large Language Models (LLMs) excel in code generation, they remain prone to replicating subtle yet critical vulnerabilities endemic to their training data. Current alignment techniques, such as Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), typically apply…

paper2026Unreviewed

GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

Paulo Ricardo Ferreira Neves, Edson Rodrigues da Cruz Filho, Paulo Henrique Eleuterio Falsetti +7

Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attacks. In addition, benchmark evaluations may be affected by contamination and partial information leakage, compromising…

paper2026Unreviewed

Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification

Thanh Luong Tuan, Abhijit Sanyal

Pre-deployment verification of enterprise artificial intelligence (AI) agents remains a critical gap between large language model (LLM) capability benchmarking and production deployment. Post-deployment monitoring, human-in-the-loop controls, and prompt-level guardrails offer…

paper2026Unreviewed

Alignment Defends LLMs from Property Inference Attacks

Pengrun Huang, Chhavi Yadav, Ruihan Wu +1

Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets that may contain sensitive, dataset-level properties. Recent work has shown that such dataset-level information can be effectively extracted through property inference attacks, posing a…

paper2026Unreviewed

A Virtuous AI is an Existential Risk

Guillermo Del Pinal, Youngchan Lee, Min Ohn

This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the…

paper2026IEEE International Conference on Circuits and Systems for CommunicationsUnreviewed

PromptShield: LoRA-Based Parameter-Efficient Refusal Alignment for Election-Targeted Adversarial LLM Attacks

Nishmitha M R, Tejakshi N S, Anshu Sharma +3

Large language models (LLMs) are increasingly deployed in public-facing information systems, where their misuse poses serious risks in high-stakes domains such as democratic elections. Despite extensive safety alignment, contemporary LLMs remain vulnerable to adversarial…

paper20262026 8th International Conference on Software Engineering and Computer Science (CSECS)Unreviewed

Double-Gaming: Jailbreak Attacks against LLM based on Red-Blue Team Game Theory

Chenlu Ma, Huairui Zhao, G. Nie +3

Despite existing security alignment mechanisms, Large Language Models (LLMs) remain vulnerable to jailbreak attacks under static defenses. To address this, we propose a novel jailbreak attack and defense optimization framework based on Red-Blue Team dynamic game theory. This…

paper2026Unreviewed

RAS: Measuring LLM Safety Through Refusal Alignment

Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu +1

Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy. Although useful, output-level evaluation is expensive, sensitive to judge choice, and easily…

paper2026Unreviewed

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs

Shivam Ratnakar, Kartikeya Vats

Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque. In this work, we investigate whether safety compliance is a deep semantic decision or a manipulable linear feature. We introduce Contrastive Logit…

paper2026Unreviewed

Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models

Hugo García Cuesta, Pablo Mateo Torrejón, Alfonso Sánchez-Macián

While Large Language Models (LLMs) have become essential productivity tools, their integration into workflows without adequate safeguards creates significant risks. This paper proposes an open-source, privacy-focused, user-facing firewall designed to secure both web-based and…