Skip to content

Guardrails

Constitutional AI, safety layers, and behavioral constraints

Resources
268
Page
4/6

Newest first

Search instead
paper2026Unreviewed

From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling

Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng

Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechanistic…

paper2026Unreviewed

Representational alignment yields generalizable safety in language models

Lingyu Li, Yan Teng, Yingchun Wang +1

Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize.…

paper2026Unreviewed

When Agent Governance Helps

Michael Ray Johnson, Linda Naimi

No specification says how a governed autotelic AI agent organization, where agents pursue self-generated goals inside guardrails, should be designed and evaluated. We answer in two parts. First, we synthesize the Governed Autotelic Multi-Agent Product Organization (GAMPO)…

paper2026Unreviewed

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari +5

As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates…

paper2026Unreviewed

An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS

Roberto Campbell, Momin Abbass, Muneeza Azmat +5

Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are costly and tightly coupled to the model, limiting…

paper2026Unreviewed

PolicyMem: Geometric Policy Memory for LLM Governance

Yuanchen Bei, Zhengzhang Chen, Yanjun Zhao +3

As large language models (LLMs) are increasingly deployed in real-world high-stakes applications, effective governance has become essential. Existing safeguards largely follow two paradigms: learning-based guards provide strong semantic discrimination but couple policy behavior…

paper2026International Journal For Multidisciplinary ResearchUnreviewed

Adversarial Robustness of Foundation Models for Intelligent Mechanical Systems: Threat Models, Benchmarks, and Defense Stacks

Vishwanath

Foundation models increasingly operate across modalities (vision, language, audio, and vision–language) and are deployed in decision-critical pipelines with tool use and retrieval. This expands the adversarial surface: small perturbations to images or audio can flip predictions,…

paper2026Unreviewed

Uncensored Open-weight Models: Redistribution as the Persistence Layer

10a Labs Juliette Garcia, Hailey May, Bobby McKenzie +5

A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We profile this ecosystem by identifying key producers, downstream reproductions, and emerging applications. Between January 2024 and March 2026, we identified 3,471…

paper2026IEEE Transactions on Dependable and Secure ComputingUnreviewed

Beyond Single-Pair Attacks: Disrupting Vision-Language Pre-Training Models With Dual-Semantic Frequency Stealth

Hai-Qi Zhang, Zi-Qiang Li, Hao Tang +1

Vision-Language Pre-training (VLP) models are highly capable in multimodal tasks but are critically vulnerable to adversarial attacks. Existing methods for creating transferable adversarial examples typically operate by modifying semantics within isolated image-text pairs. This…

paper2026Unreviewed

SingProbe Technical Report

Singg Team

Runtime guardrails are essential for reliable large language model (LLM) deployment, yet existing approaches typically rely on independent, external models that introduce additional inference cost, delayed safety signals, and a capacity mismatch with increasingly capable base…

paper2026International Journal for Educational IntegrityUnreviewed

Governing generative AI in higher education: emerging policy approaches and support ecosystems at innovative U.S. Universities

Yu-Feng Qian

This study examines how the 50 U.S. universities ranked as most innovative by U.S. News & World Report articulate policy, guidance, and support for generative artificial intelligence (GenAI) in teaching and learning. Using qualitative document analysis and inductive thematic…

GuardrailsOpen access
paper2026Unreviewed

Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems

I. Shraga, Roei Eshel, Lior Gorelik

A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is correct at check time but stale by actuation. This creates an Execute-stage time-of-check to time-of-use (TOCTOU) hazard. We study verdict freshness: whether a guardrail verdict…