Skip to content

Guardrails

Constitutional AI, safety layers, and behavioral constraints

Resources
268
Page
6/6

Newest first · 14 reviewed on this page

Search instead
paper2024North American Chapter of the Association for Computational LinguisticsUnreviewed

WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response

Tianrong Zhang, Bochuan Cao, Yuanpu Cao +3

The recent breakthrough in large language models (LLMs) such as ChatGPT has revolutionized production processes at an unprecedented pace. Alongside this progress also comes mounting concerns about LLMs' susceptibility to jailbreaking attacks, which leads to the generation of…

paper2024International Conference on Modern Problems of Radio Engineering, Telecommunications and Computer ScienceUnreviewed

Enhancing System Security: LLM-Driven Defense Against Prompt Injection Vulnerabilities

Oleksandr Muliarevych

This article examines cybersecurity vulnerabilities in systems utilizing Language Model Interfaces, focusing on the challenges of building secure systems. It provides an overview of current interfaces and their associated risks. A key contribution is the design of a prompt…

paper2023International Conference on Learning RepresentationsUnreviewed

Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models

Erfan Shayegani, Yue Dong, Nael B. Abu-Ghazaleh

We introduce new jailbreak attacks on vision language models (VLMs), which use aligned LLMs and are resilient to text-only jailbreak attacks. Specifically, we develop cross-modality attacks on alignment where we pair adversarial images going through the vision encoder with…

paper2022arXiv preprintReviewed

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai, Saurav Kadavath, Sandipan Kundu +44

Introduces Constitutional AI (CAI), a method for training AI systems to be harmless using a set of principles (a constitution) and AI-generated feedback, reducing reliance on human red teamers.

GuardrailsResponsible AIOpen access1,100 cit.