September 2026Unreviewed
An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS
Roberto Campbell, Momin Abbass, Muneeza Azmat, Michal Ulewicz, Raya Horesh, Kristjan Greenewald, Rogério Abreu de Paula, Nathalie Baracaldo
Abstract
Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are costly and tightly coupled to the model, limiting flexibility and scalability. We propose a modular correction framework that augments pretrained LLMs with Activated LoRA (aLoRA) adapters and a context-aware routing mechanism to eliminate harms from misaligned model respo
Categories
Cite
@misc{campbell2026efficient,
title = {{An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS}},
author = {Roberto Campbell and Momin Abbass and Muneeza Azmat and Michal Ulewicz and Raya Horesh and Kristjan Greenewald and Rogério Abreu de Paula and Nathalie Baracaldo},
year = {2026},
month = sep,
eprint = {2609.13624},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.13624}
}