Skip to content
Search
paperSeptember 2026Unreviewed

An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS

Roberto Campbell, Momin Abbass, Muneeza Azmat, Michal Ulewicz, Raya Horesh, Kristjan Greenewald, Rogério Abreu de Paula, Nathalie Baracaldo

Abstract

Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are costly and tightly coupled to the model, limiting flexibility and scalability. We propose a modular correction framework that augments pretrained LLMs with Activated LoRA (aLoRA) adapters and a context-aware routing mechanism to eliminate harms from misaligned model respo

Categories

Cite

@misc{campbell2026efficient,
  title = {{An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS}},
  author = {Roberto Campbell and Momin Abbass and Muneeza Azmat and Michal Ulewicz and Raya Horesh and Kristjan Greenewald and Rogério Abreu de Paula and Nathalie Baracaldo},
  year = {2026},
  month = sep,
  eprint = {2609.13624},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.13624}
}