Anqi Liu
2 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces the Configurable Safety Reward Model (CSRM), a novel reward model that can be jointly optimized for calibrated safety compliance and reward modeling, significantly improving LLM safety alignment across diverse and unseen safety configurations.
This paper studies the behavior of mixture-of-experts (MoE) models under distribution shift and proposes an adversarial reweighting method to improve their calibration.
Papers
Toward Calibrated Mixture-of-Experts Under Distribution Shift
Gina Wong, Drew Prinster, Suchi Saria, Rama Chellappa +1 more
This paper studies the behavior of mixture-of-experts (MoE) models under distribution shift and proposes an adversarial reweighting method to improve their calibration.