• Home
  • Search
  • Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
  • Cite Icon1
  • https://doi.org/10.1609/aaai.v40i44.41104Copy DOI Icon

Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment

  • Abstract
  • Literature Map
  • Citations
  • Similar Papers
Abstract

Reward-model-based fine-tuning is a central paradigm in aligning Large Language Models with human preferences. However, such approaches critically rely on the assumption that proxy reward models accurately reflect intended supervision, a condition often violated due to annotation noise, bias, or limited coverage. This misalignment can lead to undesirable behaviors, where models optimize for flawed signals rather than true human values. In this paper, we investigate a novel framework to identify and mitigate such misalignment by treating the fine-tuning process as a form of knowledge integration. We focus on detecting instances of proxy-policy conflicts, cases where the base model strongly disagrees with the proxy. We argue that such conflicts often signify areas of shared ignorance, where neither the policy nor the reward model possesses sufficient knowledge, making them especially susceptible to misalignment. To this end, we propose two complementary metrics for identifying these conflicts: a localized Proxy-Policy Alignment Conflict Score (PACS) and a global Kendall-Tau Distance measure. Building on this insight, we design an algorithm named Selective Human-in-the-loop Feedback via Conflict-Aware Sampling (SHF-CAS) that targets high-conflict QA pairs for additional feedback, refining both the reward model and policy efficiently. Experiments on two alignment tasks demonstrate that our approach enhances general alignment performance, even when trained with a biased proxy reward. Our work provides a new lens for interpreting alignment failures and offers a principled pathway for targeted refinement in LLM training.

Similar Papers
  • Conference Article

APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport

  • Jan 01, 2025
  • Zhuo Li +5
  • Conference Article
  • Citations1

Exploring Domain Robust Lightweight Reward Models based on Router Mechanism

  • Jan 01, 2024
  • Hyuk Namgoong +3
  • Conference Article

DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging

  • Jan 01, 2024
  • Tzu-Han Lin +3
  • Research Article
  • Citations1

Leveraging Large Language Models for Review Classification and Rating Estimation of Mental Health Applications

  • Jun 07, 2025
  • Proceedings of the International AAAI Conference on Web and Social Media
  • Qile Wang +9
  • Conference Article

Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Search

  • Jan 01, 2025
  • Shuo Huang +3
  • Research Article
  • Citations3

Decoding Multilingual Moral Preferences: Unveiling LLM's Biases through the Moral Machine Experiment

  • Oct 16, 2024
  • Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society
  • Karina Vida +2
  • Research Article

Model-Agnostic Sentiment Distribution Stability Analysis for Robust LLM-Generated Texts Detection

  • Mar 14, 2026
  • Siyuan Li +7
  • Conference Article
  • Citations8

RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models

  • Jan 01, 2024
  • Jiongxiao Wang +4
  • Research Article
  • Citations24

Evaluation metrics in medical imaging AI: fundamentals, pitfalls, misapplications, and recommendations

  • Sep 01, 2025
  • European Journal of Radiology Artificial Intelligence
  • Burak Kocak +10
  • Conference Article

Token-Level Accept or Reject: A Micro Alignment Approach for Large Language Models

  • Sep 01, 2025
  • Yang Zhang +10
  • Research Article
  • Citations1

PRIORITY2REWARD: Incorporating Healthworker Preferences for Resource Allocation Planning

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Shresth Verma +4
  • Research Article

Enhancing IELTS writing automated scoring with M-LoRA fine-tuned LLAMA-3 and human feedback-driven PPO reinforcement learning.

  • Mar 27, 2026
  • Scientific reports
  • Wenbo Xu +2
  • Research Article

Generating evasive payloads for assessing Web Application Firewalls with Reinforcement Learning and Pre-trained Language Models

  • Oct 02, 2025
  • Journal of Science and Technology on Information security
  • Tran Gia Bao +2
  • Conference Article

Language Models as Continuous Self-Evolving Data Engineers

  • Jan 01, 2025
  • Peidong Wang +7
  • Research Article

From Principle to Practice: Value Alignment in AI Ethics and Governance

  • Oct 01, 2025
  • German Law Journal
  • Jianfeng Cao
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.