Amazon ScienceApr 20, 2026arXiv:2604.18789

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

Jiacheng Liang, Yao Ma, Tharindu Kumarage, Tharindu Kumarage, Satyapriya Krishna, Satyapriya Krishna, Rahul Gupta, Kai-Wei Chang, A.G. Galstyan, Aram Galstyan, Charith Peris, Charith Peris

AI Summary

The paper introduces ARES, a novel framework for identifying and mitigating systemic vulnerabilities in RLHF-aligned LLMs, where both the language model and the reward model fail in tandem. ARES uses a "Safety Mentor" to generate adversarial prompts and corresponding malicious/safe responses, exposing weaknesses in both components. The framework then fine-tunes the RM and subsequently optimizes the core model using the improved RM, leading to enhanced safety robustness without sacrificing model capabilities.

Key Contribution

Current red-teaming efforts miss the forest for the trees: ARES reveals that safety failures often stem from a systemic breakdown between the LLM *and* the reward model, not just the LLM itself.

Abstract

Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an imperfect Reward Model (RM) can become a single point of failure when it fails to penalize unsafe behaviors. While existing red-teaming approaches primarily target policy-level weaknesses, they overlook what we term systemic weaknesses cases where both the core LLM and the RM fail in tandem. We present ARES, a framework that systematically discovers and mitigates such dual vulnerabilities. ARES employs a ``Safety Mentor''that dynamically composes semantically coherent adversarial prompts by combining structured component types (topics, personas, tactics, goals) and generates corresponding malicious and safe responses. This dual-targeting approach exposes weaknesses in both the core LLM and the RM simultaneously. Using the vulnerabilities gained, ARES implements a two-stage repair process: first fine-tuning the RM to better detect harmful content, then leveraging the improved RM to optimize the core model. Experiments across multiple adversarial safety benchmarks demonstrate that ARES substantially enhances safety robustness while preserving model capabilities, establishing a new paradigm for comprehensive RLHF safety alignment.

Red-Teaming & Adversarial Robustness RLHF & Preference Learning

Citation Metrics

Citations0

Influential citations0

References58

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

Related Papers