LLM Red-Teaming with Rule-Based Rewards for Safety Evaluation

S. Mansoor
Mobius Logic Inc, Virginia, United States

Keywords: large language model, red-teaming, rule-based rewards, safety evaluation, adversarial attacks

Large language models remain vulnerable to adversarial prompts designed to extract private information, yet building datasets to study and defend against such attacks typically requires costly manual red-teaming. We present an end-to-end, automated red-teaming pipeline that pairs LLM-driven attack generation with a Rule-Based Reward (RBR) signal, enabling scalable evaluation and filtering of attack quality without human annotation. Starting from only 8 seed examples drawn from the Anthropic HH-RLHF dataset, the pipeline expands this minimal corpus through multistep reinforcement learning into 325 goal–criteria pairs, 236 novel attack conversations, and a final set of 101 diversity-filtered attacks. To validate the RBR classifier, we score 280 conversations (140 private-information attacks and 140 non-private examples), observing strong separation between the two classes. These results demonstrate that a small, curated seed set can be systematically grown into a diverse, high-quality corpus of privacy-eliciting attacks, with the RBR signal serving as a reliable automated proxy for human judgment. Our approach offers a practical, scalable alternative to manual red-teaming pipelines, reducing annotation cost while preserving attack diversity and classifier reliability. We discuss implications for red-teaming methodology and for strengthening LLM defenses against privacy-violating adversarial prompts, and outline directions for extending this pipeline to other sensitive attack categories.