About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
OpenRedRL: a light-weight benchmark for reinforcement learning-based red teaming
Xiang ZHENG , Xingjun MA , Wei-Bin LEE , Cong WANG
Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (8) : 2108813
Red teaming has proven effective for identifying and mitigating vulnerabilities in Large Language Models (LLMs). Reinforcement Learning (RL) has emerged as a promising strategy among existing red teaming techniques. However, a lack of a unified benchmark hinders current RL-based red teaming methods. Implementation details, especially in Proximal Policy Optimization (PPO)-based RL, significantly affect the stability and reproducibility of outcomes. To address this issue, we introduce , a lightweight benchmark that simplifies and standardizes the implementation and evaluation of RL-based red teaming. combines the design strengths of both single-file CleanRL and highly modularized Tianshou, offering high-quality single-file red teaming implementations and modular PPO core components, such as the General Advantage Estimator. It supports a variety of token and sentence diversity metrics, featuring modularized intrinsic reward computation that facilitates plug-and-play experimentation. To clarify their influence on RL performance, we conduct an extensive ablation study of key components, including Low-Rank Adaptation (LoRA), Kullback-Leibler (KL) divergence, and Lagrange Multiplier. We hope this work contributes to 1) gaining a comprehensive understanding of the implementation nuances of RL-based red teaming algorithms, and 2) enabling rapid prototyping of innovative features for RL-based red teaming. Code for the benchmark is publicly available website at github.com/x-zheng16/OpenRedRL.
reinforcement learning / red teaming / benchmark / intrinsic motivation / diversity / large language models
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
Wang X, Chen Y, Li J, Wang Y, Yao Y, Gu T, Li J, Teng Y, Wang Y, Hu X. OpenRT: an open-source red teaming framework for multimodal LLMs. 2026, arXiv preprint arXiv:2601.01592 |
| [31] |
Laskin M, Yarats D, Liu H, Lee K, Zhan A, Lu K, Cang C, Pinto L, Abbeel P. URLB: unsupervised reinforcement learning benchmark. In: Proceedings of the Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. 2021 |
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
| [42] |
|
| [43] |
|
| [44] |
|
| [45] |
|
| [46] |
|
| [47] |
|
| [48] |
Reimers N, Gurevych I. Sentence-BERT: sentence embeddings using Siamese BERT-networks. In: Proceedings of 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing. 2019, 3982−3992 |
| [49] |
|
| [50] |
|
| [51] |
|
| [52] |
|
| [53] |
|
| [54] |
|
| [55] |
|
| [56] |
|
| [57] |
|
| [58] |
|
| [59] |
|
| [60] |
|
| [61] |
|
| [62] |
|
| [63] |
Gallego V. GPT2-alpaca (Revision dd90aae). Hugging Face, See huggingface.co/vicgalle/gpt2-alpaca website, 2023 |
Higher Education Press
/
| 〈 |
|
〉 |