OpenRedRL: a light-weight benchmark for reinforcement learning-based red teaming

Xiang ZHENG , Xingjun MA , Wei-Bin LEE , Cong WANG

Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (8) : 2108813

PDF (2858KB)
Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (8) :2108813 DOI: 10.1007/s11704-026-51865-8
Information Security
RESEARCH ARTICLE
OpenRedRL: a light-weight benchmark for reinforcement learning-based red teaming
Author information +
History +
PDF (2858KB)

Abstract

Red teaming has proven effective for identifying and mitigating vulnerabilities in Large Language Models (LLMs). Reinforcement Learning (RL) has emerged as a promising strategy among existing red teaming techniques. However, a lack of a unified benchmark hinders current RL-based red teaming methods. Implementation details, especially in Proximal Policy Optimization (PPO)-based RL, significantly affect the stability and reproducibility of outcomes. To address this issue, we introduce OpenRedRL, a lightweight benchmark that simplifies and standardizes the implementation and evaluation of RL-based red teaming. OpenRedRL combines the design strengths of both single-file CleanRL and highly modularized Tianshou, offering high-quality single-file red teaming implementations and modular PPO core components, such as the General Advantage Estimator. It supports a variety of token and sentence diversity metrics, featuring modularized intrinsic reward computation that facilitates plug-and-play experimentation. To clarify their influence on RL performance, we conduct an extensive ablation study of key components, including Low-Rank Adaptation (LoRA), Kullback-Leibler (KL) divergence, and Lagrange Multiplier. We hope this work contributes to 1) gaining a comprehensive understanding of the implementation nuances of RL-based red teaming algorithms, and 2) enabling rapid prototyping of innovative features for RL-based red teaming. Code for the benchmark is publicly available website at github.com/x-zheng16/OpenRedRL.

Graphical abstract

Keywords

reinforcement learning / red teaming / benchmark / intrinsic motivation / diversity / large language models

Cite this article

Download citation ▾
Xiang ZHENG, Xingjun MA, Wei-Bin LEE, Cong WANG. OpenRedRL: a light-weight benchmark for reinforcement learning-based red teaming. Front. Comput. Sci., 2027, 21 (8) : 2108813 DOI:10.1007/s11704-026-51865-8

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C L, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R. Training language models to follow instructions with human feedback. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2011

[2]

Hendrycks D, Carlini N, Schulman J, Steinhardt J. Unsolved problems in ML safety. 2022, arXiv preprint arXiv: 2109.13916

[3]

Ma X, Gao Y, Wang Y, Wang R, Wang X, , et al. Safety at scale: a comprehensive survey of large model and agent safety. Foundations and Trends in Privacy and Security, 2025, 8(3−4): 1−240

[4]

Ma X, Wang Y, Xu H, Wu Y, Ding Y, , et al. A safety report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 fast, nano banana pro, and seedream 4.5. 2026, arXiv preprint arXiv: 2601.10527

[5]

Ganguli D, Lovitt L, Kernion J, Askell A, Bai Y, , et al. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. 2022, arXiv preprint arXiv: 2209.07858

[6]

Andriushchenko M, Croce F, Flammarion N. Jailbreaking leading safety-aligned LLMs with simple adaptive attacks. In: Proceedings of the 13th International Conference on Learning Representations. 2025

[7]

Liu X, Xu N, Chen M, Xiao C. AutoDAN: generating stealthy jailbreak prompts on aligned large language models. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[8]

Lee D, Lee J, Ha J W, Kim J H, Lee S W, Lee H, Song H O. Query-efficient black-box red teaming via Bayesian optimization. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. 2023, 11551−11574

[9]

Yang C, Wang X, Lu Y, Liu H, Le Q V, Zhou D, Chen X. Large language models as optimizers. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[10]

Perez E, Huang S, Song F, Cai T, Ring R, Aslanides J, Glaese A, McAleese N, Irving G. Red teaming language models with language models. In: Proceedings of 2022 Conference on Empirical Methods in Natural Language Processing. 2022, 3419−3448

[11]

Perez E, Ringer S, Lukosiute K, Nguyen K, Chen E, , et al. Discovering language model behaviors with model-written evaluations. In: Proceedings of Association for Computational Linguistics. 2023, 13387−13434

[12]

Deng B, Wang W, Feng F, Deng Y, Wang Q, He X. Attack prompt generation for red teaming and defending large language models. In: Proceedings of Association for Computational Linguistics. 2023, 2176−2189

[13]

Casper S, Lin J, Kwon J, Culp G, Hadfield-Menell D. Explore, establish, exploit: red teaming language models from scratch. 2023, arXiv preprint arXiv: 2306.09442

[14]

Hong Z W, Shenfeld I, Wang T H, Chuang Y S, Pareja A, Glass J R, Srivastava A, Agrawal P. Curiosity-driven red-teaming for large language models. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[15]

Zhao A, Xu Q, Lin M, Wang S, Liu Y J, Zheng Z, Huang G. DiveR-CT: Diversity-enhanced red teaming large language model assistants with relaxing constraints. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. 2025, 26021−26030

[16]

Zheng X, Wang L, Liu Y, Ma X, Shen C, Wang C. CALM: curiosity-driven auditing for large language models. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. 2025, 27757−27764

[17]

Zheng X, Ma X, Shen C, Wang C. Constrained intrinsic motivation for reinforcement learning. In: Proceedings of the 33rd International Joint Conference on Artificial Intelligence. 2024, 5608−5616

[18]

Zheng X, Ma X, Wang S, Wang X, Shen C, Wang C. Toward evaluating robustness of reinforcement learning with adversarial policy. In: Proceedings of the 54th Annual IEEE/IFIP International Conference on Dependable Systems and Networks. 2024, 288−301

[19]

Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. 2017, arXiv preprint arXiv: 1707.06347

[20]

Chen Y, Wang X, Li J, Wang Y, Li J, Teng Y, Wang Y, Ma X. Evolve the method, not the prompts: Evolutionary synthesis of jailbreak attacks on LLMs. 2025, arXiv preprint arXiv: 2511.12710

[21]

Chao P, Robey A, Dobriban E, Hassani H, Pappas G J, Wong E. Jailbreaking black box large language models in twenty queries. In: Proceedings of 2025 IEEE Conference on Secure and Trustworthy Machine Learning. 2025, 23−42

[22]

Mehrotra A, Zampetakis M, Kassianik P, Nelson B, Anderson H, Singer Y, Karbasi A. Tree of attacks: jailbreaking black-box LLMs automatically. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 1952

[23]

Zeng Y, Lin H, Zhang J, Yang D, Jia R, Shi W. How Johnny can persuade LLMs to jailbreak them: rethinking persuasion to challenge AI safety by humanizing LLMs. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024, 14322−14350

[24]

Yu J, Lin X, Yu Z, Xing X. GPTFUZZER: red teaming large language models with auto-generated jailbreak prompts. 2024, arXiv preprint arXiv: 2309.10253

[25]

von Werra L, Belkada Y, Tunstall L, Beeching E, Thrush T, Lambert N, Huang S, Rasul K, Gallouédec Q. TRL: transformers reinforcement learning. See Github.com/huggingface/trl, 2020

[26]

Rafailov R, Sharma A, Mitchell E, Ermon S, Manning C D, Finn C. Direct preference optimization: your language model is secretly a reward model. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 2338

[27]

Hu J, Wu X, Shen W, Liu J K, Zhu Z, Wang W, Jiang S, Wang H, Chen H, Chen B, Fang W, Xianyu , Cao Y, Xu H, Liu Y. OpenRLHF: an easy-to-use, scalable and high-performance RLHF framework. 2025, arXiv preprint arXiv: 2405.11143

[28]

Shao Z, Wang P, Zhu Q, Xu R, Song J, Bi X, Zhang H, Zhang M, Li Y K, Wu Y, Guo D. DeepSeekMath: pushing the limits of mathematical reasoning in open language models. 2024, arXiv preprint arXiv: 2402.03300

[29]

Ramamurthy R, Ammanabrolu P, Brantley K, Hessel J, Sifa R, Bauckhage C, Hajishirzi H, Choi Y. Is reinforcement learning (not) for natural language processing: benchmarks, baselines, and building blocks for natural language policy optimization. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[30]

Wang X, Chen Y, Li J, Wang Y, Yao Y, Gu T, Li J, Teng Y, Wang Y, Hu X. OpenRT: an open-source red teaming framework for multimodal LLMs. 2026, arXiv preprint arXiv:2601.01592

[31]

Laskin M, Yarats D, Liu H, Lee K, Zhan A, Lu K, Cang C, Pinto L, Abbeel P. URLB: unsupervised reinforcement learning benchmark. In: Proceedings of the Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. 2021

[32]

Burda Y, Edwards H, Storkey A, Klimov O. Exploration by random network distillation. In: Proceedings of the 7th International Conference on Learning Representations. 2019

[33]

Burda Y, Edwards H, Pathak D, Storkey A J, Darrell T, Efros A A. Large-scale study of curiosity-driven learning. In: Proceedings of the 7th International Conference on Learning Representations. 2019

[34]

Zhang T, Rashidinejad P, Jiao J, Tian Y, Gonzalez J E, Russell S. MADE: exploration via maximizing deviation from explored regions. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021, 740

[35]

Liu H, Abbeel P. APS: active pretraining with successor features. In: Proceedings of the 38th International Conference on Machine Learning. 2021, 6736−6747

[36]

Hazan E, Kakade S M, Singh K, Van Soest A. Provably efficient maximum entropy exploration. In: Proceedings of the 36th International Conference on Machine Learning. 2019, 2681−2691

[37]

Mutti M, Pratissoli L, Restelli M. Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate. In: Proceedings of the 35th AAAI Conference on Artificial Intelligence. 2021, 9028−9036

[38]

Liu H, Abbeel P. Behavior from the void: unsupervised active pre-training. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021, 1411

[39]

Park S, Rybkin O, Levine S. METRA: scalable unsupervised RL with metric-aware abstraction. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[40]

Gregor K, Rezende D J, Wierstra D. Variational intrinsic control. In: Proceedings of the 5th International Conference on Learning Representations. 2017

[41]

Sharma A, Gu S, Levine S, Kumar V, Hausman K. Dynamics-aware unsupervised discovery of skills. In: Proceedings of the 8th International Conference on Learning Representations. 2020

[42]

Laskin M, Liu H, Peng X B, Yarats D, Rajeswaran A, Abbeel P. CIC: contrastive intrinsic control for unsupervised skill discovery. 2022, arXiv preprint arXiv: 2202.00161

[43]

Park S, Choi J, Kim J, Lee H, Kim G. Lipschitz-constrained unsupervised skill discovery. In: Proceedings of the 10th International Conference on Learning Representations. 2022

[44]

Eysenbach B, Zhang T, Levine S, Salakhutdinov R. Contrastive learning as goal-conditioned reinforcement learning. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2580

[45]

Yang R, Bai C, Guo H, Li S, Zhao B, Wang Z, Liu P, Li X. Behavior contrastive learning for unsupervised skill discovery. In: Proceedings of the 40th International Conference on Machine Learning. 2023, 1634

[46]

OpenAI . GPT-4o system card. 2024, arXiv preprint arXiv: 2410.21276

[47]

DeepSeek-AI . DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. 2026, arXiv preprint arXiv: 2501.12948

[48]

Reimers N, Gurevych I. Sentence-BERT: sentence embeddings using Siamese BERT-networks. In: Proceedings of 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing. 2019, 3982−3992

[49]

Vidgen B, Thrush T, Waseem Z, Kiela D. Learning from the worst: dynamically generated datasets to improve online hate detection. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. 2021, 1667−1682

[50]

Hartvigsen T, Gabriel S, Palangi H, Sap M, Ray D, Kamar E. ToxiGen: a large-scale machine-generated dataset for adversarial and implicit hate speech detection. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. 2022, 3309−3326

[51]

Jindal M. Gibberish detector. See Huggingface.co/madhurjindal/autonlp-Gibberish-Detector-492513457, 2021

[52]

Andrychowicz M, Raichuk A, Stańczyk P, Orsini M, Girgin S, Marinier R, Hussenot L, Geist M, Pietquin O, Michalski M, Gelly S, Bachem O. What matters in on-policy reinforcement learning? A Large-scale empirical study. 2020, arXiv preprint arXiv: 2006.05990

[53]

Lee S, Kim M, Cherif L, Dobre D, Lee J, Hwang S J, Kawaguchi K, Gidel G, Bengio Y, Malkin N, Jain M. Learning diverse attacks on large language models for robust red-teaming and safety tuning. In: Proceedings of the13th International Conference on Learning Representations. 2025

[54]

Zou A, Wang Z, Carlini N, Nasr M, Kolter J Z, Fredrikson M. Universal and transferable adversarial attacks on aligned language models. 2023, arXiv preprint arXiv: 2307.15043

[55]

Bianchi F, Suzgun M, Attanasio G, Röttger P, Jurafsky D, Hashimoto T, Zou J. Safety-tuned LLaMAs: lessons from improving the safety of large language models that follow instructions. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[56]

Wang Z, Zheng X, Wang X, Wang B, Ma X, Jiang Y G. GenBreak: red teaming text-to-image generators using large language models. 2025, arXiv preprint arXiv: 2506.10047

[57]

Zhao Y, Zheng X, Luo L, Li Y, Ma X, Jiang Y G. BlueSuffix: reinforced blue teaming for vision-language models against jailbreak attacks. In: Proceedings of the 13th International Conference on Learning Representations. 2025

[58]

Li J, Li Y, Huang H, Chen Y, Wang X, Wang Y, Ma X, Jiang Y G. BackdoorVLM: a benchmark for backdoor attacks on vision-language models. 2025, arXiv preprint arXiv: 2511.18921

[59]

Zheng X, Wu Y, Huang H, Li Y, Ma X, Li B, Jiang Y G, Wang C. Just ask: curious code agents reveal system prompts in frontier LLMs. 2026, arXiv preprint arXiv: 2601.21233

[60]

Wu Y, Liu X, Li Y, Gao Y, Ding Y, Ding J, Zheng X, Ma X. ADMIT: few-shot knowledge poisoning attacks on RAG-based fact checking. 2025, arXiv preprint arXiv: 2510.13842

[61]

Feng Y, Li Y, Wu Y, Tan Y, Guo Y, Ding Y, Zhai K, Ma X, Jiang Y G. BackdoorAgent: a unified framework for backdoor attacks on LLM-based agents. 2026, arXiv preprint arXiv: 2601.04566

[62]

Li Y, Li Z, Zhao W, Min N M, Huang H, Ma X, Sun J. AutoBackdoor: automating backdoor attacks via LLM agents. 2025, arXiv preprint arXiv: 2511.16709

[63]

Gallego V. GPT2-alpaca (Revision dd90aae). Hugging Face, See huggingface.co/vicgalle/gpt2-alpaca website, 2023

RIGHTS & PERMISSIONS

Higher Education Press

PDF (2858KB)

726

Accesses

0

Citation

Detail

Sections
Recommended

/