Sentence-level reward model can generalize better for aligning LLM from human preference

Wenjie QIU , Yichen LI , Xuqin ZHANG , Tianyi ZHANG , Yihang ZHANG , Zongzhang ZHANG , Yang YU

Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) : 2010382

PDF (330KB)
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) :2010382 DOI: 10.1007/s11704-026-51483-4
Artificial Intelligence
LETTER
Sentence-level reward model can generalize better for aligning LLM from human preference
Author information +
History +
PDF (330KB)

Graphical abstract

Cite this article

Download citation ▾
Wenjie QIU, Yichen LI, Xuqin ZHANG, Tianyi ZHANG, Yihang ZHANG, Zongzhang ZHANG, Yang YU. Sentence-level reward model can generalize better for aligning LLM from human preference. Front. Comput. Sci., 2026, 20 (10) : 2010382 DOI:10.1007/s11704-026-51483-4

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Llama Team. The Llama 3 herd of models. 2024, arXiv preprint arXiv: 2407. 21783

[2]

Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C L, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R. Training language models to follow instructions with human feedback. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2011

[3]

Chan A J, Sun H, Holt S, van der Schaar M. Dense reward for free in reinforcement learning from human feedback. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 6136−6154

[4]

Yang S, Zhang S, Xia C, Feng Y, Xiong C, Zhou M. Preference-grounded token-level guidance for language model fine-tuning. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 24466−24496

[5]

Guo G, Zhao R, Tang T, Zhao X, Wen J R. Beyond imitation: leveraging fine-grained quality signals for alignment. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[6]

Bradley R A, Terry M E . Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 1952, 39( 3−4): 324–345

[7]

Frohmann M, Sterner I, Vulić I, Minixhofer B, Schedl M. Segment any text: a universal approach for robust, efficient and adaptable sentence segmentation. In: Proceedings of 2024 Conference on Empirical Methods in Natural Language Processing. 2024, 11908−11941

[8]

Hu J. REINFORCE++: a simple and efficient approach for aligning large language models. 2025, arXiv preprint arXiv: 2501.03262v1

[9]

Yin Y, Yang S, Xie Y, Yang Z, Sun Y, Awadalla H H, Chen W, Zhou M. Segmenting text and learning their rewards for improved RLHF in language model. Transactions on Machine Learning Research, 2025, ISSN: 2835−8856

Rights & permissions

Higher Education Press

PDF (330KB)

Supplementary files

Highlights

Supplementary materials

332

Accesses

0

Citation

Detail

Sections
Recommended

/