DIMA: Diversity-Preserving Imitation via Minimax Adversarial Learning for Multi-Turn Agents

Xufeng ZHOU , Linjing LI , Daniel Dajun ZENG

Front. Comput. Sci. ››

PDF (3713KB)
Front. Comput. Sci. ›› DOI: 10.1007/s11704-026-52121-9
RESEARCH ARTICLE
DIMA: Diversity-Preserving Imitation via Minimax Adversarial Learning for Multi-Turn Agents
Author information +
History +
PDF (3713KB)

Abstract

Large language model (LLM) agents hold strong promise for solvingmulti-turn interactive tasks, yet their progress is hindered by two fundamental obstacles: the inefficiency and bias of long-context processing, rooted in quadratic attention complexity and cognitive decay, and the diversity collapse induced by standard cross-entropy imitation learning, which uniformly penalizes non-expert tokens and erodes semantically reasonable alternatives learned during pre-training. To jointly address these issues, we propose an integrated framework consisting of: (1) a hierarchical sliding window (HSW) architecture that maintains global task coherence through persistent planning while focusing on local context via a sliding window, thus reducing computational cost and mitigating attention bias; and (2) a diversity-preserving imitation learning method named DIMA, formulated as a minimax adversarial game, which promotes policy diversity by selectively penalizing hard negatives sampled from an entropy-regularized adversary rather than uniformly suppressing all non-expert tokens. Extensive evaluations on interactive benchmarks including ScienceWorld, AlfWorld, and Webshop demonstrate that our approach significantly improves performance over strong baselines, particularly in multi-path settings (e.g., +2.4% pass@8 on ScienceWorld), and serves as a superior initialization for post-training, boosting downstream RFT fine-tuning by an average of 13.7% in pass@1 across all environments.

Keywords

Multi-turn agents / Diversity-preserving imitation / Minimax adversarial learning / Long-horizon trajectories

Cite this article

Download citation ▾
Xufeng ZHOU, Linjing LI, Daniel Dajun ZENG. DIMA: Diversity-Preserving Imitation via Minimax Adversarial Learning for Multi-Turn Agents. Front. Comput. Sci. DOI:10.1007/s11704-026-52121-9

登录浏览全文

4963

注册一个新账户 忘记密码

References

RIGHTS & PERMISSIONS

Higher Education Press 2026

PDF (3713KB)

24

Accesses

0

Citation

Detail

Sections
Recommended

/