About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
DIMA: Diversity-Preserving Imitation via Minimax Adversarial Learning for Multi-Turn Agents
Xufeng ZHOU , Linjing LI , Daniel Dajun ZENG
Large language model (LLM) agents hold strong promise for solvingmulti-turn interactive tasks, yet their progress is hindered by two fundamental obstacles: the inefficiency and bias of long-context processing, rooted in quadratic attention complexity and cognitive decay, and the diversity collapse induced by standard cross-entropy imitation learning, which uniformly penalizes non-expert tokens and erodes semantically reasonable alternatives learned during pre-training. To jointly address these issues, we propose an integrated framework consisting of: (1) a hierarchical sliding window (HSW) architecture that maintains global task coherence through persistent planning while focusing on local context via a sliding window, thus reducing computational cost and mitigating attention bias; and (2) a diversity-preserving imitation learning method named DIMA, formulated as a minimax adversarial game, which promotes policy diversity by selectively penalizing hard negatives sampled from an entropy-regularized adversary rather than uniformly suppressing all non-expert tokens. Extensive evaluations on interactive benchmarks including ScienceWorld, AlfWorld, and Webshop demonstrate that our approach significantly improves performance over strong baselines, particularly in multi-path settings (e.g., +2.4% pass@8 on ScienceWorld), and serves as a superior initialization for post-training, boosting downstream RFT fine-tuning by an average of 13.7% in pass@1 across all environments.
Multi-turn agents / Diversity-preserving imitation / Minimax adversarial learning / Long-horizon trajectories
Higher Education Press 2026
/
| 〈 |
|
〉 |