About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding
Yiming LEI , Guozhen PENG , Zeming LIU , Hui QIU , Haitao LENG , Shaoguo LIU , Tingting GAO , Qingjie LIU , Annan LI , Yunhong WANG
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (9) : 2009377
Multimodal Large Language Models (MLLMs) have made remarkable progress in video understanding and consistently perform well on vision-centric benchmarks. However, existing benchmarks primarily evaluate factual or event-based comprehension, while neglecting audience insights. It is a critical yet underexplored dimension of video understanding, reflecting a deep comprehension of cognitive processes from the audience’s perspective. As a result, MLLMs, shaped by such benchmarks, often produce responses that are factually correct but misaligned with audience’s interests. To bridge this gap, we leverage audience insights derived from video comments as a direct proxy to guide the annotation process and introduce A3Bench, an audience-aligned benchmark for evaluating video audience insights with large-scale videos and high-quality multilingual comments. Furthermore, inspired by neuro-imaging studies, we propose Cognition Interaction of Thought (CIoT), a structured reasoning framework that emulates key aspects of cognitive processes. Extensive experiments on A3Bench reveal that current MLLMs struggle to understand audience insights, particularly compared to human-level understanding. In contrast, CIoT can improve the performance of these models, highlighting its potential to enhance the MLLMs’ capability of understanding audience insights in future research.
multilingual video understanding / audience insights / CIoT / MLLMs
| [1] |
Imba. Short-form video dominance in 2024. imbaproduction.com/blog/short-form-video-dominance/ website, 2024 |
| [2] |
|
| [3] |
Liu Z, Wang H, Niu Z Y, Wu H, Che W, Liu T. Towards conversational recommendation over multi-type dialogs. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020, 1036−1049 |
| [4] |
Ran D, Zheng W, Li Y, Bian K, Zhang J, Deng X. Revenue and user traffic maximization in mobile short-video advertising. In: Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems. 2022, 1092−1100 |
| [5] |
|
| [6] |
|
| [7] |
Xu D, Zhao Z, Xiao J, Wu F, Zhang H, He X, Zhuang Y. Video question answering via gradually refined attention over appearance and motion. In: Proceedings of the 25th ACM International Conference on Multimedia. 2017, 1645−1653 |
| [8] |
Hong W, Cheng Y, Yang Z, Wang W, Wang L, Gu X, Huang S, Dong Y, Tang J. MotionBench: benchmarking and improving fine-grained video motion understanding for vision language models. In: Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 8450−8460 |
| [9] |
Li Y, Chen X, Hu B, Wang L, Shi H, Zhang M. VideoVista: a versatile benchmark for video understanding and reasoning. 2024, arXiv preprint arXiv: 2406.11303 |
| [10] |
|
| [11] |
Ha D, Schmidhuber J. Recurrent world models facilitate policy evolution. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. 2018, 2455−2467 |
| [12] |
|
| [13] |
|
| [14] |
Punnakkal A R, Chandrasekaran A, Athanasiou N, Quirós-Ramírez A, Black M J. BABEL: bodies, action and behavior with English labels. In: Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 722−731 |
| [15] |
Bai S, Yang S, Bai J, Wang P, Zhang X, Lin J, Wang X, Zhou C, Zhou J. TouchStone: evaluating vision-language models by language models. 2023, arXiv preprint arXiv: 2308.16890 |
| [16] |
Fu C, Chen P, Shen Y, Qin Y, Zhang M, Lin X, Yang J, Zheng X, Li K, Sun X, Wu Y, Ji R, Shan C, He R. MME: a comprehensive evaluation benchmark for multimodal large language models. 2025, arXiv preprint arXiv: 2306.13394 |
| [17] |
|
| [18] |
Yu W, Yang Z, Li L, Wang J, Lin K, Liu Z, Wang X, Wang L. MM-Vet: evaluating large multimodal models for integrated capabilities. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 57730−57754 |
| [19] |
|
| [20] |
Wu B, Yu S, Chen Z, Tenenbaum J B, Gan C. STAR: a benchmark for situated reasoning in real-world videos. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021 |
| [21] |
Xiao J, Shang X, Yao A, Chua T S. NExT-QA: Next phase of question-answering to explaining temporal actions. In: Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 9772−9781 |
| [22] |
Zhang C, Lei Y, Liu Z, Leng H, Liu S, Gao T, Liu Q, Wang Y. SeriesBench: a benchmark for narrative-driven drama series understanding. In: Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 28995−29004 |
| [23] |
Lei Y, Zhang C, Liu Z, Leng H, Liu S, Gao T, Liu Q, Wang Y. GODBench: a benchmark for multimodal large language models in video comment art. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. 2025, 11884−11952 |
| [24] |
Li K, Wang Y, He Y, Li Y, Wang Y, Liu Y, Wang Z, Xu J, Chen G, Luo P, Wang L, Qiao Y. MVBench: a comprehensive multi-modal video understanding benchmark. In: Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 22195−22206 |
| [25] |
Chen X, Lin Y, Zhang Y, Huang W. AutoEval-video: an automatic benchmark for assessing large vision language models in open-ended video question answering. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 179−195 |
| [26] |
Liu Y, Li S, Liu Y, Wang Y, Ren S, Li L, Chen S, Sun X, Hou L. TempCompass: do video LLMs really understand videos? In: Proceedings of Findings of the Association for Computational Linguistics. 2024, 8731−8772 |
| [27] |
Ozaki S, Hayashi K, Oba M, Sakai Y, Kamigaito H, Watanabe T. BQA: body language question answering dataset for video large language models. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. 2025, 110−123 |
| [28] |
Brown T B, Mann B, Ryder N, Subbiah M, Kaplan J, et al. Language models are few-shot learners. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 159 |
| [29] |
OpenAI. GPT-4 technical report. 2024, arXiv preprint arXiv: 2303.08774 |
| [30] |
|
| [31] |
DeepSeek-AI. DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. 2026, arXiv preprint arXiv: 2501.12948 |
| [32] |
Comanici G, Bieber E, Schaekermann M, Pasupat I, Sachdeva N, et al. Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. 2025, arXiv preprint arXiv: 2507.06261 |
| [33] |
Yang A, Li A, Yang B, Zhang B, Hui B, et al. Qwen3 technical report. 2025, arXiv preprint arXiv: 2505.09388 |
| [34] |
|
| [35] |
|
| [36] |
Lin B, Ye Y, Zhu B, Cui J, Ning M, Jin P, Yuan L. Video-LLaVA: learning united visual representation by alignment before projection. In: Proceedings of 2024 Conference on Empirical Methods in Natural Language Processing. 2024, 5971−5984 |
| [37] |
Bai S, Chen K, Liu X, Wang J, Ge W, et al. Qwen2.5-vl technical report. 2025, arXiv preprint arXiv: 2502.13923 |
| [38] |
ERNIE Team, Baidu. ERNIE 4.5 technical report. See ernie.baidu.com/blog/publication/ERNIE_Technical_Report website, 2025 |
| [39] |
Zhu J, Wang W, Chen Z, Liu Z, Ye S, et al. InternVL3: exploring advanced training and test-time recipes for open-source multimodal models. 2025, arXiv preprint arXiv: 2504.10479 |
| [40] |
Jang Y, Song Y, Yu Y, Kim Y, Kim G. TGIF-QA: toward spatio-temporal reasoning in visual question answering. In: Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition. 2017, 1359−1367 |
| [41] |
Yu Z, Xu D, Yu J, Yu T, Zhao Z, Zhuang Y, Tao D. ActivityNet-QA: a dataset for understanding complex web videos via question answering. In: Proceedings of the 33rd AAAI Conference on Artificial Intelligence. 2019, 9127−9134 |
| [42] |
|
| [43] |
Cores D, Dorkenwald M, Mucientes M, Snoek C G M, Asano Y M. Lost in time: a new temporal benchmark for videOLLMs. In: Proceedings of the 36th British Machine Vision Conference 2025. 2025 |
| [44] |
Fang X, Mao K, Duan H, Zhao X, Li Y, Lin D, Chen K. MMBENCH-video: a long-form multi-shot benchmark for holistic video understanding. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 89098−89124 |
| [45] |
Fu C, Dai Y, Luo Y, Li L, Ren S, et al. Video-MME: the first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis. In: Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 24108−24118 |
| [46] |
|
| [47] |
NLLB Team, Costa-Jussà M R, Cross J, Çelebi O, Elbayad M, et al. No language left behind: scaling human-centered machine translation. 2022, arXiv preprint arXiv: 2207.04672 |
| [48] |
Wilie B, Vincentio K, Winata G I, Cahyawijaya S, Li X, Lim Z Y, Soleman S, Mahendra R, Fung P, Bahar S, Purwarianti A. IndoNLU: benchmark and resources for evaluating Indonesian natural language understanding. In: Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing. 2020, 843−857 |
| [49] |
|
| [50] |
Fu X, Hu Y, Li B, Feng Y, Wang H, Lin X, Roth D, Smith N A, Ma W C, Krishna R. BLINK: multimodal large language models can see but not perceive. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 148−166 |
| [51] |
|
| [52] |
|
| [53] |
|
| [54] |
|
| [55] |
Zhang Y, Wu J, Li W, Li B, Ma Z, Liu Z, Li C. LLaVA-video: video instruction tuning with synthetic data. Transactions on Machine Learning Research, 2025, 2025 |
| [56] |
Yao Y, Yu T, Zhang A, Wang C, Cui J, et al. MiniCPM-V: a GPT-4V level MLLM on your phone. 2024, arXiv preprint arXiv: 2408.01800 |
| [57] |
Zhang B, Li K, Cheng Z, Hu Z, Yuan Y, Chen G, Leng S, Jiang Y, Zhang H, Li X, Jin P, Zhang W, Wang F, Bing L, Zhao D. VideoLLaMA 3: frontier multimodal foundation models for image and video understanding. 2025, arXiv preprint arXiv: 2501.13106 |
| [58] |
OpenAI. GPT-4O system card. 2024, arXiv preprint arXiv: 2410.21276 |
| [59] |
OpenAI. Introducing GPT-4.1 in the API. See openai.com/ind ex/gpt-4-1/ website, 2025. |
| [60] |
Zheng Y, Zhang R, Zhang J, Ye Y, Luo Z. LlamaFactory: unified efficient fine-tuning of 100+ language models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024, 400−410 |
| [61] |
Shi X, Liu Z, Lei Y, Zhang C, Leng H, Wang C, Liu Q, Che W, Wang Y. KwaiChat: a large-scale video-driven multilingual mixed-type dialogue corpus. In: Proceedings of Findings of the Association for Computational Linguistics. 2025, 2279−2294 |
Higher Education Press
/
| 〈 |
|
〉 |