A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding

Yiming LEI , Guozhen PENG , Zeming LIU , Hui QIU , Haitao LENG , Shaoguo LIU , Tingting GAO , Qingjie LIU , Annan LI , Yunhong WANG

Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (9) : 2009377

PDF (7938KB)
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (9) :2009377 DOI: 10.1007/s11704-026-52159-9
Artificial Intelligence
RESEARCH ARTICLE
A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding
Author information +
History +
PDF (7938KB)

Abstract

Multimodal Large Language Models (MLLMs) have made remarkable progress in video understanding and consistently perform well on vision-centric benchmarks. However, existing benchmarks primarily evaluate factual or event-based comprehension, while neglecting audience insights. It is a critical yet underexplored dimension of video understanding, reflecting a deep comprehension of cognitive processes from the audience’s perspective. As a result, MLLMs, shaped by such benchmarks, often produce responses that are factually correct but misaligned with audience’s interests. To bridge this gap, we leverage audience insights derived from video comments as a direct proxy to guide the annotation process and introduce A3Bench, an audience-aligned benchmark for evaluating video audience insights with large-scale videos and high-quality multilingual comments. Furthermore, inspired by neuro-imaging studies, we propose Cognition Interaction of Thought (CIoT), a structured reasoning framework that emulates key aspects of cognitive processes. Extensive experiments on A3Bench reveal that current MLLMs struggle to understand audience insights, particularly compared to human-level understanding. In contrast, CIoT can improve the performance of these models, highlighting its potential to enhance the MLLMs’ capability of understanding audience insights in future research.

Graphical abstract

Keywords

multilingual video understanding / audience insights / CIoT / MLLMs

Cite this article

Download citation ▾
Yiming LEI, Guozhen PENG, Zeming LIU, Hui QIU, Haitao LENG, Shaoguo LIU, Tingting GAO, Qingjie LIU, Annan LI, Yunhong WANG. A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding. Front. Comput. Sci., 2026, 20 (9) : 2009377 DOI:10.1007/s11704-026-52159-9

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Imba. Short-form video dominance in 2024. imbaproduction.com/blog/short-form-video-dominance/ website, 2024

[2]

Cheng X, Su X, Yang B, Zarifis A, Mou J . Understanding users’ negative emotions and continuous usage intention in short video platforms. Electronic Commerce Research and Applications, 2023, 58: 101244

[3]

Liu Z, Wang H, Niu Z Y, Wu H, Che W, Liu T. Towards conversational recommendation over multi-type dialogs. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020, 1036−1049

[4]

Ran D, Zheng W, Li Y, Bian K, Zhang J, Deng X. Revenue and user traffic maximization in mobile short-video advertising. In: Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems. 2022, 1092−1100

[5]

Bassett D S, Gazzaniga M S . Understanding complexity in the human brain. Trends in Cognitive Sciences, 2011, 15( 5): 200–209

[6]

Zhang Z, Ding X, Liang X, Zhou Y, Qin B, Liu T . Brain and cognitive science inspired deep learning: a comprehensive survey. IEEE Transactions on Knowledge and Data Engineering, 2025, 37( 4): 1650–1671

[7]

Xu D, Zhao Z, Xiao J, Wu F, Zhang H, He X, Zhuang Y. Video question answering via gradually refined attention over appearance and motion. In: Proceedings of the 25th ACM International Conference on Multimedia. 2017, 1645−1653

[8]

Hong W, Cheng Y, Yang Z, Wang W, Wang L, Gu X, Huang S, Dong Y, Tang J. MotionBench: benchmarking and improving fine-grained video motion understanding for vision language models. In: Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 8450−8460

[9]

Li Y, Chen X, Hu B, Wang L, Shi H, Zhang M. VideoVista: a versatile benchmark for video understanding and reasoning. 2024, arXiv preprint arXiv: 2406.11303

[10]

Margulies D S, Ghosh S S, Goulas A, Falkiewicz M, Huntenburg J M, Langs G, Bezgin G, Eickhoff S B, Castellanos F X, Petrides M, Jefferies E, Smallwood J . Situating the default-mode network along a principal gradient of macroscale cortical organization. Proceedings of the National Academy of Sciences of the United States of America, 2016, 113( 44): 12574–12579

[11]

Ha D, Schmidhuber J. Recurrent world models facilitate policy evolution. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. 2018, 2455−2467

[12]

Burt J B, Demirtaş M, Eckner W J, Navejar N M, Ji J L, Martin W J, Bernacchia A, Anticevic A, Murray J D . Hierarchy of transcriptomic specialization across human cortex captured by structural neuroimaging topography. Nature Neuroscience, 2018, 21( 9): 1251–1259

[13]

Kober H, Barrett L F, Joseph J, Bliss-Moreau E, Lindquist K, Wager T D . Functional grouping and cortical–subcortical interactions in emotion: a meta-analysis of neuroimaging studies. NeuroImage, 2008, 42( 2): 998–1031

[14]

Punnakkal A R, Chandrasekaran A, Athanasiou N, Quirós-Ramírez A, Black M J. BABEL: bodies, action and behavior with English labels. In: Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 722−731

[15]

Bai S, Yang S, Bai J, Wang P, Zhang X, Lin J, Wang X, Zhou C, Zhou J. TouchStone: evaluating vision-language models by language models. 2023, arXiv preprint arXiv: 2308.16890

[16]

Fu C, Chen P, Shen Y, Qin Y, Zhang M, Lin X, Yang J, Zheng X, Li K, Sun X, Wu Y, Ji R, Shan C, He R. MME: a comprehensive evaluation benchmark for multimodal large language models. 2025, arXiv preprint arXiv: 2306.13394

[17]

Xu P, Shao W, Zhang K, Gao P, Liu S, Lei M, Meng F, Huang S, Qiao Y, Luo P . LVLM-EHUB: a comprehensive evaluation benchmark for large vision-language models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47( 3): 1877–1893

[18]

Yu W, Yang Z, Li L, Wang J, Lin K, Liu Z, Wang X, Wang L. MM-Vet: evaluating large multimodal models for integrated capabilities. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 57730−57754

[19]

Jang Y, Song Y, Kim C D, Yu Y, Kim Y, Kim G . Video question answering with spatio-temporal reasoning. International Journal of Computer Vision, 2019, 127( 10): 1385–1412

[20]

Wu B, Yu S, Chen Z, Tenenbaum J B, Gan C. STAR: a benchmark for situated reasoning in real-world videos. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021

[21]

Xiao J, Shang X, Yao A, Chua T S. NExT-QA: Next phase of question-answering to explaining temporal actions. In: Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 9772−9781

[22]

Zhang C, Lei Y, Liu Z, Leng H, Liu S, Gao T, Liu Q, Wang Y. SeriesBench: a benchmark for narrative-driven drama series understanding. In: Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 28995−29004

[23]

Lei Y, Zhang C, Liu Z, Leng H, Liu S, Gao T, Liu Q, Wang Y. GODBench: a benchmark for multimodal large language models in video comment art. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. 2025, 11884−11952

[24]

Li K, Wang Y, He Y, Li Y, Wang Y, Liu Y, Wang Z, Xu J, Chen G, Luo P, Wang L, Qiao Y. MVBench: a comprehensive multi-modal video understanding benchmark. In: Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 22195−22206

[25]

Chen X, Lin Y, Zhang Y, Huang W. AutoEval-video: an automatic benchmark for assessing large vision language models in open-ended video question answering. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 179−195

[26]

Liu Y, Li S, Liu Y, Wang Y, Ren S, Li L, Chen S, Sun X, Hou L. TempCompass: do video LLMs really understand videos? In: Proceedings of Findings of the Association for Computational Linguistics. 2024, 8731−8772

[27]

Ozaki S, Hayashi K, Oba M, Sakai Y, Kamigaito H, Watanabe T. BQA: body language question answering dataset for video large language models. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. 2025, 110−123

[28]

Brown T B, Mann B, Ryder N, Subbiah M, Kaplan J, et al. Language models are few-shot learners. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 159

[29]

OpenAI. GPT-4 technical report. 2024, arXiv preprint arXiv: 2303.08774

[30]

Xu Y, Hu L, Zhao J, Qiu Z, Xu K, Ye Y, Gu H . A survey on multilingual large language models: corpora, alignment, and bias. Frontiers of Computer Science, 2025, 19( 11): 1911362

[31]

DeepSeek-AI. DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. 2026, arXiv preprint arXiv: 2501.12948

[32]

Comanici G, Bieber E, Schaekermann M, Pasupat I, Sachdeva N, et al. Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. 2025, arXiv preprint arXiv: 2507.06261

[33]

Yang A, Li A, Yang B, Zhang B, Hui B, et al. Qwen3 technical report. 2025, arXiv preprint arXiv: 2505.09388

[34]

Qin L, Chen Q, Feng X, Wu Y, Zhang Y, Li Y, Li M, Che W, Yu P S . Large language models meet NLP: a survey. Frontiers of Computer Science, 2026, 20( 11): 2011361

[35]

Li K, He Y, Wang Y, Li Y, Wang W, Luo P, Wang Y, Wang L, Qiao Y . VideoChat: chat-centric video understanding. Science China Information Sciences, 2025, 68( 10): 200102

[36]

Lin B, Ye Y, Zhu B, Cui J, Ning M, Jin P, Yuan L. Video-LLaVA: learning united visual representation by alignment before projection. In: Proceedings of 2024 Conference on Empirical Methods in Natural Language Processing. 2024, 5971−5984

[37]

Bai S, Chen K, Liu X, Wang J, Ge W, et al. Qwen2.5-vl technical report. 2025, arXiv preprint arXiv: 2502.13923

[38]

ERNIE Team, Baidu. ERNIE 4.5 technical report. See ernie.baidu.com/blog/publication/ERNIE_Technical_Report website, 2025

[39]

Zhu J, Wang W, Chen Z, Liu Z, Ye S, et al. InternVL3: exploring advanced training and test-time recipes for open-source multimodal models. 2025, arXiv preprint arXiv: 2504.10479

[40]

Jang Y, Song Y, Yu Y, Kim Y, Kim G. TGIF-QA: toward spatio-temporal reasoning in visual question answering. In: Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition. 2017, 1359−1367

[41]

Yu Z, Xu D, Yu J, Yu T, Zhao Z, Zhuang Y, Tao D. ActivityNet-QA: a dataset for understanding complex web videos via question answering. In: Proceedings of the 33rd AAAI Conference on Artificial Intelligence. 2019, 9127−9134

[42]

Ning M, Zhu B, Xie Y, Lin B, Cui J, Yuan L, Chen D, Yuan L . Video-bench: a comprehensive benchmark and toolkit for evaluating video-based large language models. Computational Visual Media, 2026, 12( 1): 71–84

[43]

Cores D, Dorkenwald M, Mucientes M, Snoek C G M, Asano Y M. Lost in time: a new temporal benchmark for videOLLMs. In: Proceedings of the 36th British Machine Vision Conference 2025. 2025

[44]

Fang X, Mao K, Duan H, Zhao X, Li Y, Lin D, Chen K. MMBENCH-video: a long-form multi-shot benchmark for holistic video understanding. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 89098−89124

[45]

Fu C, Dai Y, Luo Y, Li L, Ren S, et al. Video-MME: the first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis. In: Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 24108−24118

[46]

Xu D, Chen W, Peng W, Zhang C, Xu T, Zhao X, Wu X, Zheng Y, Wang Y, Chen E . Large language models for generative information extraction: a survey. Frontiers of Computer Science, 2024, 18( 6): 186357

[47]

NLLB Team, Costa-Jussà M R, Cross J, Çelebi O, Elbayad M, et al. No language left behind: scaling human-centered machine translation. 2022, arXiv preprint arXiv: 2207.04672

[48]

Wilie B, Vincentio K, Winata G I, Cahyawijaya S, Li X, Lim Z Y, Soleman S, Mahendra R, Fung P, Bahar S, Purwarianti A. IndoNLU: benchmark and resources for evaluating Indonesian natural language understanding. In: Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing. 2020, 843−857

[49]

Zhang Q, Chen P, Feng L, Liu S, Li J, Zhao H, Chen M, Li H, Wang Y . PediaBench: a comprehensive Chinese pediatric dataset for benchmarking large language models. Frontiers of Computer Science, 2026, 20( 3): 2003902

[50]

Fu X, Hu Y, Li B, Feng Y, Wang H, Lin X, Roth D, Smith N A, Ma W C, Krishna R. BLINK: multimodal large language models can see but not perceive. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 148−166

[51]

Grill-Spector K, Malach R . The human visual cortex. Annual Review of Neuroscience, 2004, 27: 649–677

[52]

Binder J R, Desai R H, Graves W W, Conant L L . Where is the semantic system? A critical review and meta-analysis of 120 functional neuroimaging studies. Cerebral Cortex, 2009, 19( 12): 2767–2796

[53]

Pessoa L . On the relationship between emotion and cognition. Nature Reviews Neuroscience, 2008, 9( 2): 148–158

[54]

Miller E K, Cohen J D . An integrative theory of prefrontal cortex function. Annual Review of Neuroscience, 2001, 24: 167–202

[55]

Zhang Y, Wu J, Li W, Li B, Ma Z, Liu Z, Li C. LLaVA-video: video instruction tuning with synthetic data. Transactions on Machine Learning Research, 2025, 2025

[56]

Yao Y, Yu T, Zhang A, Wang C, Cui J, et al. MiniCPM-V: a GPT-4V level MLLM on your phone. 2024, arXiv preprint arXiv: 2408.01800

[57]

Zhang B, Li K, Cheng Z, Hu Z, Yuan Y, Chen G, Leng S, Jiang Y, Zhang H, Li X, Jin P, Zhang W, Wang F, Bing L, Zhao D. VideoLLaMA 3: frontier multimodal foundation models for image and video understanding. 2025, arXiv preprint arXiv: 2501.13106

[58]

OpenAI. GPT-4O system card. 2024, arXiv preprint arXiv: 2410.21276

[59]

OpenAI. Introducing GPT-4.1 in the API. See openai.com/ind ex/gpt-4-1/ website, 2025.

[60]

Zheng Y, Zhang R, Zhang J, Ye Y, Luo Z. LlamaFactory: unified efficient fine-tuning of 100+ language models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024, 400−410

[61]

Shi X, Liu Z, Lei Y, Zhang C, Leng H, Wang C, Liu Q, Che W, Wang Y. KwaiChat: a large-scale video-driven multilingual mixed-type dialogue corpus. In: Proceedings of Findings of the Association for Computational Linguistics. 2025, 2279−2294

Rights & permissions

Higher Education Press

PDF (7938KB)

Supplementary files

highlights

520

Accesses

0

Citation

Detail

Sections
Recommended

/