Consistent-point: consistent pseudo-points for semi-supervised crowd counting and localization

Yuda ZOU , Zelong LIU , Yuliang GU , Bo DU , Yongchao XU

Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (7) : 2107339

PDF (4723KB)
Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (7) :2107339 DOI: 10.1007/s11704-026-51063-6
Artificial Intelligence
RESEARCH ARTICLE
Consistent-point: consistent pseudo-points for semi-supervised crowd counting and localization
Author information +
History +
PDF (4723KB)

Abstract

Crowd counting and localization are critical for applications such as public security and traffic management. While existing methods have achieved impressive results, they rely heavily on extensive manual annotations. This paper proposes a novel point-localization-based semi-supervised crowd counting and localization method termed Consistent-Point. We identify and address two key inconsistencies of pseudo-points that have not been adequately explored. To enhance their position consistency, we aggregate the positions of neighboring auxiliary proposal-points, while an instance-wise uncertainty calibration is proposed to alleviate the class consistency of pseudo-points. By generating higher-quality pseudo-points with enhanced consistency, Consistent-Point provides more stable and effective supervision during training, yielding superior crowd counting and localization performance. Extensive experiments across five widely used datasets and three different labeled ratio settings demonstrate that our method achieves state-of-the-art performance in crowd localization while also attaining impressive crowd counting results.

Graphical abstract

Keywords

crowd counting / object counting / crowd localization / semi-supervised learning

Cite this article

Download citation ▾
Yuda ZOU, Zelong LIU, Yuliang GU, Bo DU, Yongchao XU. Consistent-point: consistent pseudo-points for semi-supervised crowd counting and localization. Front. Comput. Sci., 2027, 21 (7) : 2107339 DOI:10.1007/s11704-026-51063-6

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Liang D, Chen X, Xu W, Zhou Y, Bai X . TransCrowd: weakly-supervised crowd counting with transformers. Science China Information Sciences, 2022, 65( 6): 160104

[2]

Han T, Bai L, Gao J, Wang Q, Ouyang W. DR.VIC: decomposition and reasoning for video individual counting. In: Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, 3083−3092

[3]

Hui X, Wu Q, Rahmani H, Liu J. Class-agnostic object counting with text-to-image diffusion model. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 1−18

[4]

Zhu H, Yuan J, Yang Z, Guo Y, Wang Z, Zhong X, He S. Zero-shot object counting with good exemplars. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 368−385

[5]

Ðukić N, Lukežič A, Zavrtanik V, Kristan M. A low-shot object counting network with iterative prototype adaptation. In: Proceedings of 2023 IEEE/CVF International Conference on Computer Vision. 2023, 18826−18835

[6]

Sui J, Ding S, Huang X, Yu Y, Liu R, Xia B, Ding Z, Xu L, Zhang H, Yu C, Bu D . A survey on deep learning-based algorithms for the traveling salesman problem. Frontiers of Computer Science, 2025, 19( 6): 196322

[7]

Guo M, Sheng H, Zhang Z, Huang Y, Chen X, Wang C, Zhang J . CW-YOLO: joint learning for mask wearing detection in low-light conditions. Frontiers of Computer Science, 2023, 17( 6): 176710

[8]

Shang C, Ai H, Yang Y . Crowd counting via learning perspective for multi-scale multi-view Web images. Frontiers of Computer Science, 2019, 13( 3): 579–587

[9]

Ma Z, Wei X, Hong X, Lin H, Qiu Y, Gong Y. Learning to count via unbalanced optimal transport. In: Proceedings of the 35th AAAI Conference on Artificial Intelligence. 2021, 2319−2327

[10]

Abousamra S, Hoai M, Samaras D, Chen C. Localization in the crowd with topological constraints. In: Proceedings of the 35th AAAI Conference on Artificial Intelligence. 2021, 872−881

[11]

Liu W, Salzmann M, Fua P. Context-aware crowd counting. In: Proceedings of 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019, 5099−5108

[12]

Han T, Bai L, Liu L, Ouyang W. STEERER: resolving scale variations for counting and localization via selective inheritance learning. In: Proceedings of 2023 IEEE/CVF International Conference on Computer Vision. 2023, 21848−21859

[13]

Huang Z K, Chen W T, Chiang Y C, Kuo S Y, Yang M H. Counting crowds in bad weather. In: Proceedings of 2023 IEEE/CVF International Conference on Computer Vision. 2023, 23308−23319

[14]

Liu X, Li G, Qi Y, Han Z, van den Hengel A, Sebe N, Yang M H, Huang Q. Consistency-aware anchor pyramid network for crowd localization. IEEE Transactions on Pattern Analysis and Machine Intelligence, advance online publication, Apr. 29, 2024. doi:10.1109/TPAMI.2024.3392013

[15]

Wan J, Chan A B. Modeling noisy annotations for crowd counting. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 285

[16]

Wan J, Kumar N S, Chan A B . Fine-grained crowd counting. IEEE Transactions on Image Processing, 2021, 30: 2114–2126

[17]

Sun G, An Z, Liu Y, Liu C, Sakaridis C, Fan D P, Van Gool L. Indiscernible object counting in underwater scenes. In: Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023, 13791−13801

[18]

Zhang Q, Chan A B. Wide-area crowd counting via ground-plane density maps and multi-view fusion CNNs. In: Proceedings of 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019, 8297−8306

[19]

Zhang Q, Lin W, Chan A B. Cross-view cross-scene multi-view crowd counting. In: Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 557−567

[20]

Ranasinghe Y, Nair N G, Bandara W G C, Patel V M. CrowdDiff: multi-hypothesis crowd density estimation using diffusion models. In: Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 12809−12819

[21]

Peng Z, Chan S H G. Single domain generalization for crowd counting. In: Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 28025−28034

[22]

Guo M, Yuan L, Yan Z, Chen B, Wang Y, Ye Q. Regressor-segmenter mutual prompt learning for crowd counting. In: Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 28380−28389

[23]

Song Q, Wang C, Jiang Z, Wang Y, Tai Y, Wang C, Li J, Huang F, Wu Y. Rethinking counting and localization in crowds: a purely point-based framework. In: Proceedings of 2021 IEEE/CVF International Conference on Computer Vision. 2021, 3365−3374

[24]

Sam D B, Peri S V, Sundararaman M N, Kamath A, Babu R V . Locate, size, and count: accurately resolving people in dense crowds via detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43( 8): 2739–2751

[25]

Liu C, Lu H, Cao Z, Liu T. Point-query quadtree for crowd counting, localization, and more. In: Proceedings of 2023 IEEE/CVF International Conference on Computer Vision. 2023, 1676−1685

[26]

Liang D, Xu W, Bai X. An end-to-end transformer model for crowd localization. In: Proceedings of the 17th European Conference on Computer Vision. 2022, 38−54

[27]

Liang D, Xie J, Zou Z, Ye X, Xu W, Bai X. CrowdCLIP: unsupervised crowd counting via vision-language model. In: Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023, 2893−2903

[28]

Knobel L, Han T, Asano Y M. Learning to count without annotations. In: Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 22924−22934

[29]

Pelhan J, Lukežič A, Zavrtanik V, Kristan M. DAVE-a detect-and-verify paradigm for low-shot counting. In: Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 23293−23302

[30]

D’Alessandro A, Mahdavi-Amiri A, Hamarneh G. AFreeCA: annotation-free counting for all. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 75−91

[31]

Zhao H, Min W, Xu J, Wang Q, Zou Y, Fu Q . Scene-adaptive crowd counting method based on meta learning with dual-input network DMNet. Frontiers of Computer Science, 2023, 17( 1): 171304

[32]

Jiang X, Liu H, Zhang L, Li G, Xu M, Lv P, Zhou B . Transferring priors from virtual data for crowd counting in real world. Frontiers of Computer Science, 2022, 16( 3): 163314

[33]

Wan J, Wu Q, Lin W, Chan A. Robust zero-shot crowd counting and localization with adaptive resolution SAM. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 478−495

[34]

Liu X, van de Weijer J, Bagdanov A D . Exploiting unlabeled data in CNNs by self-supervised learning to rank. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, 41( 8): 1862–1878

[35]

Sindagi V A, Yasarla R, Babu D S, Babu R V, Patel V M. Learning to count in the crowd from limited labeled data. In: Proceedings of the 16th European Conference on Computer Vision. 2020, 212−229

[36]

Liu Y, Liu L, Wang P, Zhang P, Lei Y. Semi-supervised crowd counting via self-training on surrogate tasks. In: Proceedings of the 16th European Conference on Computer Vision. 2020, 242−259

[37]

Lin H, Ma Z, Hong X, Wang Y, Su Z. Semi-supervised crowd counting via density agency. In: Proceedings of the 30th ACM International Conference on Multimedia. 2022, 1416−1426

[38]

Lin H, Ma Z, Ji R, Wang Y, Su Z, Hong X, Meng D . Semi-supervised counting via pixel-by-pixel density distribution modeling. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47( 5): 3625–3638

[39]

Meng Y, Zhang H, Zhao Y, Yang X, Qian X, Huang X, Zheng Y. Spatial uncertainty-aware semi-supervised crowd counting. In: Proceedings of 2021 IEEE/CVF International Conference on Computer Vision. 2021, 15549−15559

[40]

Wang X, Zhan Y, Zhao Y, Yang T, Ruan Q . Semi-supervised crowd counting with spatial temporal consistency and pseudo-label filter. IEEE Transactions on Circuits and Systems for Video Technology, 2023, 33( 8): 4190–4203

[41]

Li C, Hu X, Abousamra S, Chen C. Calibrating uncertainty for semi-supervised crowd counting. In: Proceedings of 2023 IEEE/CVF International Conference on Computer Vision. 2023, 16685−16695

[42]

Lin W, Chan A B. Optimal transport minimization: crowd localization on density maps for semi-supervised counting. In: Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023, 21663−21673

[43]

Zhang S, Ke W, Liu S, Hong X, Zhang T. Boosting semi-supervised crowd counting with scale-based active learning. In: Proceedings of the 32nd ACM International Conference on Multimedia. 2024, 8681−8690

[44]

Zou Y, Xiao X, Zhou P, Sun Z, Du B, Xu Y. Shifted Autoencoders for point annotation restoration in object counting. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 113−130

[45]

Ma Z, Wei X, Hong X, Gong Y. Bayesian loss for crowd count estimation with point supervision. In: Proceedings of 2019 IEEE/CVF International Conference on Computer Vision. 2019, 6142−6151

[46]

Wan J, Wu Q, Chan A B . Modeling noisy annotations for point-wise supervision. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45( 12): 15065–15080

[47]

Lin H, Ma Z, Ji R, Wang Y, Hong X. Boosting crowd counting via multifaceted attention. In: Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, 19628−19637

[48]

Tarvainen A, Valpola H. Mean teachers are better role models: weight-averaged consistency targets improve semi-supervised deep learning results. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 1195−1204

[49]

Cheng Z Q, Dai Q, Li H, Song J, Wu X, Hauptmann A G. Rethinking spatial invariance of convolutional networks for object counting. In: Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, 19606−19616

[50]

Wan J, Chan A. Adaptive density map generation for crowd counting. In: Proceedings of 2019 IEEE/CVF International Conference on Computer Vision. 2019, 1130−1139

[51]

Wan J, Wang Q, Chan A B . Kernel-based density map generation for dense object counting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44( 3): 1357–1370

[52]

Li Y, Zhang X, Chen D. CSRNet: dilated convolutional neural networks for understanding the highly congested scenes. In: Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2018, 1091−1100

[53]

Wang B, Liu H, Samaras D, Hoai M. Distribution matching for crowd counting. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 135

[54]

Ma Z, Wei X, Hong X, Gong Y. Learning scales from points: a scale-aware probabilistic model for crowd counting. In: Proceedings of the 38th ACM International Conference on Multimedia. 2020, 220−228

[55]

Mo H, Zhang X, Tan J, Yang C, Gu Q, Hang B, Ren W. CountFormer: multi-view crowd counting transformer. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 20−40

[56]

Wu J, Li Z, Qu W, Zhou Y . One shot crowd counting with deep scale adaptive neural network. Electronics, 2019, 8( 6): 701

[57]

Xu C, Qiu K, Fu J, Bai S, Xu Y, Bai X. Learn to scale: generating multipolar normalized density maps for crowd counting. In: Proceedings of 2019 IEEE/CVF International Conference on Computer Vision. 2019, 8382−8390

[58]

Chen I H, Chen W T, Liu Y W, Yang M H, Kuo S Y. Improving point-based crowd counting and localization based on auxiliary point guidance. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 428−444

[59]

Chen J, Wang Z . Multi-task semi-supervised crowd counting via global to local self-correction. Pattern Recognition, 2023, 140: 109506

[60]

Liu X, van de Weijer J, Bagdanov A D. Leveraging unlabeled data for crowd counting by learning to rank. In: Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2018, 7661−7669

[61]

Kuhn H W . The Hungarian method for the assignment problem. Naval Research Logistics Quarterly, 1955, 2( 1-2): 83–97

[62]

Qian Y, Zhang L, Guo Z, Hong X, Arandjelović O, Donovan C R . Perspective-assisted prototype-based learning for semi-supervised crowd counting. Pattern Recognition, 2025, 158: 111073

[63]

Feng W, Wang Y, Ma L, Yuan Y, Zhang C. Temporal knowledge consistency for unsupervised visual representation learning. In: Proceedings of 2021 IEEE/CVF International Conference on Computer Vision. 2021, 10170−10180

[64]

Zhang Y, Zhou D, Chen S, Gao S, Ma Y. Single-image crowd counting via multi-column convolutional neural network. In: Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition. 2016, 589−597

[65]

Idrees H, Tayyab M, Athrey K, Zhang D, Al-Maadeed S, Rajpoot N, Shah M. Composition loss for counting, density map estimation and localization in dense crowds. In: Proceedings of the 15th European Conference on Computer Vision. 2018, 532−546

[66]

Sindagi V A, Yasarla R, Patel V M . JHU-CROWD++: large-scale crowd counting dataset and a benchmark method. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44( 5): 2594–2609

[67]

Wang Q, Gao J, Lin W, Li X . NWPU-crowd: a large-scale benchmark for crowd counting and localization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43( 6): 2141–2149

[68]

Qian Y, Hong X, Guo Z, Arandjelović O, Donovan C R . Semi-supervised crowd counting with contextual modeling: facilitating holistic understanding of crowd scenes. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34( 9): 8230–8241

[69]

Liang D, Xu W, Zhu Y, Zhou Y . Focal inverse distance transform maps for crowd localization. IEEE Transactions on Multimedia, 2023, 25: 6040–6052

[70]

Wei X, Qiu Y, Ma Z, Hong X, Gong Y . Semi-supervised crowd counting via multiple representation learning. IEEE Transactions on Image Processing, 2023, 32: 5220–5230

[71]

Xu Y, Zhong Z, Lian D, Li J, Li Z, Xu X, Gao S. Crowd counting with partial annotations in an image. In: Proceedings of 2021 IEEE/CVF International Conference on Computer Vision. 2021, 15570−15579

[72]

Wang X, Zhan Y, Zhao Y, Yang T, Ruan Q . Hybrid perturbation strategy for semi-supervised crowd counting. IEEE Transactions on Image Processing, 2024, 33: 1227–1240

[73]

Chen X, Yuan Y, Zeng G, Wang J. Semi-supervised semantic segmentation with cross pseudo supervision. In: Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 2613−2622

RIGHTS & PERMISSIONS

Higher Education Press

PDF (4723KB)

Supplementary files

Highlights

378

Accesses

0

Citation

Detail

Sections
Recommended

/