Perception-guided accuracy estimation: a universal framework for robust model evaluation

Hao SUN , Zhongyi HAN , Yilong YIN

Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (7) : 2107340

PDF (5181KB)
Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (7) :2107340 DOI: 10.1007/s11704-026-51697-6
Artificial Intelligence
RESEARCH ARTICLE
Perception-guided accuracy estimation: a universal framework for robust model evaluation
Author information +
History +
PDF (5181KB)

Abstract

Model evaluation is crucial for ensuring machine learning models meet performance standards, becoming especially vital under distribution shifts where reliable deployment in dynamic, non-stationary environments requires robust evaluation strategies. Given the inaccessibility of supervised information about target domains, existing methods utilize statistical metrics or common patterns based on restrictive mathematical assumptions. Consequently, these conventional approaches often lead to task-specific overfitting and sub-optimal evaluation performance. To overcome these limitations, we propose Human-like Perception Training (HPT), a novel and universal framework that approaches model evaluation from a human-like visual perspective by focusing on feature-level insights. This approach offers strong universality and robustness while minimizing the reliance on strict mathematical assumptions. Specifically, HPT incorporates two novel modules: 1) The Human-like Perception Representing (HPR) module quantifies a given model’s representational capability by mimicking human visual perception, offering a distinct evaluation perspective. 2) Building on this perception representation, the Human-like Perception Mentoring (HPM) module guides the regression model to emulate human-like decisions through the incorporation of local perception priors and a novel coherent contrastive learning loss. Extensive experiments on standard benchmarks demonstrate that HPT achieves a strong correlation with true model accuracy, precisely estimates model performance, and significantly outperforms prior state-of-the-art methods.

Graphical abstract

Keywords

model evaluation / distribution shift / accuracy estimation

Cite this article

Download citation ▾
Hao SUN, Zhongyi HAN, Yilong YIN. Perception-guided accuracy estimation: a universal framework for robust model evaluation. Front. Comput. Sci., 2027, 21 (7) : 2107340 DOI:10.1007/s11704-026-51697-6

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Quiñonero-Candela J, Sugiyama M, Schwaighofer A, Lawrence N D. Dataset Shift in Machine Learning. Cambridge: MIT Press, 2009, 5

[2]

Finlayson S G, Bowers J D, Ito J, Zittrain J L, Beam A L, Kohane I S . Adversarial attacks on medical machine learning: emerging vulnerabilities demand new conversations. Science, 2019, 363( 6433): 1287–1289

[3]

Kouw W M, Loog M . A review of domain adaptation without target labels. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43( 3): 766–785

[4]

Sun T, Segu M, Postels J, Wang Y, van Gool L, Schiele B, Tombari F, Yu F. SHIFT: A synthetic driving dataset for continuous multi-task domain adaptation. In: Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, 21339−21350

[5]

Liu H, Wang J, Long M. Cycle self-training for domain adaptation. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021

[6]

Li J, Chen E, Ding Z, Zhu L, Lu K, Shen H T . Maximum density divergence for domain adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43( 11): 3918–3930

[7]

Ramdas A, Reddi S J, Poczos B, Singh A, Wasserman L. On the decreasing power of kernel and distance based nonparametric hypothesis tests in high dimensions. In: Proceedings of the 29th AAAI Conference on Artificial Intelligence. 2015, 3571−3577

[8]

Rabanser S, Günnemann S, Lipton Z C. Failing loudly: an empirical study of methods for detecting dataset shift. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. 2019

[9]

Garg S, Balakrishnan S, Lipton Z C, Neyshabur B, Sedghi H. Leveraging unlabeled data to predict out-of-distribution performance. In: Proceedings of the 10th International Conference on Learning Representations. 2022

[10]

Chen J, Liu F, Avci B, Wu X, Liang Y, Jha S. Detecting errors and estimating accuracy on unlabeled data with self-training ensembles. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021

[11]

Yu Y, Yang Z, Wei A, Ma Y, Steinhardt J. Predicting out-of-distribution error with the projection norm. In: Proceedings of the 39th International Conference on Machine Learning. 2022, 25721−25746

[12]

Baek C, Jiang Y, Raghunathan A, Kolter Z. Agreement-on-the-line: Predicting the performance of neural networks under distribution shift. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022

[13]

van der Maaten L, Hinton G . Visualizing data using t-SNE. Journal of Machine Learning Research, 2008, 9( 86): 2579–2605

[14]

Natekar P, Sharma M. Representation based complexity measures for predicting generalization in deep learning. 2020, arXiv preprint arXiv: 2012.02775

[15]

Jiang Y, Natekar P, Sharma M, Aithal S K, Kashyap D, Subramanyam N, Lassance C, Roy D M, Dziugaite G K, Gunasekar S, Guyon I, Foret P, Yak S, Mobahi H, Neyshabur B, Bengio S. Methods and analysis of the first competition in predicting generalization of deep learning. In: Proceedings of NeurIPS 2020 Competition and Demonstration Track. 2021, 170−190

[16]

Deng W, Zheng L. Are labels always necessary for classifier accuracy evaluation? In: Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, 15064−15073

[17]

Deng W, Gould S, Zheng L. What does rotation prediction tell us about classifier accuracy under varying testing environments? In: Proceedings of the 38th International Conference on Machine Learning. 2021, 2579−2589

[18]

Guillory D, Shankar V, Ebrahimi S, Darrell T, Schmidt L. Predicting with confidence on unseen distributions. In: Proceedings of 2021 IEEE/CVF International Conference on Computer Vision. 2021, 1114−1124

[19]

Yosinski J, Clune J, Bengio Y, Lipson H. How transferable are features in deep neural networks? In: Advances In: Proceedings of the 28th Annual Conference on Neural Information Processing Systems. 2014, 3320−3328

[20]

Szegedy C, Zaremba W, Sutskever I, Bruna J, Erhan D, Goodfellow I J, Fergus R. Intriguing properties of neural networks. In: Proceedings of the 2nd International Conference on Learning Representations. 2014

[21]

Shiming C, Bowen D, Salman K, Fahad Shahbaz K. Interpretable zero-shot learning with locally-aligned vision-language model. In: ICCV. 2025

[22]

Chen S, Hong Z, You X, Shao L . Semantics-conditioned generative zero-shot learning via feature refinement. International Journal of Computer Vision, 2025, 133( 7): 4504–4521

[23]

Long M, Zhu H, Wang J, Jordan M I. Deep transfer learning with joint adaptation networks. In: Proceedings of the 34th International Conference on Machine Learning. 2017, 2208−2217

[24]

Ganin Y, Ustinova E, Ajakan H, Germain P, Larochelle H, Laviolette F, Marchand M, Lempitsky V . Domain-adversarial training of neural networks. The Journal of Machine Learning Research, 2016, 17( 1): 2096–2030

[25]

Chen S, Hong Z, Hou W, Xie G S, Song Y, Zhao J, You X, Yan S, Shao L . TransZero++: cross attribute-guided transformer for zero-shot learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45( 11): 12844–12861

[26]

Ben-David S, Blitzer J, Crammer K, Pereira F. Analysis of representations for domain adaptation. In: Proceedings of the 20th Annual Conference on Neural Information Processing Systems. 2006

[27]

Mansour Y, Mohri M, Rostamizadeh A. Domain adaptation: learning bounds and algorithms. In: Proceedings of the 22nd Conference on Learning Theory. 2009

[28]

Cortes C, Mohri M, Medina A M . Adaptation based on generalized discrepancy. The Journal of Machine Learning Research, 2019, 20( 1): 1–30

[29]

Zhang Y, Liu T, Long M, Jordan M. Bridging theory and algorithm for domain adaptation. In: Proceedings of the 36th International Conference on Machine Learning. 2019, 7404−7413

[30]

Ben-David S, Luu T, Lu T, Pál D. Impossibility theorems for domain adaptation. In: Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. 2010

[31]

Zeiler M D, Fergus R. Visualizing and understanding convolutional networks. In: Proceedings of the 13th European Conference on Computer Vision. 2014, 818−833

[32]

Pearson K . LIII. On lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 1901, 2( 11): 559–572

[33]

McInnes L, Healy J, Melville J. UMAP: uniform manifold approximation and projection for dimension reduction. 2020, arXiv preprint arXiv: 1802.03426.

[34]

Wang Y, Jiang Y, Li J, Ni B, Dai W, Li C, Xiong H, Li T. Contrastive regression for domain adaptation on gaze estimation. In: Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, 19354−19363

[35]

Krizhevsky A. Learning multiple layers of features from tiny images. Toronto: University of Toronto, 2009

[36]

Netzer Y, Wang T, Coates A, Bissacco A, Wu B, Ng A Y. Reading digits in natural images with unsupervised feature learning. In: Proceedings of NIPS Workshop on Deep Learning and Unsupervised Feature Learning. 2011

[37]

Hull J J . A database for handwritten text recognition research. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1994, 16( 5): 550–554

[38]

Hendrycks D, Dietterich T G. Benchmarking neural network robustness to common corruptions and perturbations. In: Proceedings of the 7th International Conference on Learning Representations. 2019

[39]

Recht B, Roelofs R, Schmidt L, Shankar V. Do CIFAR-10 classifiers generalize to CIFAR-10? 2018, arXiv preprint arXiv: 1806.00451

[40]

Hendrycks D, Gimpel K. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In: Proceedings of the 5th International Conference on Learning Representations. 2017

[41]

Jiang Y, Nagarajan V, Baek C, Kolter J Z. Assessing generalization of SGD via disagreement. In: Proceedings of the 10th International Conference on Learning Representations. 2022

RIGHTS & PERMISSIONS

Higher Education Press

PDF (5181KB)

Supplementary files

Highlights

237

Accesses

0

Citation

Detail

Sections
Recommended

/