Latent space climber: progressive latent space exploration for desirable sample discovery

Hongyang WANG , Zhangnan WANG , Jie LI

Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) : 2010714

PDF (4840KB)
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) :2010714 DOI: 10.1007/s11704-026-51649-0
Image and Graphics
RESEARCH ARTICLE
Latent space climber: progressive latent space exploration for desirable sample discovery
Author information +
History +
PDF (4840KB)

Abstract

Generative model (GM)-assisted product design has become increasingly popular. The key is to find generated samples (GSs) satisfying the design goal from the GM’s latent space (GLS). Most existing works rely on defining objective functions or writing prompts to search for GSs in the GLS. Unlike them, we propose a progressive approach that relies on multiple rounds of neighborhood exploration to choose desirable GSs from the GLS. Notably, the approach allows users to concretize and refine their goals during the exploration, thus applying to abstract or unspecific goals, which is unavailable for all existing techniques. The approach integrates two techniques to solve the challenges of achieving and applying it. First, many GSs are highly similar or irrelevant to the neighborhood center. Those GSs do not allow users to make comprehensive comparisons for rational choices and cannot be excluded by classic methods. Thus, we propose a method to avoid collecting them, which makes collected GSs have representative feature variations from the neighborhood center. Second, we need a system for applying the approach. The system should fulfill many visualization requirements to efficiently drive exploration and keep it always in the right direction. Thus, we followed the mountain-climbing metaphor to design the system and developed a series of visual and quantitative techniques to achieve these requirements. Cases on multiple real-world datasets and GMs, results of quantitative experiments, and performance and feedback of participants in user studies prove the approach’s effectiveness and usability.

Graphical abstract

Keywords

latent space / desirable sample / product design / generative model / GANs / visualization / interactive exploration

Cite this article

Download citation ▾
Hongyang WANG, Zhangnan WANG, Jie LI. Latent space climber: progressive latent space exploration for desirable sample discovery. Front. Comput. Sci., 2026, 20 (10) : 2010714 DOI:10.1007/s11704-026-51649-0

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Goodfellow I J, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y. Generative adversarial nets. In: Proceedings of the 28th International Conference on Neural Information Processing Systems. 2014, 2672−2680

[2]

Kingma D P, Welling M. Auto-encoding variational bayes. In: Proceedings of the 2nd International Conference on Learning Representations (ICLR). 2014, 1−14

[3]

Bengesi S, El-Sayed H, Sarker M K, Houkpati Y, Irungu J, Oladunni T . Advancements in generative AI: a comprehensive review of GANs, GPT, autoencoders, diffusion model, and transformers. IEEE Access, 2024, 12: 69812–69837

[4]

Nauata N, Chang K H, Cheng C Y, Mori G, Furukawa Y. House-GAN: relational generative adversarial networks for graph-constrained house layout generation. In: Proceedings of the 16th European Conference on Computer Vision. 2020, 162−177

[5]

Li J, Yang J, Hertzmann A, Zhang J, Xu T . LayoutGAN: synthesizing graphic layouts with vector-wireframe adversarial networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43( 7): 2388–2399

[6]

Abbasi M, Santos B P, Pereira T C, Sofia R, Monteiro N R C, Simes C J V, Brito R M M, Ribeiro B, Oliveira J L, Arrais J P . Designing optimized drug candidates with generative adversarial network. Journal of Cheminformatics, 2022, 14( 1): 40

[7]

Huang Q, Hong C, Wawrzynek J, Subedar M, Shao Y S. Learning a continuous and reconstructible latent space for hardware accelerator design. In: Proceedings of 2022 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). 2022, 277−287

[8]

Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. In: Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, 10674−10685

[9]

Fernandes P, Correia J, Machado P. Evolutionary latent space exploration of generative adversarial networks. In: Proceedings of the 23rd European Conference on Applications of Evolutionary Computation. 2020, 595−609

[10]

Machín B, Nesmachnow S, Toutouh J. Evolutionary latent space search for driving human portrait generation. In: Proceedings of 2021 IEEE Latin American Conference on Computational Intelligence (LA-CCI). 2021, 1−6

[11]

Kikuchi K, Simo-Serra E, Otani M, Yamaguchi K. Constrained graphic layout generation via latent optimization. In: Proceedings of the 29th ACM International Conference on Multimedia. 2021, 88−96

[12]

Nichol A Q, Dhariwal P, Ramesh A, Shyam P, Mishkin P, McGrew B, Sutskever I, Chen M. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In: Proceedings of the 39th International Conference on Machine Learning. 2022, 16784−16804

[13]

Saharia C, Chan W, Saxena S, Li L, Whang J, Denton E, Ghasemipour S K S, Ayan B K, Mahdavi S S, Gontijo-Lopes R, Salimans T, Ho J, Fleet D J, Norouzi M. Photorealistic text-to-image diffusion models with deep language understanding. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2643

[14]

Hao Y, Chi Z, Dong L, Wei F. Optimizing prompts for text-to-image generation. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 2923

[15]

Wu Z, Gao H, Wang Y, Zhang X, Wang S. Universal prompt optimizer for safe text-to-image generation. In: Proceedings of 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024, 6340−6354

[16]

Hertz A, Mokady R, Tenenbaum J, Aberman K, Pritch Y, Cohen-Or D. Prompt-to-prompt image editing with cross-attention control. In: Proceedings of the 11th International Conference on Learning Representations (ICLR). 2023

[17]

Zhou C, Zhong F, Öztireli C. CLIP-PAE: projection-augmentation embedding to extract relevant features for a disentangled, interpretable and controllable text-guided face manipulation. In: Proceedings of the ACM SIGGRAPH 2023 Conference Proceedings. 2023, 57

[18]

Chen C, Yuan J, Lu Y, Liu Y, Su H, Yuan S, Liu S . OoDAnalyzer: interactive analysis of out-of-distribution samples. IEEE Transactions on Visualization and Computer Graphics, 2021, 27( 7): 3335–3349

[19]

Bertucci D, Hamid M M, Anand Y, Ruangrotsakun A, Tabatabai D, Perez M, Kahng M . DendroMap: visual exploration of large-scale image datasets for machine learning with treemaps. IEEE Transactions on Visualization and Computer Graphics, 2023, 29( 1): 320–330

[20]

Fried O, DiVerdi S, Halber M, Sizikova E, Finkelstein A . IsoMatch: creating informative grid layouts. Computer Graphics Forum, 2015, 34( 2): 155–166

[21]

Cabezon Pedroso T, Byrne D. Browsing the latent space: a new approach to interactive design exploration for volumetric generative systems. In: Proceedings of the 15th Conference on Creativity and Cognition. 2023, 330−333

[22]

Zhang Y, Li J, Xu C. Graph-based latent space traversal for new molecules discovery. In: Proceedings of the 16th International Symposium on Visual Information Communication and Interaction. 2023, 26

[23]

Härkönen E, Hertzmann A, Lehtinen J, Paris S. GANSpace: discovering interpretable GAN controls. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 825

[24]

Jeong S, Li M, Berger M, Liu S. Concept lens: visually analyzing the consistency of semantic manipulation in GANs. In: Proceedings of 2023 IEEE Visualization and Visual Analytics (VIS). 2023, 221−225

[25]

Jeong S, Liu S, Berger M . Interactively assessing disentanglement in GANs. Computer Graphics Forum, 2022, 41( 3): 85–95

[26]

Kwon O H, Ma K L . A deep generative model for graph layout. IEEE Transactions on Visualization and Computer Graphics, 2020, 26( 1): 665–675

[27]

Kwon O H, Kao C H, Chen C H, Ma K L . A deep generative model for reordering adjacency matrices. IEEE Transactions on Visualization and Computer Graphics, 2023, 29( 7): 3195–3208

[28]

Dang H, Buschek D. GestureMap: supporting visual analytics and quantitative analysis of motion elicitation data by learning 2D embeddings. In: Proceedings of 2021 CHI Conference on Human Factors in Computing Systems. 2021, 317

[29]

Feng Y, Wang X, Wong K, Wang S, Lu Y, Zhu M, Wang B, Chen W . PromptMagician: interactive prompt engineering for text-to-image creation. IEEE Transactions on Visualization and Computer Graphics, 2024, 30( 1): 295–305

[30]

Brade S, Wang B, Sousa M, Oore S, Grossman T. Promptify: text-to-image generation through interactive prompt exploration with large language models. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 2023, 96

[31]

Chen C, Lv F, Guan Y, Wang P, Yu S, Zhang Y, Tang Z . Human-guided image generation for expanding small-scale training image datasets. IEEE Transactions on Visualization and Computer Graphics, 2025, 31( 6): 3809–3821

[32]

Padala M, Das D, Gujar S. Effect of input noise dimension in GANs. In: Proceedings of the 28th International Conference on Neural Information Processing. 2021, 558−569

[33]

Feng R, Zhao D, Zha Z J. Understanding noise injection in GANs. In: Proceedings of the 38th International Conference on Machine Learning. 2021, 3284−3293

[34]

Choi J, Hwang G, Cho H, Kang M . Analyzing the latent space of GAN through local dimension estimation for disentanglement evaluation. Pattern Recognition, 2025, 157: 110914

[35]

Barthel K U, Hezel N, Jung K, Schall K . Improved evaluation and generation of grid layouts using distance preservation quality and linear assignment sorting. Computer Graphics Forum, 2023, 42( 1): 261–276

[36]

Radford A, Kim J W, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I. Learning transferable visual models from natural language supervision. In: Proceedings of the 38th International Conference on Machine Learning. 2021, 8748−8763

[37]

Patil A G, Li M, Fisher M, Savva M, Zhang H. LayoutGMN: neural graph matching for structural layout similarity. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 11043−11052

[38]

Wang Z, Bovik A C, Sheikh H R, Simoncelli E P . Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 2004, 13( 4): 600–612

[39]

Karras T, Aittala M, Hellsten J, Laine S, Lehtinen J, Aila T. Training generative adversarial networks with limited data. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 1015

[40]

Karras T, Laine S, Aila T. A style-based generator architecture for generative adversarial networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019, 4396−4405

[41]

Heyrani Nobari A, Rashad M F, Ahmed F. CreativeGAN: editing generative adversarial networks for creative design synthesis. In: Proceedings of the International Design Engineering Technical Conferences and Computers and Information in Engineering Conference. 2021, V03AT03A002

[42]

Kiyota Y. Promoting open innovations in real estate tech: provision of the LIFULL HOME’S data set and collaborative studies. In: Proceedings of 2018 ACM on International Conference on Multimedia Retrieval (ICMR). 2018, 6

[43]

Deng L . The MNIST database of handwritten digit images for machine learning research [Best of the Web]. IEEE Signal Processing Magazine, 2012, 29( 6): 141–142

[44]

Krizhevsky A. Learning multiple layers of features from tiny images. Toronto: University of Toronto, 2009

[45]

Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 6629−6640

[46]

Zhang R, Isola P, Efros A A, Shechtman E, Wang O. The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2018, 586−595

[47]

Brock A, Donahue J, Simonyan K. Large scale GAN training for high fidelity natural image synthesis. In: Proceedings of the 7th International Conference on Learning Representations. 2019

[48]

Deng J, Dong W, Socher R, Li L J, Li K, Fei-Fei L. ImageNet: a large-scale hierarchical image database. In: Proceedings of 2009 IEEE Conference on Computer Vision and Pattern Recognition. 2009, 248−255

[49]

Miyato T, Kataoka T, Koyama M, Yoshida Y. Spectral normalization for generative adversarial networks. In: Proceedings of the 6th International Conference on Learning Representations (ICLR). 2018

[50]

Li Z, Wang C, Zheng H, Zhang J, Li B. FakeCLR: exploring contrastive learning for solving latent discontinuity in data-efficient GANs. In: Proceedings of the 17th European Conference on Computer Vision. 2022, 598−615

[51]

Zhang Y, Li J, Zeng W . Latent space map for visual utilization of generated data. IEEE Transactions on Visualization and Computer Graphics, 2025, 31( 12): 10746–10761

[52]

Song Y, Dhariwal P, Chen M, Sutskever I. Consistency models. In: Proceedings of the 40th International Conference on Machine Learning. 2023, 32211−32252

Rights & permissions

Higher Education Press

PDF (4840KB)

Supplementary files

Highlights

418

Accesses

0

Citation

Detail

Sections
Recommended

/