About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
Latent space climber: progressive latent space exploration for desirable sample discovery
Hongyang WANG , Zhangnan WANG , Jie LI
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) : 2010714
Generative model (GM)-assisted product design has become increasingly popular. The key is to find generated samples (GSs) satisfying the design goal from the GM’s latent space (GLS). Most existing works rely on defining objective functions or writing prompts to search for GSs in the GLS. Unlike them, we propose a progressive approach that relies on multiple rounds of neighborhood exploration to choose desirable GSs from the GLS. Notably, the approach allows users to concretize and refine their goals during the exploration, thus applying to abstract or unspecific goals, which is unavailable for all existing techniques. The approach integrates two techniques to solve the challenges of achieving and applying it. First, many GSs are highly similar or irrelevant to the neighborhood center. Those GSs do not allow users to make comprehensive comparisons for rational choices and cannot be excluded by classic methods. Thus, we propose a method to avoid collecting them, which makes collected GSs have representative feature variations from the neighborhood center. Second, we need a system for applying the approach. The system should fulfill many visualization requirements to efficiently drive exploration and keep it always in the right direction. Thus, we followed the mountain-climbing metaphor to design the system and developed a series of visual and quantitative techniques to achieve these requirements. Cases on multiple real-world datasets and GMs, results of quantitative experiments, and performance and feedback of participants in user studies prove the approach’s effectiveness and usability.
latent space / desirable sample / product design / generative model / GANs / visualization / interactive exploration
| [1] |
Goodfellow I J, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y. Generative adversarial nets. In: Proceedings of the 28th International Conference on Neural Information Processing Systems. 2014, 2672−2680 |
| [2] |
Kingma D P, Welling M. Auto-encoding variational bayes. In: Proceedings of the 2nd International Conference on Learning Representations (ICLR). 2014, 1−14 |
| [3] |
|
| [4] |
Nauata N, Chang K H, Cheng C Y, Mori G, Furukawa Y. House-GAN: relational generative adversarial networks for graph-constrained house layout generation. In: Proceedings of the 16th European Conference on Computer Vision. 2020, 162−177 |
| [5] |
|
| [6] |
|
| [7] |
Huang Q, Hong C, Wawrzynek J, Subedar M, Shao Y S. Learning a continuous and reconstructible latent space for hardware accelerator design. In: Proceedings of 2022 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). 2022, 277−287 |
| [8] |
Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. In: Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, 10674−10685 |
| [9] |
Fernandes P, Correia J, Machado P. Evolutionary latent space exploration of generative adversarial networks. In: Proceedings of the 23rd European Conference on Applications of Evolutionary Computation. 2020, 595−609 |
| [10] |
Machín B, Nesmachnow S, Toutouh J. Evolutionary latent space search for driving human portrait generation. In: Proceedings of 2021 IEEE Latin American Conference on Computational Intelligence (LA-CCI). 2021, 1−6 |
| [11] |
Kikuchi K, Simo-Serra E, Otani M, Yamaguchi K. Constrained graphic layout generation via latent optimization. In: Proceedings of the 29th ACM International Conference on Multimedia. 2021, 88−96 |
| [12] |
Nichol A Q, Dhariwal P, Ramesh A, Shyam P, Mishkin P, McGrew B, Sutskever I, Chen M. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In: Proceedings of the 39th International Conference on Machine Learning. 2022, 16784−16804 |
| [13] |
Saharia C, Chan W, Saxena S, Li L, Whang J, Denton E, Ghasemipour S K S, Ayan B K, Mahdavi S S, Gontijo-Lopes R, Salimans T, Ho J, Fleet D J, Norouzi M. Photorealistic text-to-image diffusion models with deep language understanding. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2643 |
| [14] |
Hao Y, Chi Z, Dong L, Wei F. Optimizing prompts for text-to-image generation. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 2923 |
| [15] |
Wu Z, Gao H, Wang Y, Zhang X, Wang S. Universal prompt optimizer for safe text-to-image generation. In: Proceedings of 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024, 6340−6354 |
| [16] |
Hertz A, Mokady R, Tenenbaum J, Aberman K, Pritch Y, Cohen-Or D. Prompt-to-prompt image editing with cross-attention control. In: Proceedings of the 11th International Conference on Learning Representations (ICLR). 2023 |
| [17] |
Zhou C, Zhong F, Öztireli C. CLIP-PAE: projection-augmentation embedding to extract relevant features for a disentangled, interpretable and controllable text-guided face manipulation. In: Proceedings of the ACM SIGGRAPH 2023 Conference Proceedings. 2023, 57 |
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
Cabezon Pedroso T, Byrne D. Browsing the latent space: a new approach to interactive design exploration for volumetric generative systems. In: Proceedings of the 15th Conference on Creativity and Cognition. 2023, 330−333 |
| [22] |
Zhang Y, Li J, Xu C. Graph-based latent space traversal for new molecules discovery. In: Proceedings of the 16th International Symposium on Visual Information Communication and Interaction. 2023, 26 |
| [23] |
Härkönen E, Hertzmann A, Lehtinen J, Paris S. GANSpace: discovering interpretable GAN controls. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 825 |
| [24] |
Jeong S, Li M, Berger M, Liu S. Concept lens: visually analyzing the consistency of semantic manipulation in GANs. In: Proceedings of 2023 IEEE Visualization and Visual Analytics (VIS). 2023, 221−225 |
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
Dang H, Buschek D. GestureMap: supporting visual analytics and quantitative analysis of motion elicitation data by learning 2D embeddings. In: Proceedings of 2021 CHI Conference on Human Factors in Computing Systems. 2021, 317 |
| [29] |
|
| [30] |
Brade S, Wang B, Sousa M, Oore S, Grossman T. Promptify: text-to-image generation through interactive prompt exploration with large language models. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 2023, 96 |
| [31] |
|
| [32] |
Padala M, Das D, Gujar S. Effect of input noise dimension in GANs. In: Proceedings of the 28th International Conference on Neural Information Processing. 2021, 558−569 |
| [33] |
Feng R, Zhao D, Zha Z J. Understanding noise injection in GANs. In: Proceedings of the 38th International Conference on Machine Learning. 2021, 3284−3293 |
| [34] |
|
| [35] |
|
| [36] |
Radford A, Kim J W, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I. Learning transferable visual models from natural language supervision. In: Proceedings of the 38th International Conference on Machine Learning. 2021, 8748−8763 |
| [37] |
Patil A G, Li M, Fisher M, Savva M, Zhang H. LayoutGMN: neural graph matching for structural layout similarity. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 11043−11052 |
| [38] |
|
| [39] |
Karras T, Aittala M, Hellsten J, Laine S, Lehtinen J, Aila T. Training generative adversarial networks with limited data. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 1015 |
| [40] |
Karras T, Laine S, Aila T. A style-based generator architecture for generative adversarial networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019, 4396−4405 |
| [41] |
Heyrani Nobari A, Rashad M F, Ahmed F. CreativeGAN: editing generative adversarial networks for creative design synthesis. In: Proceedings of the International Design Engineering Technical Conferences and Computers and Information in Engineering Conference. 2021, V03AT03A002 |
| [42] |
Kiyota Y. Promoting open innovations in real estate tech: provision of the LIFULL HOME’S data set and collaborative studies. In: Proceedings of 2018 ACM on International Conference on Multimedia Retrieval (ICMR). 2018, 6 |
| [43] |
|
| [44] |
Krizhevsky A. Learning multiple layers of features from tiny images. Toronto: University of Toronto, 2009 |
| [45] |
Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 6629−6640 |
| [46] |
Zhang R, Isola P, Efros A A, Shechtman E, Wang O. The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2018, 586−595 |
| [47] |
Brock A, Donahue J, Simonyan K. Large scale GAN training for high fidelity natural image synthesis. In: Proceedings of the 7th International Conference on Learning Representations. 2019 |
| [48] |
Deng J, Dong W, Socher R, Li L J, Li K, Fei-Fei L. ImageNet: a large-scale hierarchical image database. In: Proceedings of 2009 IEEE Conference on Computer Vision and Pattern Recognition. 2009, 248−255 |
| [49] |
Miyato T, Kataoka T, Koyama M, Yoshida Y. Spectral normalization for generative adversarial networks. In: Proceedings of the 6th International Conference on Learning Representations (ICLR). 2018 |
| [50] |
Li Z, Wang C, Zheng H, Zhang J, Li B. FakeCLR: exploring contrastive learning for solving latent discontinuity in data-efficient GANs. In: Proceedings of the 17th European Conference on Computer Vision. 2022, 598−615 |
| [51] |
|
| [52] |
Song Y, Dhariwal P, Chen M, Sutskever I. Consistency models. In: Proceedings of the 40th International Conference on Machine Learning. 2023, 32211−32252 |
Higher Education Press
/
| 〈 |
|
〉 |