1 Introduction
Globalization has long been proven to have a significant impact on the development of tourism. Once a built or natural environment becomes a destination, globalization inevitably marks the endpoint of its evolutionary process (
Xu and Lu, 2024). Tourism is one of the largest markets, accounting for more than 10% of the world’s GDP. The number of trips made each year worldwide exceeds the total global population (
Hall, 2015;
Rasoolimanesh et al., 2023;
Uǧur and Akbıyık, 2020). Although the impact of COVID-19 in 2020 caused a decline of over 70% in the tourism industry, by the first quarter of 2023, the industry had recovered to 80% of the pre-pandemic levels for the same period (
UNWTO, 2020,
2023). It underscores the immense scale of the potential of this industry. Globalization, in turn, transforms human environments into destinations through the lens of consumer-driven postmodernism, drawing inspiration from global and historical treasures to create visual spectacles that have secured a prominent place in the experience-driven economy (
Ibelings, 2002;
Urry and Larsen, 2011). Cities like Bilbao, started with the Guggenheim Museum, have undergone a comprehensive transformation of their urban image through the regeneration of the Abandoibarra area (
Sepe and Di Trapani, 2010). Towns such as Zhouzhuang have created a strong image by focusing on the theme of “Small Bridge, Flowing Water, and Residents” (
Gao et al., 2016). Architectures such as CCTV’s Headquarters, initially designed to be “a convincing architecture of image” (
Rodeš, 2021). These images are leveraged as tools to enhance the competitiveness of local spaces in the dynamic global tourism market (
Afshardoost and Eshaghi, 2020).
With the explosive growth of online services in recent decades, the impact of digitalization has become increasingly diverse. Social networks play a central role in information dissemination, public opinion formation and polarization, among other areas (
Pagan et al., 2021). The combination of globalization and digitalization has resulted in a continuous transformation of “Spaces of Places” into “Spaces of Flows”, intensifying the battle over the construction of meaning among destinations (
Castells, 2009). Traditional built environments possess competitive cultural resources such as history and tradition, abundant natural resources like rivers, mountains, and climate, as well as digital resources from user-generated content on social media and film promotions (
Baral et al., 2017;
Kubickova and Martin, 2020;
Ritchie and Crouch, 2003). These attributes make them uniquely suited for destination development, but also prone to face more challenges, including the degradation of natural resources and the loss of cultural heritage, which could diminish their perceived value in the process of development (
Zhang et al., 2021).
The water towns of the Jiangnan, China represent a typical example of traditional built environments influenced by tourism. The picturesque landscapes of “Small Bridge, Flowing Water, and Residents”, the leisure lifestyle, and the rich cultural heritage, coupled with their proximity to major metropolitan areas like Shanghai (
Liu, 2002;
Porfyriou, 2019), made them popular global destinations after the tourism industry began to grow in the 1980s. Issues such as creative destruction, commercialization, and homogenization, emerged simultaneously and became significant obstacles to the development of these ancient towns (
Bei Huang et al., 2007;
Zhang et al., 2021). The profitability of the entire water town tourism market has reached a state of equilibrium (
Ma et al., 2015). Due to competition, some water towns, such as Zhouzhuang and Wuzhen, have solidified their leading positions, while others, despite efforts to differentiate themselves through local cultural heritage, have struggled to surpass these top destinations. Homogenization has thus become a prevalent characteristic among these towns. In the era of globalization and digitalization, visual experiences have taken center stage, and the collective experiences of visitors cannot be overlooked. In response to the homogenization of water towns, this study reconsiders the issue from a tourist perspective. By analyzing user-uploaded photos from two water towns in Shanghai, we aim to reassess the landscape perception of Jiangnan water towns and reevaluate the challenges they face.
1.1 Analyzing landscape perception by utilizing visual content
Perception is both a response to external stimuli in terms of sensation and the active, deliberate imprinting of specific phenomena in the mind after disregarding or excluding other occurrences. The objects of perception are valuable to humans, serving either as necessities for survival or sources of cultural satisfaction. Over time, perceptions of the environment accumulate to form more stable attitudes (
Tuan, 1990). Landscape perception has been widely applied in research on human-environment relationships, land use and management (
Dorning et al., 2017), visual quality assessment (
Ma et al., 2021;
Zhao et al., 2024), ecosystem services (
Enrica et al., 2023), cross-cultural comparisons (
Matijošaitienė et al., 2014), and walkability (
Tiitu et al., 2024) as examples.
Web 2.0 has transformed users from passive consumers into active participants. A vast amount of photos, text, and geographic location data is being shared, creating an unprecedented volume of user-generated content, which has opened up new avenues for research in landscape perception (
Ghermandi et al., 2023;
Kaplan and Haenlein, 2010). Visual content, such as photographs, provides rich information that reflects users’ perceptions of their surroundings and their appreciation of specific spatial characteristics (
Oteros-Rozas et al., 2018), which becomes vital recourses of the research of landscape perception.
Humans are visually dominant creatures, primarily perceiving the world through vision. The visual-first characteristic actually originates from vision’s historical struggle to differentiate itself from other senses, eventually becoming the predominant mode of perception in modern society (
Urry and Larsen, 2011). The significant shift in perception during the 19th century was most evident in the gradual rise of vision as the predominant mode of sensory experience. However, vision did not replace other senses, rather, it became the dominant organizer of sensory experience within the context of emerging technologies and societal changes (
Crary, 1992). “The mode of human sense perception changes with humanity’s entire mode of existence,” photography has fundamentally transformed our perception of space in modern life (
Benjamin, 2018;
Nnany, 2015). Photography offers a means for people to convey experiences and ideas through images that possess unique characteristics, requiring minimal effort and cultural background to comprehend (
Kislinger and Kotrschal, 2021). From Kodak to smartphones, photography has become increasingly accessible to the general public, and the rise of the internet and social media has made photo sharing a global phenomenon. The pervasive presence of photographs online has profoundly transformed our perceptual experiences, making them more visually dominated than they were during the early modern era.
The essence of a photograph resides in its content. Therefore, the primary task of photographic analysis is to identify and categorize this content by extracting key information from visual data and assigning appropriate labels. This approach aligns with the fundamental principles of content analysis (
Li and Yang, 2022;
Neuendorf, 2017;
Walden-Schreiner et al., 2018). Owing to its adaptability, this framework has gained increasing adoption in visual content research (
Sheydayi and Dadashpoor, 2023). Machine learning streamlines large-scale photographic analysis by automating content processing (
Athey, 2017). Computer vision creates cognition and understanding of scenes by observing visual objects and generating images based on specific tasks. For example, researchers used a convolutional neural network model to perform image segmentation on street view photos of Tianjin, calculating the proportion of sky and vegetation (
Wang et al., 2022).
Arefieva et al. (2021) utilized Google Cloud Vision to segment images collected from Instagram and applied K-means clustering and Word2Vec to classify the associated keywords, thereby uncovering perceptions of urban imagery in Austria. Recently, the introduction of the Transformer architecture has enabled parallel training, significantly enhancing the accuracy and efficiency of data processing for both individuals and institutions. Today, the use of pre-trained Transformer models has nearly become a standard practice (
Pelicon et al., 2020).
Luo et al. (2022) employed a Transformer-based model for semantic segmentation of panoramic river views captured by drones. While perceptual studies analyzing photo content have developed mature methods and technologies, current photo classification primarily focuses on the elements within the image. This approach implies that environmental perception is largely based on the intuitive information reflected in the photograph. For example, semantic segmentation in a photograph captures two-dimensional information, such as buildings, water bodies, and plants, but it fails to convey three-dimensional aspects like spatial structure. However, when people are immersed in a specific environment, the overall spatial structure and organizational patterns of objects become crucial factors (
Hunter and Askarinejad, 2015). Further exploration is needed from the perspective of spatial structure to understand human perception of the environment.
1.2 Spatial perception
The binocular vision provided by the human eyes allows for the overlapping of two focal points, helping to establish a clear three-dimensional space. This innate ability enables humans, even as newborns, to use perspective and parallax to recognize faces. Eight-weeks old infants can distinguish depth and the horizon line. It is through the depth cues that humans can judge the organization of space. As a result, we have become accustomed to observing the world in a three-dimensional and depth-perceiving manner (
Bower, 1966;
Torralba and Oliva, 2002;
Tuan, 1979). Visual content research predominantly emphasizes functional aspects, such as spatial attributes or dominant elements, to classify spaces, while largely neglecting their inherent abstract structures and organizational patterns.
A photograph, though a two-dimensional plane, still contains numerous spatial cues that convey depth (
Zakia and Suler, 2017). Depth has been proven to have a significant effect on landscape perception (
Appleton, 1996;
Zhang et al., 2021). From an evolutionary perspective, environments with extensive spatial depth are perceived as advantageous because they enhance the ability to monitor resources and detect threats (
Appleton, 1996). Depth information primarily helps humans perceive the proportional relationships and sizes of objects within the environment (
Hunter and Askarinejad, 2015). Further understanding of spatial features is possible when depth can be observed in a two-dimensional photograph. As the foreground subjects recede into the background, this shift signifies an increase in depth and a transition in focus from capturing the subject alone to emphasizing the overall environment (
Mattens, 2011).
Additionally, horizontal and vertical constitute an important pair of concepts in spatial perception that have fundamental significance in how we observe the environment. The vertical direction represents the sacred and spiritual aspects, while the horizontal direction emphasizes the pursuit of the secular and material (
Harries, 1998).
Peckham (1967) argues that spaces with shallow depth, which are enclosed and have little variation in layers, tend to evoke a fixed and restrained feeling. In contrast, spaces with greater depth, which are open and rich in variation, evoke a dynamic and expansive sensation. Pantheon is considered a paradigm of vertical space, creating an axis of the heavens through the singular oculus at its top. In Villa Rotonda, Palladio employs a “lantern” to seal the oculus of the dome, thereby transforming the room into a frame that captures the expansive landscape beyond. The Verticality is supplanted by a horizontal “house of views” (
Ruan, 2015a). In 17th-century Dutch painting, vertical space is secularized. Despite the alluring views outside the windows, the figures in the scenes remain absorbed in their indoor activities (
Ruan, 2015b). Large floor-to-ceiling windows in modern homes, designed to offer views of the captivating scenery outside, turn to provide a horizontal, open experience (
Tuan, 1977). Verticality and horizontality respectively correspond to an individual’s focused attention on a singular object versus environmental distraction through spatial perception. The shallower the depth, the more vertical the space tends to be, conversely, the greater the depth, the more horizontal the space becomes.
In the field of landscape perception research, studies increasingly rely on images as primary data, often employing semantic segmentation techniques. This 2D-focused approach has limitations, as it fails to capture the spatial structures that fundamentally embody human activities. When examining historic environments today, semantic segmentation alone may identify surface-level material similarities. Yet, by exploring spatial structures, we can uncover whether the deeper functional, cultural, and social meanings still resonate with contemporary observers. Depth estimation now enables the extraction of depth information from two-dimensional photographs. Previous studies have utilized depth estimation to analyze street view image (SVI) data, achieving significant progress in the development of landscape perception frameworks (
Cao et al., 2025). However, SVI usually collected by a vehicle equipped with multiple cameras and sensors, assigned by a specific institution. This method provides a valuable large-scale source of urban data, enabling the examination of visual features from a human perspective (
Biljecki and Ito, 2021). Nevertheless, SVI are objectively collected and do not represent an individual’s subjective experience of the environment, maintaining a certain distance from user-generated content.
This study analyzes the spatial perception and element characteristics of two water towns in Shanghai, Zhujiajiao and Fengjing, using depth estimation and semantic segmentation to analyze photos uploaded on tourism platforms. It reconsiders the landscape perception of traditional human settlements from the perspective of spatial structure.
2 Methodology
2.1 Study area
This study selects two historic water towns in Shanghai as case study areas: Zhujiajiao and Fengjing (Fig. 1). They are relatively representative among the ancient towns in the suburban areas of Shanghai. In terms of location, both Zhujiajiao and Fengjing are situated far from the city center, near the borders of neighboring provinces. Zhujiajiao is one of the oldest historical and cultural towns in Shanghai, with a long history of tourism development and a mature achievement. Fengjing, on the other hand, is the first “Chinese Historical and Cultural Town” in Shanghai, with a shallower level development, but relatively well-preserved. The water towns of Jiangnan are celebrated for their picturesque scenes of “Small Bridges, Flowing Water, and Residents,” which often give them a similar appearance. However, younger generations are increasingly migrating away, leaving behind an aging population. Consequently, traditions such as festivals and customs are struggling to survive without the support of a vibrant community or the transmission to new generations. As a result, the unique cultural distinctions that once defined these towns are fading, leaving only their physical appearances intact. This homogenization is particularly problematic in an era where tourism is dominated by visual appeal.
Nestled in western Qingpu District with a millennium-long history, Zhujiajiao straddles the tri-provincial junction of Shanghai, Jiangsu, and Zhejiang near Dianshan Lake. Renowned for its intricate canal network and historical role as a commercial hub, this “Land of Rice and Fish” combines aquatic abundance with well-preserved Ming-Qing architecture. Its core area retains original street layouts and heritage sites, epitomizing Jiangnan’s water town legacy.
Fengjing situated in southwestern Shanghai’s Jinshan District at the Shanghai-Zhejiang border. It embodies the cultural fusion of Wu and Yue traditions, earning its title as the “Wu-Yue Cultural Confluence”. Designated among China’s first National Historical-Cultural Villages in 2005, it preserves classic Jiangnan landscapes through its labyrinthine waterways and arched bridges, offering a living archive of regional architectural heritage.
2.2 Landscape perception mining and analysis based on photo data
The study obtains landscape perception data from tourist photos through three steps (Fig. 2). The first step is data collection. The second step involves semantic segmentation to identify landscape elements and classify them. Subsequently, depth estimation and calculation of the average absolute depth are used for the spatial classification of photos. After obtaining the perception of elements and space, statistical analysis and correlation analysis are performed.
2.2.1 Data Collection
Ctrip and Mafengwo are two prominent tourism platforms in China, each with a substantial user base. These platforms serve both functional and social purposes, facilitating user aggregation for various objectives. The social aspect is achieved by creating a virtual community where users can upload opinions and comments, effectively functioning as a form of social media (
Huai et al., 2023;
Zhu et al., 2022). This study explores the landscape perception of Jiangnan historic water towns under the influence of tourism, making these two platforms particularly relevant for our research. We collected 1987 photos of Zhujiajiao and 1994 photos of Fengjing from Ctrip, spanning from 2011 to 2022. From Mafengwo, 970 photos of Zhujiajiao and 1282 photos of Fengjing are gathered, with a time span from 2015 to 2022. A total of 6233 photos were collected from the two platforms, with 2957 of Zhujiajiao and 3276 of Fengjing.
2.2.2 Semantic segmentation and element categorizing
Element recognition in the photos was achieved through semantic segmentation, a process that involves dividing an image into multiple segments and performing pixel-level labeling to identify elements such as trees, water, and sky (
De Berg, 2000;
Minaee et al., 2022). Several established semantic segmentation architectures have been developed, including FCN, SegNet, PSPNet, and SegFormer. Among these, SegFormer demonstrates superior performance due to its hierarchical Transformer encoder that captures multi-scale contextual features through overlapped patch merging, and an efficient lightweight MLP decoder that eliminates the need for computationally expensive operations. This architectural innovation enables SegFormer-b5 to achieve state-of-the-art performance of 84.0% mIoU on the Cityscapes validation set, outperforming FCN (76.6%), SegNet (57.0%), and PSPNet (78.5%) under comparable training conditions (
Badrinarayanan et al., 2017;
Luo et al., 2022;
Xie et al., 2021;
Zhao et al., 2017). Compared to the more recent Mask2Former, although it outperforms SegFormer in terms of overall performance, it requires higher image resolution. In contrast, SegFormer exhibits less sensitivity to low-resolution images, making it more suitable for UGC studies (
Gibril et al., 2024;
Xie et al., 2021). We selected the “nvidia/segformer-b5-finetuned-ade-640-640” model, which has been fine-tuned on the ADE20k dataset, due to its exceptional performance in urban scene parsing, where structural precision is crucial. A total of 136 landscape elements were identified in all photos of two towns, which were subsequently classified. After labeling the landscape elements, the Landmark Detection API provided by Baidu AI was used to identify potential landmarks in the photos (Fig. 3).
The labels returned by semantic segmentation were manually classified into 16 categories, as shown in Table 1, which explains each of the 16 landscape elements. Walls were classified under the “Infrastructure” category, primarily because the segmentation results indicated that walls appeared not only in photos containing buildings but also in photos of bridges, embankments, and other structures. Additionally, some wall surfaces were directly captured by tourists. Photos featuring food, plates, or other serving items often appear alongside the food, but the primary focus of these photos is the food itself. Therefore, the serving items were classified under “Foods” rather than “Artifacts”.
2.2.3 Depth estimation and space categorizing
Two techniques, depth estimation and structural similarity algorithms, are utilized to access the spatial classification of photographs (Fig. 4).
The enhanced ability of understanding the 3-dimensional space in reality makes depth estimation models a crucial role in many applications such as augmented reality and autonomous driving (
Ming et al., 2021;
Valentin et al., 2018). Depth is typically obtained using laser, structured light, or other surface reflections which incur high costs. As a result, monocular depth estimation has become the more widely used and mainstream approach for depth estimation (
Bhoi, 2019;
Mertan et al., 2022;
Zhang et al., 2020). Traditional monocular depth estimation methods are often based on neural network architectures (
Goodfellow et al., 2014;
Krizhevsky et al., 2012;
Rumelhart et al., 1986). Transformer based models, such as Intel/dpt-large, has shown a 28% improvement in performance compared to the most advanced fully convolutional networks (
Ranftl et al., 2021). In this study, all photos were converted into depth maps and stored as arrays by using dpt-large. The first photo was sequentially matched with the remaining photos, which involved calculating structural similarity:
x and y represent two images, μ is the pixel average value, and are the pixel variances of x and y, respectively, σxy is the covariance between x and y, C1 and C2 are constants used for stabilizing the calculation. Structural similarity index was set to 0.85. A range of values between 0.75 and 0.95, with increments of 0.05, was tested, and 0.85 yielded the best balance of computational performance and calculation effectiveness.
Photos with a similarity score above the threshold are marked as similar and classified into the same folder, and they no longer participate in subsequent calculations. This process continues until the last photo is processed. After completing the calculations, all the photos are grouped into 256 groups for Zhujiajiao and 189 groups for Fengjing. The groups containing similar photos are then manually labeled again.
Based on the earlier discussion of spatial perception, the classification criterion is as follows: horizontal space refers to the observation of the environment, while vertical space focuses on the attention directed at a specific element (Fig. 5).
The initial classification, which integrates the concepts of horizontal and vertical space with the actual content of the photos, divides the space into four categories. Vertical: The classification of vertical-type images reveals an intriguing pattern, which can be divided into two distinct spatial categories: (1) Vertical Space, where tourists photograph objects head-on, creating a 90-degree alignment between themselves and the subject. This suggests a focused, almost immersive connection with the object; and (2) Oblique Space, where photos often frame the subject at an angle, producing an angular perspective. This mimics the dynamic perception of objects while in motion, indicating that the photographer was likely moving. Alternatively, the oblique angle may result from practical constraints―either the subject was too large to capture frontally, or the photographer lacked sufficient space to step back for a full frontal shot. Horizontal: (1) a clearly directed space delineated by artificial or natural elements on both sides, where the human gaze naturally extends along the axis rather than fixating on any single object, facilitating movement in that direction; (2) Panoramic Space, in a panoramic view, the lens captures the entire expansive scene, providing a comprehensive overview of vast environments.
The absolute mean depth refers to the ratio of the total depth value in the depth map to the total number of pixels, reflecting the spatial structure of the photo. For example, a landscape photo with a range of several hundred meters, a city street scene with a range of tens of meters, and a close-up shot all have different spatial structures, and their absolute mean depth values will differ significantly (
Lee et al., 2001;
Torralba and Oliva, 2002). Except for panoramic space, the other three types of photos can have varying depths, so the absolute mean depth was calculated for them. By setting different depth value ranges they were further classified into multiple subcategories (Table 2). Together with panoramic space, this results in a total of 10 subcategories of space.
Tourist photos in this study were ultimately classified into four categories, including ten subcategories: 1) Corridor space, including Corridor Space_Wide, Corridor Space_Regular, and Corridor Space_Narrow; 2) Panoramic space; 3) Vertical space, containing Vertical Space_Distant, Vertical Space_Middle, Vertical Space_Foreground, and Vertical Space_Close up; 4) Oblique space, in which contains Oblique Space_Distant and Oblique Space_Middle. The first two belong to horizontal space, meaning that when tourists are photographing, they focus more on the space itself, guiding their line of sight and movement direction. The latter two belong to vertical space, where tourists focus more on the objects, which attracts them to pause and observe. Table 3 explains the categories contained in the four types of space, and Fig. 6 shows representative photos of each space.
2.3 Analysis of elemental and spatial perception
After the perception extraction and classification described above, this study conducted descriptive statistics on the perception of elements and space, followed by two types of analysis. The first is a comparison of the perceptions of elements and space between Zhujiajiao and Fengjing. Given our relatively large sample size and the proportional nature of the data meeting the conditions
np ≥ 10 and
n(1 –
p) ≥ 10, the two-sample
z-test was selected as the testing method. The second is a correlation analysis, which provides a feasible approach to better understand the relationship between elements and space in different towns. The correlation analysis was divided into two steps: the first is the correlation between elements, and the second is the correlation between elements and space. The goal is to explore how elements are perceived in water towns and whether the perception of elements is related to spatial structure. For the analysis between elements, we used the Dice Coefficient as the indicator of correlation. By giving higher weight to higher frequencies, this method can more effectively capture the co-occurrence frequency and closeness of landscape elements within the same photo. The calculation method is as follows (
Dice, 1945):
Dice Coefficient calculates the correlation between elements by estimating the co-occurrence frequency and individual frequency. In this formula, x and y represent the sets of elements extracted from different photos, |x ꓵy| indicates the size of the intersection between x and y, that is, the number of common elements between the two photos. |x| and |y| represent the sizes of the two sets, i. e., the number of elements in each of the two photos.
For the correlation analysis between elements and space, Pointwise Mutual Information was used as the metric. This method is better suited to highlight those landscape elements that are significant in specific spaces, to uncover the potential relationships between elements and space (
Church and Hanks, 1990):
P(x, y) represents the co-occurrence probability of element x and space y, P(x) and P(y) are the probabilities of element x and space y occurring independently. A higher PMI value indicates that the co-occurrence of x and y occurs more frequently than expected based on their independent probabilities, meaning there is a strong positive correlation between the two. When PMI equals 0, it indicates that x and y are independent and there is no particular relationship. A negative PMI value suggests that the co-occurrence is less frequent than expected, potentially indicating a negative correlation. To uncover the underlying patterns in the data, Principal Component Analysis (PCA) is employed.
3 Results
3.1 Descriptive statistics
3.1.1 Perception of elements
The landscape element perception statistics based on photo data for the ancient towns are shown in Fig. 7. In Zhujiajiao, the proportion of artificial elements is 64.39%, and natural elements account for 35.61%. Among the artificial elements, Architecture and Infrastructure have the highest proportions at 12.90% and 12.06%, followed by Transportation (6.40%), Artifacts (5.96%), Structure (5.68%), and Information (5.65%). The proportions of Foods, Geomaterials, and Landmark are relatively low at 4.94%, 3.66%, and 3.14%, respectively. The proportion of Furniture, as well as Art and culture, are the lowest at 2.86% and 1.13%. Among the natural elements, Vegetation and Sky make up the largest proportion at 10.53% and 10.33%, followed by Human and Water features at 7.93% and 6.69%. The proportion of Animal is the lowest at 0.13%.
In Fengjing, artificial elements account for 65.26%, while natural elements make up 34.74%. Among the artificial elements, Infrastructure (13.25%) and Architecture (13.11%) have the highest proportions, followed by Artifacts (6.51%), Information (6.18%), Furniture (6.10%), Structure (5.14%), and Transportation (4.02%). The proportions of Geomaterials and Art and Culture are relatively low at 3.40% and 3.16%, while Landmarks and Foods have the lowest proportions, at 2.94% and 1.43%. Among the natural elements, Vegetation and Sky dominate with proportions of 10.61% and 10.22%, followed by Human and Water features at 7.51% and 6.24%. The proportion of Animal is the lowest at 0.16%.
3.1.2 Perception of space
The spatial perception statistics of two towns are shown in Fig. 8. In Zhujiajiao, horizontal space accounts for 22.74%, while vertical space accounts for 77.26%. Among the Horizontal spaces, Corridor Space_Regular has the highest proportion, 11.13%, followed by Wide Corridor Space and Panoramic Space, at 6.90% and 2.54%. The proportion of Corridor Space_Narrow is the lowest, at 1.01%. In Vertical space, Vertical Space_Foreground and Vertical Space_Middle have the highest proportions, at 20.60% and 16.77%, followed by Oblique Space_Distant, Vertical Space_Close Up, and Oblique Space_Middle, at 14.31%, 10.01%, and 9.77%, respectively. The proportion of Vertical Space_Distant is the lowest, at 6.97%.
In the result of Fengjing, horizontal space accounts for 19.67%, while vertical space accounts for 80.33%. Among the horizontal spaces, Corridor Space_Regular has the highest proportion, at 13.74%, followed by Corridor Space_Narrow, at 5.46%. Corridor Space_Wide accounts for 0.37%, which is nearly absent in Fengjing Ancient Town, while Panoramic Space is completely absent, at 0.00%. In the vertical space, Vertical Space_Foreground and Vertical Space_Close Up have the largest proportions, at 25.98% and 19.66%, followed by Oblique Space_Distant, Vertical Space_Middle, and Oblique Space_Middle, at 13.71%, 10.32%, and 8.79%. The proportion of Vertical Space_Distant is the lowest, at 1.98%.
3.2 Comparative analysis
The results of the two-sample z-test comparing the perception of elements and space in Zhujiajiao and Fengjing are shown in Table 4. Since the Panoramic Space in Fengjing is 0, this type of space was not included in the z-test. In the comparison of element perception, half of the elements show no significant differences. Only eight elements―Structure, Infrastructure, Artifacts, Transportation, Furniture, Information, Art and culture, Foods―demonstrate significant differences, suggesting that in tourist photos, the perception of landscape elements in both towns tends to be similar. This result could provide reasonable evidence for the idea of homogenization. However, the results of the two-sample z-test on spatial perception show that Zhujiajiao and Fengjing have significant differences in most space types, such as Corridor Space_Wide, Corridor Space_Regular, Corridor Space_Narrow, Vertical Space_Distant, Vertical Space_Middle, Vertical Space_Foreground, Vertical Space_Close up, with only Oblique Space_Distant and Oblique Space_Middle showing no significant differences in perception (Table 4). Therefore, from the perspective of statistical comparison, the towns show a certain similarity in element perception but significant differences in spatial perception, which seems to challenge the notion of homogenization.
Evidence from spatial perceptions shows that the main space types perceived in Zhujiajiao are Corridor Space_Wide, Panoramic Space, and Vertical Space_Distant, while the frequently perceived space types in Fengjing are Corridor Space_Regular, Corridor Space_Narrow, Vertical Space_Middle, and Vertical Space_Close Up. The spatial layout of Zhujiajiao determines that visitors pay more attention to Corridor Space_Wide and Panoramic Space centered around the Dianpu River. However, when the focus is primarily on expansive areas such as the Fangsheng Bridge and Dianpu River, smaller-scale corridor spaces tend to be overlooked. In contrast, Fengjing Ancient Town does not have the same spatial configuration as Zhujiajiao, making it difficult to capture wide corridor spaces and panoramic views. The spatial layout of Zhujiajiao determines that visitors pay more attention to the broader corridor spaces and panoramic spaces centered around Dianpu River. However, when the focus is primarily on expansive areas such as the Fangsheng Bridge and Dianshan Lake, smaller-scale corridor spaces tend to be overlooked. In contrast, Fengjing Ancient Town does not have the same spatial configuration as Zhujiajiao, making it difficult to capture wide corridor spaces and panoramic views. While Corridor Space_Narrow in Fengjing do not attract as much attention as Corridor Space_Regular, they are not entirely neglected. On one hand, the spatial layout of the ancient town is primarily characterized by Corridor Space_Regular, which is filled with shops and people and serves as the main public space. On the other hand, Corridor Space_Narrow requires a deep familiarity with the town to evoke a sense of resonance. For instance, residents who live near these narrow corridors use them daily as essential routes to those main thoroughfares, whereas visitors generally do not engage in activities in these spaces.
3.3 Correlation analysis
The results of the correlation analysis are primarily presented in the form of matrix heatmaps. Two aspects of correlations are calculated: first, the correlation between elements, to explore which landscape elements are perceived and combined in different towns; second, the correlation between elements and space, to explore which types of spaces are associated with specific landscape elements in the two towns.
3.3.1 Correlation of elements
The correlation analysis first calculated the relationships between landscape elements. Overall, the comparative analysis results indicate that the element perception in different towns shares certain similarities. Therefore, further analysis of the relationships between elements can help move beyond simple statistical and comparative analysis, providing a more detailed understanding of the similarities and differences in the specific composition of elements in the two towns. Figure 9 respectively illustrates the element correlations in Zhujiajiao and Fengjing.
In both Zhujiajiao and Fengjing, several key elements show significant correlations with the other elements, including Architecture (Zhujiajiao 0.66 and Fengjing 0.64 with Water, 0.86 and 0.87 with Vegetation, 0.87 and 0.87 with Sky, 0.73 and 0.69 with Human, 0.58 and 0.56 with Furniture, 0.62 and 0.61 with Artifacts, 0.58 and 0.60 with Information, 0.60 and 0.56 with Structure, 0.91 and 0.93 with Infrastructure), Infrastructure (0.60 and 0.61 with Water, 0.81 and 0.84 with Vegetation, 0.79 and 0.81 with Sky, 0.74 and 0.70 with Human, 0.64 and 0.62 with Furniture, 0.67 and 0.64 with Artifacts, 0.59 and 0.60 with Information, 0.59 and 0.53 with Structure, 0.91 and 0.93 with Architecture), and Structure (0.58 and 0.56 with Water, 0.60 and 0.58 with Vegetation, 0.62 and 0.59 with Sky, 0.60 and 0.56 with Architecture, 0.59 and 0.53 with Infrastructure) in artificial elements. Water (beside Architecture, Infrastructure and Structure, 0.72 and 0.69 with Sky, 0.72 and 0.70 with Vegetation), Vegetation (0.66 and 0.64 with Human, 0.85 and 0.86 with Sky, 0.72 and 0.70 with Water), Sky (0.68 and 0.64 with Human, 0.85 and 0.86 with Vegetation, 0.72 and 0.69 with Water), and Human (0.68 and 0.64 with Sky, 0.66 and 0.64 with Vegetation) in natural elements (Fig. 9). These elements have relatively significant correlations with all other elements because they constitute the fundamental landscape of water towns. The critical role of Vegetation in water towns and its high correlation with Human provides a new perspective for understanding ancient towns––traditionally, the image of Jiangnan water towns has often been summed up with the phrase “A tapestry of arched bridges, meandering canals, and traditional residences.” However, our study demonstrates the significant role that Vegetation plays in these traditional built environments. This highlights that Jiangnan’s traditional settlements coexist not only with Water as a natural element but also with Vegetation, reflecting the ecological wisdom of a harmonious relationship between humans and nature.
Other interesting relationships include the notable positive correlations between Artifacts and Information (0.57 and 0.55), Furniture (0.66 and 0.67), and Human (0.58 and 0.57) in the perceptions of both towns (Fig. 9). Additionally, in Zhujiajiao, the relationship between Transportation and other elements is closer than in Fengjing. This element shows stronger positive correlations with Water (0.69), Vegetation (0.57), Sky (0.59), and Human (0.56) (Fig. 9(a)). This difference echoes the results from the comparative analysis.
Upon reviewing the elements in the photos, we found that the reason Transportation plays a more prominent role in Zhujiajiao is due to the significant presence of “boats”. In the results of the semantic segmentation, the most frequent element in Transportation is “boat” (Fig. 10), which also explains why Transportation shows the highest correlation with Water (0.69) in Zhujiajiao.
3.3.2 Correlation of elements and space
In order to assess whether space has an influence on the perception of elements, or whether elements can to some extent determine people’s choice of space, this study calculated the correlation matrix between elements and space (Fig. 11).
To reduce the dimensionality and visualize inherent data structure-based positional relationships between the 16 elements and special perceptions, PCA is utilized (Fig. 12). The PCA visualization elucidates how people perceive spaces and elements in the two ancient towns. Dots represent different spatial types, positioned according to their principal component scores, with those near the center indicating less distinctive features. Arrows indicate the influence of each element on the components, where longer arrows signify stronger contributions. The angle between an arrow and a dot reflects their relationship: sharper angles denote tighter correlations. Overall, both towns exhibit similar patterns. Corridor Space_Narrow are closely aligned with Architecture, while natural elements such as Water and Vegetation are more associated with open, expansive spaces like Corridor Space_Wide and Panoramic Space. Conversely, tight vertical spaces, such as Vertical Space_Close up and Vertical Space_Foreground, tend to be linked with humanmade elements like Art and culture, Furniture and Foods. Notable differences also emerge. In Zhujiajiao, Transportation is strongly associated with Corridor Space_Wide and Panoramic Space (Fig. 12(a)), whereas in Fengjing, cultural elements such as Art and culture, Furniture are more tightly connected with Vertical Space_Close up (Fig. 12(b)).
The key relationships between elements and space can be elaborated as follows.
(1) Transportation: Consistent with the results reflected in the correlation analysis between elements, the characteristic of being perceived through boats as the main transportation element in Zhujiajiao is also reflected in its space. Transportation has a positive correlation with Corridor Space_Wide (0.42), Corridor Space_Regular (0.23), Panoramic Space (0.57), Vertical Space_Distant (0.22), and Oblique Space_Distant (0.36). Notably, the strong correlations with Corridor Space_Wide and Panoramic Space indicate that large-scale water bodies like the Dianpu River not only facilitate active boating but also command significant sensory attention from visitors (Fig. 13). The river’s expansive surface and panoramic vistas create an immediate visual draw, capturing the gaze even from distant viewpoints. The movement of boats further enhances intuitive engagement, reinforcing the water’s role as a perceptual focal point.
Unlike Zhujiajiao, the perception of transportation in Fengjing is more prevalent in smaller-scale spaces, such as Corridor Space_Narrow (0.45) and Vertical Space_Middle (0.23). Although there are boats in Fengjing, their perceived presence is not as strong as in Zhujiajiao. Instead, bicycles and minibikes are more frequently observed in Fengjing’s transportation element, which reflects the preserved living atmosphere of the town (Fig. 14).
(2) Artifacts, Furniture, Information, Art and Culture: An observation of these elements reveals a noticeable tendency, where they are mostly perceived in smaller-scale spaces, such as Corridor Space_Narrow, Vertical Space_Middle, Vertical Space_Foreground, and Vertical Space_Close up (Fig. 15). This suggests that these elements are typically observed at closer distances. This trend is even more apparent in Fengjing, where these elements are almost exclusively correlated with Vertical Space_Foreground and Vertical Space_Close up. It seems that in Fengjing, people tend to focus their lenses on shops or local life. Many of the shops in Fengjing sell bamboo products, which are recognized by Segformer as baskets, and many hang signs, which are detected as flags. Much like bicycles and minibikes, Fengjing has retained more traces of local life. As a result, these traces of local life, such as Artifacts and Furniture, appear to become objects of focus for tourists in Fengjing. Everyday items like trash cans are also detected. Additionally, Fengjing’s murals and books in shops―representing Art and Culture―also become objects of careful observation by visitors. These elements are captured at a close distance in Fengjing, forming a correlation with Vertical Space_Foreground and Vertical Space_Close up.
(3) Landmark: In contrast to Artifacts, Landmark are always more closely related to larger-scale spaces (Fig. 16). In Zhujiajiao, Landmark has a stronger correlation with Corridor Space_Wide (0.82), Panoramic Space (0.33), and Vertical Space_Distant (0.43). The main reasons for this association are the open water surfaces and landmarks such as the Fangsheng Bridge which is located along the waterways. In Fengjing, Landmark correlates more strongly with Corridor Space_Wide (0.91), Corridor Space_Regular (0.58), Vertical Space_Distant (0.47), Vertical Space_Middle (0.36), and Oblique Space_Distant (0.46). This reflects the distribution of landmarks in the alleyway spaces, and the correlation with Oblique Space_Distant also indicates that people tend to observe Fengjing’s landmarks while moving through the space.
(4) Foods, Animal: Both Foods and Animals show a significant positive correlation with Vertical Space_Close up in Zhujiajiao and Fengjing (1.22 and 1.0, 1.36 and 0.98) (Fig. 17). This aligns with common sense, as people are accustomed to observing animals and food through close-up shots.
(5) Vegetation, Water, Sky: The distribution of natural elements in spatial perception also follows the pattern that, in the correlation analysis with more open spaces, natural elements show higher values. Water has some differences in its spatial relationship between the two towns. In Zhujiajiao, the positive correlations of Corridor Space_Wide (0.56), and Panoramic Space (0.65) with Water are more significant. In Fengjing, Water is more significantly correlated with Corridor Space_Wide (0.70), Corridor Space_Regular (0.44), and Oblique Space_Distant (0.57). From a perceptual perspective, natural elements―especially water bodies―consistently exhibit high attractiveness in both towns. Despite their distinct spatial configurations, human preferences for water-related landscapes reveal a shared inclination. The differences in spatial perception are primarily due to the distinct water body patterns in the two towns: Zhujiajiao’s expansive water surfaces versus Fengjing’s linear river corridors. Additionally, people may tend to take photos with water scenes while moving, highlighting the directional nature of the rivers.
4 Discussion
When all potential forms of communication are learned and respected, multiple and infinite forms of knowledge can be further understood (
Innes, 1995;
Sheydayi and Dadashpoor, 2023). By analyzing user-generated visual content, this study offers a deeper understanding of spatial perception and provides insights into the preservation of traditional environments.
4.1 Combining depth estimation and semantic segmentation provides more comprehensive landscape perception
This study employs depth estimation and semantic segmentation to analyze UGC from traditional settlements, offering fresh insights into landscape perception. Segformer, the visual processing tool employed, excels in image segmentation while maintaining computational efficiency and robustness to varying image quality, making it highly suitable for UGC analysis. The integration of depth estimation enhances the traditional semantic segmentation approach by adding a spatial dimension, thereby revealing subtle details in the UGC.
The immediacy of social media ensures that most tourist expressions are brief, spontaneous, and immediate, often conveyed through photographs rather than detailed text. This phenomenon reflects the growing performative nature of modern society, vividly described as an “online playground” where interactions tend to be less reflective (
Munar, 2010;
Munar and Ooi, 2012). However, the significance of UGC should not be underestimated. Not only is its volume substantial, but people often assume that what they share represents the totality of their knowledge, when in fact, individuals know much more than they can articulate (
Tuan, 1977). Only through thorough and meticulous analysis of UGC can we gain deeper insights.
This study reveals new insights into the ancient towns of Zhujiajiao and Fengjing that were previously unexplored in prior research. The comprehensive analysis of their elements and spaces highlights the social, cultural, and customary forces that have shaped these towns, viewed through contemporary lenses. For example, the perception of Zhujiajiao’s extensive waterways, rich aquatic landscapes, and boating activities is deeply rooted in its longstanding socio-cultural context. Local traditions such as “rocking fast boats,” “fist boats,” “lantern boats,” and “silk-bamboo boats” encapsulate the local saying, “With a single oar in hand, the water becomes your road.” Historically, residents depended on boating for trade and social interactions, which has earned it the modern nickname of “Venice of Shanghai.”
In Fengjing, local life and artistic creations accentuate its vertical spatial characteristics. Works by native artists like Ding Cong and Cheng Shifa, along with murals by French street artist Julien Malland, imbue the water town with an artistic atmosphere, influencing visitors’ perceptions. The presence of hanging signs is closely linked to the corridor spaces. These signs, including traditional wine banners, have largely disappeared from commercial districts and neighborhoods across China, signaling a loss of traditional street aesthetics (
Crane, 1926), or the decline of corridor spaces as commercial centers.
4.2 Homogenization is absent in spatial perception
This study has prompted reflections on the visual homogenization of ancient water towns. The findings highlight the need for further consideration of the “thousand towns with one face” phenomenon in water towns. When discussing the uniformity of water towns, one often envisions a few streets, several bridges, buildings with strikingly similar styles, and nearly identical boats. Admittedly, the element perception derived from semantic segmentation in this study clearly confirms homogenization as an undeniable fact. However, it is important to note that the homogenization demonstrated here is limited to the element level. By incorporating spatial analysis, the study not only refutes homogenization from a spatial perspective but also shows that the negation of homogenization holds true when space and elements are interrelated.
From a visual perspective, ordinary tourists may find it challenging to distinguish one water town from another, as these towns, distributed across the Taihu Lake basin, share similar architectural styles and natural landscapes. However, when examining the underlying spatial structures beyond the visible material forms, the social, cultural, and customary dimensions associated with the built environment become evident. The findings, based on UGC analysis, reveal that even in spontaneous photography, tourists intuitively grasp spatial structures, which resonate with broader social and cultural themes. This challenges the notion of homogenization.
The findings indirectly highlight a fundamental aspect of China’s settlements: whether imperial cities or courtyards, their forms have remained largely unchanged over millennia because the visible form is not the primary focus, the underlying meaning holds greater significance (
Ruan and Huang, 2020). What truly endows water towns with meaning is the socio-cultural life and the spatial structures they shape, while the physical appearance is secondary. This approach introduces a novel method for conserving traditional settlements by emphasizing the comprehensive preservation of spatial structures to support socio-cultural practices, rather than focusing exclusively on the restoration of physical appearances.
4.3 Transformation of the way of perception
The nostalgic appeal of ancient towns’ social spaces cannot obscure their contemporary reality: most no longer retain multigenerational residents, and few ancestral homes remain owner-operated, with rental commercialization now dominating. While outsiders idealize these environments, locals pragmatically prioritize leasing old properties for income and relocating to modern apartments―a choice reflecting China’s architectural philosophy that historically rejected material preservation as essential (
Ruan and Huang, 2020). Ordinary families consistently value present livelihoods over architectural heritage, rendering unrealistic any expectation for them to invest resources in maintaining obsolete material forms.
This dilemma arises from economic transformation, which also shapes the changing perceptions of ancient towns. Historically, water towns emerged as rural exchange hubs, evolving from temporary markets to permanent commercial communities (
Fei, 1996). As nodal points in China’s pre-modern rural economy, they formed a system enabling self-sufficient lifestyles through commercial interdependence. Water towns specifically developed as transport-oriented trade communities along waterways. Their defining “exchange function” structured traditional life: farmers routinely visited towns to trade in corridor-style commercial spaces such as shops, taverns, and teahouses, whose architectural forms mattered less than their transactional purpose. In essence, individuals who lived in or visited ancient towns in the past did not primarily experience them through visual means. Rather, their perception was deeply embodied, shaped by rituals, customs, and a range of multisensory interactions such as movement, touch, conversation, scents, sounds, and emotional connections. These elements collectively created a rich, holistic engagement with these spaces. However, as this study underscores, the contemporary globalized and digital world, saturated with photography and social media, has shifted our perception of ancient towns to one that is predominantly visual.
Experience-driven tourists seek novelty over traditional lifestyles and can only visually appreciate the surface aspects of environments they cannot functionally engage with. The future of ancient towns depends not on rigidly preserving traditions, but on adaptive modernization. The revitalization of ancient towns necessitates the infusion of both functional and social vitality. While preserving their spatial structure, it is essential to introduce more participatory activities, whether traditional or emerging. The ultimate goal is to transform mere visual perception into embodied perception.
4.4 Limitations of this study
This study offers new insights into landscape perception, but it is crucial to acknowledge its limitations. Firstly, the sources of UGC are somewhat restricted. The selected platforms do not encompass all user demographics, particularly excluding older adults and those without internet access. Nonetheless, the vast volume of UGC compensates for the limitations of traditional data in terms of quantity and collection methods, providing more diverse, detailed, and dynamic insights. Additionally, while we focused on typical water towns in Shanghai, the comparison between the two towns may still be inadequate. Future research should aim to include a broader range of water towns for in-depth analysis and incorporate a wider array of social media platforms to enhance data diversity.
Another limitation is the variability in the quality of social media images, especially those from a decade ago, which may include low-resolution photos. To mitigate this issue, this study utilized SegFormer, which is less sensitive to image resolution. Future research could benefit from obtaining higher-quality images and employing more advanced models to conduct more precise and in-depth analyses of spatial structures.
Furthermore, this study did not account for the temporal dimension. However, it is noteworthy that temporal factors, such as seasons, months, and hours of the day, can significantly influence perception due to their relationship with natural elements. Future research should consider incorporating these temporal factors into perception studies, which may yield novel insights.
5 Conclusions
This study performs a landscape perception analysis of UGC in water towns by integrating depth estimation technology and a spatial analysis framework that considers both vertical and horizontal dimensions, alongside semantic segmentation. The research addresses the limitations of prior UGC-based landscape perception studies, which were limited to semantic segmentation alone. By providing detailed insights at the element level and incorporating the concept of spatial structure, the study categorizes the spatial perception of ancient water towns into 10 distinct types, thus broadening the scope of landscape perception research.
The integration of semantic segmentation and depth estimation has shown promising applications in UGC studies. Selecting an appropriate Segformer model ensures accurate segmentation results with minimal sensitivity to image quality, making it well-suited for social media image analysis. Depth estimation using the dpt-large model provides valuable insights into the spatial content of images. Further correlation analysis between spatial and elemental features reveals how social, cultural, and other factors are manifested spatially through these elements. This methodology can be widely applied to future image-related perception research.
Through a joint analysis of element and spatial perception, this study highlights the need to reflect on the concept of homogenization. Semantic segmentation results reveal that the two ancient towns share similar element perceptions, reflecting homogenization characteristics under visually dominant perception modes. However, significant differences in spatial perception suggest that homogenization is primarily a phenomenon of visual perception. The true essence of human settlement is shaped by social, cultural, and lifestyle factors embedded in spatial structures, which are more deeply perceived through embodied experiences. These findings offer valuable insights into the preservation of traditional settlements, highlighting the importance of studying spatial structures and integrating their functional values to enhance modern users’ embodied perception.
2095-2635/2025 The Authors. Publishing services by Elsevier B.V. on behalf of KeAi Communications Co. Ltd.