DeepTagPhoto: Expert-guided unsupervised clustering of Thai architectural photography using pre-trained CNN models

Walaiporn Nakapan , Chi-tathon Kupwiwat , Chavanont Khosakitchalert

Front. Archit. Res. ›› 2026, Vol. 15 ›› Issue (4) : 1160 -1173.

PDF (10277KB)
Front. Archit. Res. ›› 2026, Vol. 15 ›› Issue (4) :1160 -1173. DOI: 10.1016/j.foar.2025.09.012
RESEARCH ARTICLE
DeepTagPhoto: Expert-guided unsupervised clustering of Thai architectural photography using pre-trained CNN models
Author information +
History +
PDF (10277KB)

Abstract

This study presents an unsupervised image clustering approach to support expertled analysis and tagging of traditional Thai architectural photography. This study evaluates how effectively deep-feature clustering derived from pre-trained CNN embeddings reveal meaningful groupings within Thai stupa photography without relying on existing labels. A subset of 824 images depicting stupas was selected from an original academic archival 7696-image dataset in order to investigate the feasibility of clustering based on deep features extracted from four pre-trained Convolutional Neural Networks (CNNs): VGG19, ResNet50, DenseNet, and InceptionV3. These embeddings were clustered using the K-means algorithm, and results were validated qualitatively by a domain expert. Among the models, ResNet50 achieved the lowest rejection rate and highest alignment with expert interpretations. Notably, the expert’s engagement with the clustering outputs led to the revision of prior annotations, revealing instances of label noise and enhancing intra-cluster specificity. These findings suggest that CNN-based clustering can mitigate the limitations of expert-dependent tagging, namely incompleteness, subjectivity, and inconsistency while offering a generalizable framework for similar heritage datasets with sparse or noisy labels. The study contributes a hybrid human-AI methodology for visual cultural heritage analysis and highlights opportunities in supervised style classification and automated archiving workflows.

Graphical abstract

Keywords

Convolutional neural network / Unsupervised learning / Traditional art and architecture / Visual clustering / Human-AI collaboration

Cite this article

Download citation ▾
Walaiporn Nakapan, Chi-tathon Kupwiwat, Chavanont Khosakitchalert. DeepTagPhoto: Expert-guided unsupervised clustering of Thai architectural photography using pre-trained CNN models. Front. Archit. Res., 2026, 15 (4) : 1160-1173 DOI:10.1016/j.foar.2025.09.012

登录浏览全文

4963

注册一个新账户 忘记密码

1 Introduction

1.1 Background and research challenges

1.1.1 Context and recent advances

The integration of machine learning into cultural heritage studies has significantly enhanced tasks such as image classification, retrieval, and large-scale archival analysis. Convolutional Neural Networks (CNNs) have proven effective for automatic feature extraction from visual data, enabling the classification of artworks, architectural styles, and artifacts with notable success (Cetinic et al., 2018; Llamas et al., 2017). These advancements are especially pronounced in well-structured datasets from Western collections, which benefit from consistent metadata and expert annotations.

However, the success of supervised learning in these contexts is heavily contingent on the quality of expert labeling. Manual annotation in the heritage domain is time-consuming and often shaped by interpretive ambiguity, stylistic complexity, and cultural context (Foka et al., 2025a, b; Ottoni and Ottoni, 2024a). This problem is exacerbated in underrepresented regions like Thailand, where domain-specific expertise is essential but inconsistently applied. Even expert-curated datasets may exhibit incompleteness, label noise, and subjectivity―issues that undermine the assumptions of standard supervised learning workflows (Reshetnikov et al., 2023; West, 2017; Zhang and Cao, 2025).

1.1.2 Challenges in expert tagging of cultural heritage images

Despite expert involvement, heritage image datasets often remain problematic for machine learning due to three key issues:

(1) Incompleteness: Many cultural heritage datasets suffer from partial or missing labels due to the laborious nature of annotation and the requirement for deep domain expertise. Belhi et al. (2018) identify incomplete tagging as a common obstacle in heritage collections. In our dataset, while elements like stupas or stucco work are frequently labeled, more nuanced iconographic or stylistic features are omitted, often requiring further scholarly inquiry. This results in annotation gaps that weaken the training of supervised models and risk introducing classification biases.

(2) Label Noise: Label noise refers to incorrect or inconsistent annotations that misrepresent the true category of an image. Song et al. (2022) emphasize how such noise adversely affects deep learning performance. In our context, label noise arises when images of real architectural elements (e.g., pagodas) are tagged similarly to symbolic representations in paintings or wood carvings. This semantic discrepancy introduces misleading signals during model training, compromising classification accuracy.

(3) Subjectivity: Labeling in the domain of cultural heritage is inherently interpretive, shaped by the individual scholar’s methodological lens, regional expertise, and cultural framework. Gupta and Naaz (2023) emphasize that subjective interpretation is central to the annotation of art and architecture, with different experts potentially assigning divergent labels to the same object. For instance, one scholar might classify a structure as “Lanna,” while another might describe it as “Chiang Mai style,” depending on their preferred stylistic taxonomy. Ottoni and Ottoni (2024b) further argue that expert labeling in cultural heritage exists in a constant tension between the pursuit of taxonomic precision and the interpretive flexibility demanded by culturally embedded artifacts. This interpretive nature of tagging is corroborated by West (2017) and Belhi et al. (2018), who reveal how both expert and public annotation practices can introduce inconsistency across large-scale visual collections, ultimately complicating the reliability of supervised learning systems.

1.1.3 Research gap

Popular CNNs such as VGG-19, ResNet-50, DenseNet, and Inception-v3 were trained with supervised learning on the labeled ImageNet-1K dataset, yet their use still centers mostly on Western art collections (He et al., 2016; Huang et al., 2017; Simonyan and Zisserman, 2015; Szegedy et al., 2016). Consequently, unsupervised approaches remain underexplored in Southeast Asian heritage domains. Prior studies focus on datasets with well-structured labels (Caron et al., 2018; Parisotto et al., 2022), but few investigate clustering techniques that circumvent label dependence altogether. The potential of human-in-the-loop evaluation in such settings also remains understudied (Goldfinch et al., 2025).

This study addresses these gaps by proposing a hybrid methodology: CNN-based feature extraction combined with K-means clustering, followed by qualitative validation by an expert curator. While the work focuses on a curated dataset of Thai architectural photography, the method offers a generalizable approach for organizing under-annotated visual collections across diverse heritage contexts.

1.2 Research objective and question

This study aims to evaluate the viability of unsupervised learning as an alternative to traditional supervised approaches for classifying visual datasets in cultural heritage. Specifically, we examine whether features extracted from pre-trained Convolutional Neural Networks (CNNs) can support image clustering that aligns with expert interpretations of visual and stylistic characteristics in traditional Thai art and architecture through subsequent expert qualitative review.

The dataset used in this study comprises a curated photographic collection developed over the past decade for academic archival purposes. Although annotated by a domain expert, the dataset reflects inherent challenges typical of expert tagging, namely incompleteness, label noise, and subjectivity. These issues limit their direct applicability for supervised learning.

To address these limitations, we adopt an unsupervised clustering approach using the K-means algorithm applied to visual features extracted from four established CNN architectures: VGG19, ResNet50, DenseNet, and InceptionV3. We then assess the interpretive validity of the resulting clusters through qualitative feedback from the same expert who originally labeled the dataset.

This study asks: Which CNN model most effectively produces visual groupings that align with expert categorization in Thai heritage photography? Addressing this question can help overcome the limitations of noisy, incomplete, and subjective human labels through scalable computational means.

1.3 Research contributions

The contributions of this paper are threefold:

(1) it presents a novel unsupervised learning workflow for analyzing traditional Thai heritage imagery;

(2) it demonstrates how human expert review can enrich the interpretability of machine-generated visual clusters; and

(3) it offers a generalizable methodological model for applying human-AI collaboration in other cultural heritage contexts where labeled data are incomplete or unreliable.

The remainder of this paper is organized as follows: Section 2 reviews relevant literature on CNN-based classification, clustering, and expert validation in the digital humanities. Section 3 describes the dataset, CNN feature extraction process, and the unsupervised clustering methodology. Section 4 presents the results of the clustering and expert evaluation. Section 5 discusses the implications and limitations of the findings, and Section 6 concludes the paper with directions for future research.

2 Related research on CNN-based classification of art and architectural images

2.1 CNNs as tools for image classification

Convolutional Neural Networks (CNNs) have become instrumental in the classification of art and architectural images due to their ability to learn hierarchical feature representations. Karayev et al. (2014) pioneered the use of pre-trained CNNs, specifically AlexNet, to classify paintings by style and genre, achieving superior performance over traditional hand-crafted feature methods. Building upon this, Bar et al. (2015) utilized CNN-derived features to classify a large dataset of paintings by artist and style, demonstrating the scalability of CNNs in art classification tasks. Tan et al. (2016) further enhanced classification accuracy by fine-tuning CNN models on extensive art datasets, underscoring the importance of domain-specific training.

Cetinic et al. (2018) conducted comprehensive experiments, fine-tuning CNNs for multiple art-related classification tasks, including artist, style, genre, period, and national context. Their findings highlighted that CNNs pre-trained on general image datasets could be effectively adapted for art classification, with scene recognition models outperforming those focused-on object recognition. Tan et al. (2018) introduced Domain-Adversarial Training to improve CNN efficiency across varied art datasets, while Kondo and Hasegawa (2020) explored artist classification, revealing that CNNs capture more than just stylistic features. Chen et al. (2021) provided a comprehensive review of various CNN architectures, including LeNet and AlexNet, discussing their applicability in art classification tasks.

2.2 Representative CNN models for visual feature extraction

Several Convolutional Neural Network (CNN) models have played a foundational role in advancing image classification tasks, particularly in applications involving complex visual datasets such as those in art and architecture. Among them, VGG19, developed by Simonyan and Zisserman (2015), is distinguished by its architectural simplicity and uniform use of small convolutional filters. Comprising 19 layers, VGG19 is widely adopted for feature extraction and transfer learning due to its straightforward design and strong performance in tasks requiring fine-grained visual representation.

Another influential model is ResNet50, introduced by He et al. (2016), which employs residual learning to address the vanishing gradient problem encountered in deep architectures. By incorporating skip connections that enable the network to bypass certain layers during training, ResNet50 allows for deeper, more accurate models while maintaining computational efficiency.

DenseNet, proposed by Huang et al. (2017), builds on this advancement by connecting each layer to every other layer in a feed-forward fashion. This design encourages feature reuse and enhances gradient flow, enabling the model to achieve high accuracy with a reduced number of parameters.

Finally, InceptionV3, introduced by Szegedy et al. (2016), represents a shift toward architectural efficiency using factorized convolutions and auxiliary classifiers. Its inception modules capture multi-scale features simultaneously, making the network highly effective for large-scale image classification tasks that require both precision and computational scalability.

2.3 Unsupervised clustering and expert-guided interpretation in cultural heritage imaging

Recent advances in unsupervised learning have demonstrated the effectiveness of CNN-based clustering methods for organizing large-scale image collections in cultural heritage domains, particularly where manual annotation is impractical. Caron et al. (2018) introduced DeepCluster, a technique that iteratively applies K-means clustering to CNN-extracted features, allowing networks to learn visual representations without labeled data. This foundational method has inspired a range of heritage applications. For example, Parisotto et al. (2022) applied a variational autoencoder-based clustering pipeline to Roman pottery profiles, with expert archaeologists evaluating the results to ensure that the automatically generated clusters reflected meaningful typological groupings. Similarly, Castellano and Vessio (2022) proposed DELIUS, a deep embedded clustering framework for unlabeled painting collections, successfully uncovering visual similarities and stylistic trends. Gultepe et al. (2018) also demonstrated that CNN-based unsupervised feature learning could distinguish artistic styles in digitized paintings, even in the absence of predefined categories. These approaches have extended beyond fine arts into broader cultural heritage contexts: Llamas et al. (2017) applied deep CNNs to automate the classification of architectural heritage photographs, while Wang (2024) used clustering methods to organize vast museum artifact collections, showing the potential of unsupervised learning to support curatorial workflows. In the realm of architectural image classification, Xu et al. (2014) employed a Multinomial Latent Logistic Regression (MLLR) model to classify architectural styles, creating a dataset encompassing 25 styles and analyzing the relationships between styles and building features. Despite the effectiveness of CNNs in art and architecture classification, challenges persist, such as handling limited datasets, improving interpretability for art historians, classifying subtle styles and influences, and incorporating contextual and historical information.

Crucially, these studies highlight the value of expert feedback in validating and interpreting machine-generated groupings, reinforcing the importance of domain knowledge in bridging computational results with historical and cultural significance. This interpretive layer becomes especially relevant in regional contexts that have received limited computational attention. For instance, Goldfinch et al. (2025) applied deep learning to prehistoric rock art and ceramics from Thailand’s Khorat Plateau, demonstrating that CNN models can uncover visual patterns previously difficult to identify manually. Similarly, Kuntitan and Chaowalit (2022) classified Sukhothai-era ceramic motifs using fine-tuned CNNs, showing high accuracy in detecting regional stylistic traits based on expert-labeled data. These examples from Southeast Asia underscore the growing applicability of unsupervised and expert-validated machine learning in underrepresented cultural domains and directly inform the present study’s use of CNN-based clustering in the context of traditional Thai art and architecture. Despite the demonstrated effectiveness of both traditional and CNN-based classification methods, key challenges remain, particularly in managing limited training datasets, improving model interpretability for domain experts, classifying nuanced or overlapping styles, and incorporating historical and contextual metadata into machine learning workflows.

3 The dataset

3.1 The importance of the dataset

The dataset used in this study holds significant historical and cultural value, serving as a critical resource for the study of Thai art and architecture. It comprises photographic documentation captured by Assoc. Prof. Dr. Wichai Posayachinda, a Doctor of Medicine who travelled extensively across Thailand in the 1970s during his field investigations into drug-related issues. Notably, many of these photographs were taken in the same locations and from similar perspectives as those of Mr. Prayul Uluchata (also known as Nor Na Paknam), a celebrated Thai artist who was named National Artist in Fine Art (Painting) in 1992. This intentional mirroring of visual perspectives suggests that Dr. Wichai possessed a deep appreciation and understanding of Thai art and architectural composition. As such, the dataset not only documents historical sites but also reflects a unique visual dialogue between a medical professional and an acclaimed artist. This collection, curated by the Center for Visual Study, Chulalongkorn University (2008—2013), was enriched by expert tagging over the past decade. It represents a valuable yet underexplored archive for the academic community. This image collection, comprising thousands of photographs of temples, stupas, mural paintings, and decorative elements across Thailand, was initially assembled for scholarly and archival purposes. It later served as the foundational visual material for the Following the Old Images book series, a five-volume publication that juxtaposes black-and-white images captured in the 1970s with color photographs taken between 2011 and 2012 (Chanthawilasawong, 2010a—e). These books―including Temples in Ayutthaya, Buddhist Art of the South, Magnificent Lanna, Sukhothai Craftsmanship, and Art Beyond the Old Capital―reflect Dr. Wichai’s vision of tracing stylistic and material transformations in Thai religious architecture over time. The curated nature of this dataset, informed by expert photographic practice and thematic organization, offers a rare opportunity to assess how unsupervised machine learning methods can support the visual classification and interpretation of traditional cultural heritage imagery.

3.2 The primary dataset for investigation

The original dataset on art and architecture comprises 7696 images, each tagged by a human expert with Level 1 and Level 2 keywords in Thai. Level 1 keywords are for visible items such as temples, pagodas, and stucco; Level 2 for elements noticeable only by experts, like specific roof ornamentals and indented corners of stupas; and Level 3 for relationships between elements. However, the collection has not been tagged with Level 3 vocabularies due to the deeper understanding and research required. The most frequently used tags (cf. Fig. 1) include stucco (2229 occurrences), painting (1,843), Ayutthaya (1,399), mural painting (1,369), wood carving (1,054), Bangkok (927), pagoda (793), angel (742), stupa (701), Mahathat temple (662), Buddha statue (646), Buddha’s story (621), Buddha (620), Phetchaburi (599), Naga (587), Chiang Mai (543), lotus petal pattern (525), stripe pattern (518), and tympanum (513). Keywords are ranked in descending frequency, with 204 tags appearing only once. For this study, a subset of 824 images tagged with “stupa” and related terms, such as “stupa style,” was selected to form the primary dataset for investigation.

4 Methodology

The methodology combines quantitative and qualitative analyses to evaluate and compare the clustering efficacy of four Convolutional Neural Network (CNN) models applied to a Thai art and architecture dataset.

Quantitatively, the study performs unsupervised image clustering with four CNN models, aiming to assess and compare their performance. From the original dataset of 7696 image-text pairs, a subset of 824 images labeled as “stupa” was selected for detailed analysis. For each image in this subset, feature embeddings were generated using pre-trained CNNs, where each image is encoded as a vector in an embedding space, capturing essential visual features that facilitate clustering and comparison of images based on visual similarity. These embeddings were then clustered using K-means to group the images.

The qualitative evaluation involves expert reviews to gauge the classification accuracy and model effectiveness for tagging and categorization, with expert feedback providing insights into the models’ interpretability and relevance. Statistical analysis was subsequently employed to examine the consistency of images within each cluster. Detailed descriptions of each methodological step are provided as follows:

4.1 Image embedding

Four pre-trained CNN models are utilized to embed the image into vector data for unsupervised clustering. Despite using CNNs, these models have different neural architectures and are trained under different datasets and training experiment settings. Therefore, they may create different embeddings which eventually lead to different clustering, which needs to be verified. The CNN models studied in this research are as follows: VGG19, ResNet50, Densenet, InceptionV3. These pre-trained CNN models were obtained from TensorFlow (Abadi et al., 2016) and modified by removing the original fully connected layers for prediction. Note that the outputs of these modified models are still embedded tensors which need to be modified into a vector. The global average pooling (Alzubaidi et al., 2020) is applied after the last convolutional layer. This reduces the spatial dimensions of the output to a single value per feature map, creating an embedded 1-D vector, a data point in high dimensional space, that can be classified using the K-Means algorithm.

4.2 Unsupervised K-Means clustering

K-Means clustering is an unsupervised learning approach used to divide data points into K different clusters based on similarity, enabling the discovery of intrinsic groupings within the dataset. This research uses K-Means on the embedding vectors produced by pre-trained CNN models, rather than on the raw images themselves. These embeddings offer a concise, high-level representation of each image by encapsulating critical visual elements in a lower-dimensional vector space, efficiently summarizing intricate patterns while minimizing computing requirements.

K-Means clustering aims at minimizing the sum of squared distances between data points and their respective cluster centroids. Let X={x1,x2, …,xn} be the dataset containing data point xi∈ℝd where d is the feature or dimension of the data. K-Means clustering partition X into K cluster {C1,C2, …,CK} by minimizing the cost function defined as

(1)J=∑j=1K∑xi∈Cj‖xi−μj‖2,

where ‖xi−μj‖2 is the squared Euclidean distance between a data point and the centroid μj of the cluster Cj. The centroid μj is computed as

(2)μj=1|Cj|∑xi∈Cjxi.

The algorithm is initialized by randomly selecting K datapoint from X as the initial centroids {μ1, μ2, …, μK}. Then the iterative step of:

1) assigning each datapoint xi to the cluster Cj with the nearest centroid as

(3)Cj={xi:‖xi−μj‖2≤‖xi−μk‖2| ∀ k=1,..,K},

and 2) after all points have been assigned, recompute each centroid μj from the mean of all points in Cj using Eq. (2). The algorithm keeps iterating until the centroids no longer change or when a maximum number of iterations is reached.

4.3 Human expert evaluation

Because unsupervised clustering lacks interpretative capacity, qualitative validation was conducted by a senior historian and archaeologist with extensive expertise in Thai art and architecture. This expert, who also carried out the original annotation of the full 7696-image dataset using Adobe Bridge, possesses long-standing curatorial experience and has previously applied a consistent tagging framework to the collection.

4.3.1 Review environment

To examine the outputs of each clustering model, the images were exported into folders according to their assigned cluster labels (K = 3, 4, and 10). The expert systematically reviewed each folder in grid format, which enabled both holistic comparison and rapid anomaly detection. For every cluster, she first articulated an overall descriptive concept that captured the predominant characteristics of the group. For example: “Stupas with 75% height visible in close-up, without the pointed top” (VGG19, K = 4, cluster 4). These cluster-level summaries served as benchmarks against which individual images were judged.

4.3.2 Evaluation criteria

The expert applied a consistent set of visual and stylistic criteria, including (i) view type (full stupa, partial close-up, wide-angle, or landscape context), (ii) stupa type (bell-shaped, prang/tower, or star-fruit form), (iii) stupa shape (slender, flat-topped, indented corners, or tiered bases), and (iv) ornamentation and decorative motifs (stucco, statues, brass sheet gilding). A schematic of these categories, developed by the expert, is presented in Fig. 3 (adapted from her notes; see “Key Elements” diagram). This framework provided a structured lens for determining intra-cluster coherence.

4.3.3 Rejection process

Images that diverged from the dominant cluster concept were flagged as “rejected.” Rejections fell into several categories: (1) images not depicting a stupa as an edifice (e.g., mural paintings or wood carvings containing stupa motifs), (2) incorrectly tagged images from the original dataset, and (3) poor-quality or blurred photographs that obscured architectural features. Each rejection was recorded, and percentages were calculated at cluster and model levels to quantify alignment between machine-generated groupings and expert interpretation.

4.3.4 Qualitative refinement

In addition to accept/reject judgments, the expert’s engagement revealed opportunities to refine prior labels. Certain clusters surfaced consistent visual subtypes that were originally grouped under generic “stupa” labels. For instance, ResNet50 clusters highlighted bell-shaped stupas with distinct lotus-petal bases, prompting reclassification of several images. These refinements illustrate the value of unsupervised clustering in stimulating more granular expert categorization.

4.3.5 Reliability

Although only a single expert participated in this study, intra-rater consistency was strengthened through a reflexive protocol. Each cluster was revisited after a two-week interval, and judgments were cross-checked against the initial evaluation. Any discrepancies were reconciled in consultation with the research team, thereby enhancing the robustness of the qualitative findings.

5 Results

The embedding process utilizing CNN models is performed on a LANTA supercomputer equipped with NVIDIA A100 GPUs. Additional computations are conducted on a PC featuring a 2.3 GHz Dual-Core Intel Core i5 CPU and an Intel Iris Plus Graphics 640 GPU with 1536 MB of memory.

The K-Means implementation in Scikit-Learn (Pedregosa et al., 2011) is employed through embedded data from all CNN models using K clusters of 3, 4, and 10. These number K clusters are obtained by observing the final cost function J in Eq. (1) when changing K, known as the elbow method. Figure 2 shows the cost function of each embedded data from different CNN models. The effective K is around 3 and 4. However, K = 10 is also selected for its potential to improve classification differentiation, despite the continued improvement in the cost function at this value.

Figure 3 illustrates the key elements identified by the expert, including viewpoint, ornament, shape, and stupa type. Examples of cluster definitions based on these elements are provided below:

• "Stupas with 75% height visible in close-up, without the pointed top." (VGG19, K = 4, cluster 4),

• "Wide-angle views of stupas, some with extensive landscapes around, and others showing the entire stupa." (ResNet50, K = 3, cluster 3),

• "Bell-shaped stupas and bell shapes atop stupa towers; some images show middle or base parts, with some having a star fruit-like top" (Densenet, K = 10, cluster 4),

• "Wide-angle views of pointed stupas" (InceptionV3, K = 4, cluster 3),

• "Tall and slender stupas (full view)" (InceptionV3, K = 10, cluster 1).

Figures 4—7 provide examples of image clusters generated from four different CNN models (VGG19, ResNet50, DenseNet, InceptionV3) using K = 3 (a), 4 (b), and 10 (c), with images categorized by visual elements such as the stupa, its ornaments, or stupa with surrounding landscape.

Figure 4 presents examples of images obtained from VGG19 clustering. At K = 3, the images are grouped into [i] All types of stupas, some showing full views from a distance and some parts, but without landscape, [ii] General stucco work on architectural elements and low-relief sculptures, and [iii] Overall stupa views, with distant views showing surrounding landscapes and full stupas. With an increase to K = 4, the partitioning becomes more detailed, resulting in clusters of [i] Stucco work on architectural elements and sculptures, [ii] General stucco designs, [iii] Distant views of stupas showing both landscape and the stupa, and [iv] Stupas with 75% height visible in close-up, without the pointed top. At K = 10, the algorithm further divides the data into 10 distinct categories, including [i] Niches, [ii] General stucco, [iii] Stucco in niches alongside indented stupas, [iv] General stucco patterns decorating stupas, [v] Pointed stupas, [vi] Distant stupa views showing surrounding landscapes and full stupas, [vii] General stucco, [viii] Various stupa styles, [ix] Pointed stupas, and [x] General stucco.

Figure 5 shows how RestNet50 clusters the images at different values of K. For K = 3, three broad groupings emerge, namely [i] Full and close-up views of stupas without nearby elements, [ii] Decorative stucco on stupas and sculptures in niches, and [iii] Wide-angle views of stupas, some with extensive landscapes around, and others showing the entire stupa. When K = 4 is used, these groups split into more specific clusters, including of [i] General stucco decorations on stupas, [ii] Wide-angle views of stupas in full and close-up, [iii] Decorative stucco on stupas and low- and high-relief sculptures, and [iv] Wide-angle views of stupas with extensive landscape and full views. By K = 10, the clustering becomes even more granular, producing 10 categories of [i] General stucco decorations on stupas, [ii] General stucco decorations on stupas, [iii] Wide-angle views of stupas with surrounding landscapes and full views, [iv] General stucco decorations on stupas, [v] Stucco in niches, possibly with Buddha statues, [vi] Overall and close-up views of stupas in various styles, [vii] Stupas with indented corners and flat-topped stupas, such as prang and star-fruit styles, [viii] Close-up views of various stupas, [ix] Wide-angle views of stupas with sharp-pointed tops, and [x] Decorative stucco on stupas and sculptures in or plain niches.

In Fig. 6, the clustering results of DenseNet are illustrated. Under K = 3, the images can be classified into [i] Stucco ornamentation on stupas, [ii] Overall views of stupas, mostly in wide angles, and [iii] General stupa structure with no wood carvings or paintings. Increasing the cluster count to K = 4 leads to more precise separations, highlighting [i] Stupas that are slender and pointed, or bell-shaped with sharp proportions, [ii] Overview of various types of stupas and their elements, such as stucco of elephants around stupas, [iii] Stucco decorations on stupas, and [iv] Stupas in prang (tower) or temple style with multiple facets or angles. At K = 10, the clustering process differentiates the dataset into 10 distinct groups, including [i] Pointed stupas with or without a tiered base; the lower part may or may not have a bell, [ii] Stucco decoration on stupas with indented corners (mostly tiered bases), [iii] Full-body sculptures on walls or in niches or plain niches between decorative columns, [iv] Bell-shaped stupas and bell shapes atop stupa towers; some images show middle or base parts, with some having a star fruit-like top, [v] Wide-angle views of full stupas showing the surrounding landscape or temple setting, [vi] Stupas in prang style, star-fruit style, or closed-lotus-bud style (no bell shapes), [vii] General stucco decorations on stupas, [viii] Niches with or possibly with Buddha statues inside, [ix] General stucco decorations on stupas, and [x] Various types of stupas with pointed tops.

Figure 7 depicts the outcomes of InceptionV3 clustering across different values of K. With K = 3, the results show groupings of [i] Stupas of all types with both close-up and distant views, [ii] Wide-angle views of various stupas, and [iii] Various stucco reliefs and decorative low-relief and high-relief sculptures on stupas. When the cluster number is increased to K = 4, the division of images becomes more fine-grained, yielding [i] Wide-angle views of various stupas, [ii] Various stucco reliefs and decorative low-relief and high-relief sculptures on stupas, [iii] Wide-angle views of pointed stupas, and [iv] Overview of various types of stupas (temple towers, bell-shaped, prang style) with stucco decoration. At K = 10, the method captures even subtler distinctions of [i] Stucco decorations on stupas, [ii] Tall and slender stupas (full view), [iii] Stupa corners with indentations in different types of stupas, [iv] Niches (stucco) decorating various types of stupas, [v] Stupas in stocky forms, [vi] General stucco patterns decorating stupas, [vii] General stucco patterns decorating stupas, [viii] Wide-angle views of stupas that emphasize the surrounding landscape, [ix] Wide-angle views of various stupas, showing them more closely, and [x] General stucco decorations on stupas.

The results suggest that the combination of pre-trained CNN models and unsupervised clustering can be a valuable tool for the expert’s initial screening process. After classification, images are organized into directories by cluster, and the expert reviews each group to ensure accurate categorization.

Figure 8 presents the expert evaluation of image clusters for each K value (K = 3, 4, and 10) across four CNN models. For instance, K = 3 groups images into three clusters (labeled 1 to 3), while K = 10 organizes them into ten clusters (labeled 1 to 10). Within each cluster, the expert identified the dominant theme or subject matter; for example, with ResNet50 at K = 3, cluster 2 was characterized by "wide-angle views of stupas, some with extensive landscapes, and others displaying the entire stupa." The expert systematically reviewed each cluster, rejecting images that did not match the identified theme. The figure shows the number and percentage of rejected images per cluster, identifying statistics such as minimum, maximum, and average rejection percentages for each K value.

This evaluation provides insight into the clustering model’s effectiveness in aligning thematically consistent images with expert judgment. Results indicate that the unsupervised pre-trained embedding method achieved strong thematic consistency, evidenced by low rejection rates in most clusters. Among the models, ResNet50 achieved the lowest rejection rate at K = 10. At K = 3 and K = 4, its rejection rates were slightly higher than those of other models, except for InceptionV3, which consistently exhibited the highest rejection rates across all K values. Across all models, smaller K values corresponded with lower average rejection rates, while larger K values increased misalignment. Although InceptionV3 did not follow this trend, the overall pattern suggests that further fine-tuning of pre-trained CNN models is necessary for optimal performance in art and architectural classification tasks.

6 Conclusion and discussion

This study has demonstrated the efficacy of unsupervised clustering models―specifically VGG19, ResNet50, DenseNet, and InceptionV3―in revealing meaningful architectural patterns within a corpus of Thai stupa imagery, with validation provided by a senior historian/archaeologist. By structuring cluster coherence according to view type, stupa form, ornamentation, and decorative motifs (see Fig. 3, “Key Elements”), our approach not only confirmed the conceptual consistency of machine-generated groupings but also facilitated the discovery of refined stylistic categories that had previously gone unrecognized.

A particularly influential recent work is Li (2025), who reports over 93% average classification accuracy using a supervised, modified CNN on a curated dataset of 5000 artworks, outperforming both ResNet50 and VGG16 while also employing tools like t-SNE, PCA, and Grad-CAM to unpack learned style features. These findings resonate strongly with the present research: while Li’s work is grounded in supervised learning with curated datasets, our unsupervised approach highlights how meaningful structure can also emerge in data with incomplete or noisy labels. Taken together, the two studies suggest a promising trajectory―hybrid methodologies that begin with unsupervised grouping to surface latent patterns and then fine-tune supervised models for improved accuracy and interpretability.

6.1 Addressing labeling challenges through unsupervised clustering

This study investigates the use of unsupervised clustering, powered by CNN-based feature extraction and K-means grouping, as a hybrid human-AI strategy to support the classification of photographic datasets in the domain of traditional Thai art and architecture. Through a focused analysis of 824 images labeled as “stupa,” the research demonstrates how this approach can effectively mitigate three persistent challenges in expert-driven tagging: incompleteness, label noise, and subjectivity.

6.1.1 Minimizing incomplete tagging through reviewing clusters

Incomplete labeling, resulting from the time-consuming and fragmented nature of manual annotation, has long posed a barrier to reliable supervised learning in the humanities. In our case, the expert who tagged the dataset over a decade did so in segmented sessions organized by geographic region (e.g., province), which prevented holistic cross-examination. As a result, many significant visual features remained under-tagged or entirely unrecognized.

Unsupervised clustering addresses this limitation by grouping images based solely on learned visual similarity, independent of pre-existing tags. By embedding the images using four pre-trained CNN models and applying K-means clustering, this study surfaced previously overlooked thematic consistencies across the image set. The expert’s evaluation of clusters, particularly at lower K values (K = 3 and 4), revealed broad visual coherence that surpassed the granularity of the original manual labels. For example, ResNet50 clusters aggregated full-view stupas across different provinces, enabling the expert to perceive patterns that were not evident during the original tagging process. This suggests that clustering can aid in completing visual categorization where manual efforts fall short.

6.1.2 Identifying and filtering noisy labels through visual verification

Label noise in this context refers to incorrect or misleading annotations that do not accurately represent the visual content. This was particularly evident in cases where the tag “stupa” was assigned to images of painted murals or carved wooden thrones that featured stupa-like motifs but were not actual architectural structures.

The unsupervised clustering process helped reveal these inconsistencies by positioning visually anomalous images within unexpected clusters. Upon reviewing the groupings, the expert identified and rejected such misclassified items, demonstrating that the algorithmic clustering provided a critical mechanism for quality control. In effect, the AI system served not just as a tool for categorization, but as a mirror that enabled the expert to re-evaluate and refine earlier decisions. This iterative feedback loop minimizes the propagation of noisy labels and enhances the semantic clarity of the dataset.

6.1.3 Mitigating subjectivity through pattern emergence

Subjectivity in expert tagging arises from the interpretive nature of cultural heritage classification, where decisions may vary depending on disciplinary focus, regional expertise, or personal framing. In our study, for instance, some stupas were labeled based on stylistic impression (“Lanna” vs. “Chiang Mai style”) rather than strict architectural criteria.

By abstracting visual relationships through machine learning, unsupervised clustering introduces a complementary perspective to human interpretation. The expert noted that certain clusters, particularly at higher K values (e.g., K = 10), allowed for more nuanced distinctions between subtypes of stupas, such as star-fruit shaped tops or decorative stucco in niches. These machine-generated groupings encouraged the expert to adopt a more consistent evaluative lens, thereby reducing interpretive drift. The hybrid model thus positions AI not as a replacement for scholarly judgment, but as a tool to scaffold more structured and replicable interpretations.

6.2 Methodological limitations

While the study demonstrates the promise of combining CNN-based unsupervised clustering with expert validation for cultural heritage image classification, several methodological limitations warrant consideration. First, the clustering pipeline relies solely on visual features extracted from pre-trained CNNs without domain-specific fine-tuning. Although this approach enhances generalizability, it may limit the model’s sensitivity to stylistic subtleties specific to Thai art and architecture. Second, the K-means algorithm presumes spherical clusters and equal variance among groups, assumptions that may not hold for complex, high-dimensional visual data, potentially affecting cluster granularity and interpretability.

Third, the qualitative evaluation was performed by a single expert―the same individual who originally annotated the dataset. While this ensures continuity and contextual depth, it introduces potential bias and limits the intersubjective validation of clustering results. Future studies could benefit from incorporating multiple expert evaluations to enhance reliability and triangulate interpretive findings.

Lastly, the study focused exclusively on images labeled as stupa, which, while culturally significant, represents only a subset of the broader dataset. This scope limitation constrains the immediate generalizability of findings to other architectural or artistic motifs. Further validation across additional typologies and larger, more diverse image sets is necessary to confirm the robustness and transferability of the proposed method.

6.3 Generalizability of the method

While this study focuses on a specific dataset of Thai traditional art and architecture, the methodological framework it proposes―unsupervised image clustering using pre-trained Convolutional Neural Networks (CNNs) combined with expert validation―is broadly applicable across other domains of cultural heritage and beyond. The use of feature embeddings extracted from well-established CNN models (e.g., ResNet50, VGG19, DenseNet, InceptionV3) enables transferability to a wide range of image-based datasets, even in contexts where labeled data are limited, noisy, or inconsistent.

The generalizability of this approach lies in its hybrid structure. On the one hand, the unsupervised clustering process is model-agnostic and does not rely on domain-specific fine-tuning, making it suitable for initial explorations in under-annotated image corpora. On the other hand, the involvement of a human expert in interpreting the clusters ensures cultural and contextual relevance, a key consideration in heritage informatics. This collaborative loop between algorithm and domain expertise can be replicated in various disciplines, including archaeology, art history, museum studies, and visual media archiving.

6.4 Future directions and practical applications

This hybrid methodology, combining CNN-based clustering with expert review, demonstrates a scalable framework for enhancing visual classification in the digital humanities. Based on expert feedback, future implementations should include clustering tools that allow for (1) differentiation of architectural styles (e.g., Sukhothai vs. Lanna stupas), (2) detection of representational artifacts (e.g., murals vs. structures), and (3) categorization by functional typologies (e.g., pagodas, ordination halls). These requirements underscore the curatorial nature of the task, with direct implications for digital archiving and academic resource development.

Looking forward, we aim to expand this work toward automatic stylistic detection via supervised learning. Differentiating stylistic lineages such as Sukhothai and Lanna will demand richer feature sets and more robust training data, introducing new challenges related to model interpretability and dataset fidelity. Nonetheless, the insights gained from this unsupervised phase provide a strong foundation for that next step. This aligns with prior efforts such as Li et al. (2021), who demonstrate how deep learning enables classification across architectural styles using CNNs trained on curated imagery.

Interpretability remains a major concern in AI-assisted workflows, particularly in cultural domains where symbolic content matters (Chen et al., 2022). By bridging algorithmic clustering with human expertise, this study offers a replicable model for cultural heritage contexts worldwide, where labeled data are incomplete, inconsistent, or interpretively fluid, and advances the broader mission of building more inclusive and intelligent digital archives.

References

[1]

Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Zheng, X., 2016. TensorFlow: a system for large-scale machine learning. In: Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), pp. 265—283.

[2]

Alzubaidi, L., Al-Sabaawi, A., Ibrahim, H.M., Mohsin, Z., Al-Amidie, M., 2020. Amended convolutional neural network with global average pooling for image classification. Adv. Intell. Syst. Comput. 1194, 1—10.

[3]

Bar, Y., Levy, N., Wolf, L., 2015. Classification of artistic styles using binarized features derived from a deep neural network. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer, pp. 71—84.

[4]

Belhi, A., Bouras, A., Foufou, S., 2018. Automatic annotation of cultural heritage images using a multimodal classification approach. Appl. Sci. 8 (10), 1768.

[5]

Caron, M., Bojanowski, P., Joulin, A., Douze, M., 2018. Deep clustering for unsupervised learning of visual features. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer, pp. 132—149.

[6]

Castellano, G., Vessio, G., 2022. DELIUS: a deep learning pipeline for unsupervised clustering of digitized paintings. Multimed. Tool. Appl. 81 (19), 27849—27876.

[7]

Cetinic, E., Lipic, T., Grgic, S., 2018. Fine-tuning convolutional neural networks for fine art classification. Expert Syst. Appl. 114, 107—118.

[8]

Chanthawilasawong, S. (Ed.), 2010a. Temples in Ayutthaya. Visual Archive and Rattanakosin Exhibition Hall Project. Bangkok (in Thai).

[9]

Chanthawilasawong, S. (Ed.), 2010b. Buddhist Art of the South. Visual Archive and Rattanakosin Exhibition Hall Project. Bangkok (in Thai).

[10]

Chanthawilasawong, S. (Ed.), 2010c. Magnificent Lanna. Visual Archive and Rattanakosin Exhibition Hall Project. Bangkok (in Thai).

[11]

Chanthawilasawong, S. (Ed.), 2010d. Sukhothai Craftsmanship. Visual Archive and Rattanakosin Exhibition Hall Project. Bangkok (in Thai).

[12]

Chanthawilasawong, S. (Ed.), 2010e. Art Beyond the Old Capital. Visual Archive and Rattanakosin Exhibition Hall Project. Bangkok (in Thai).

[13]

Chen, C.-C., Su, Y.-H., Tsai, D.-M., Wang, Y.-C.F., 2021. A review on deep learning in artistic image analysis. ACM Comput. Surv. 54 (8), 1—29.

[14]

Chen, L., Wu, Z., Wang, M., 2022. Bridging algorithmic design and cultural narrative: interpretability in machine-assisted architectural workflows. Front. Architect. Res. 11 (1), 104—118.

[15]

Foka, A., Griffin, G., Ortiz Pablo, D., Rajkowska, P., 2025a. Reimagining digital cultural heritage: critical perspectives on bias and interpretation in expert annotations. J. Digit. Human. 14 (1), 33—48.

[16]

Foka, A., Griffin, G., Ortiz Pablo, D., Rajkowska, P., Badri, S., 2025. Tracing the bias loop: AI, cultural heritage and bias-mitigating in practice. AI Soc 1—13.

[17]

Goldfinch, L., Chang, C., Silpachai, R., 2025. Machine learning for the recognition of prehistoric rock art in Southeast Asia. SEAMEO SPAFA J. 6 (1), 1—17.

[18]

Gultepe, B., Ozturk, A., Cetin, M., 2018. Unsupervised feature learning for paintings style classification. Signal Image Video Proc. 12 (6), 1125—1132.

[19]

Gupta, V., Naaz, S., 2023. Interpretation of art and architecture. ShodhKosh: J. Vis. Perform. Arts 4 (2), 42—59.

[20]

He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770—778.

[21]

Huang, G., Liu, Z., van der Maaten, L., Weinberger, K.Q., 2017. Densely connected convolutional networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4700—4708.

[22]

Karayev, S., Hertzmann, A., Winnemoeller, H., Darrell, T., 2014. Recognizing image style. In: Proceedings of the British Machine Vision Conference (BMVC).

[23]

Kondo, K., Hasegawa, T., 2020. Artist classification using convolutional neural networks. J. Cult. Herit. 41, 212—220.

[24]

Kuntitan, P., Chaowalit, O., 2022. Using deep learning for the image recognition of motifs on the center of Sukhothai ceramics. Curr. Appl. Sci. Technol. 22 (2), 1—15.

[25]

Li, J., Sun, X., Xu, Y., 2021. Image-based classification of architectural styles using deep learning techniques. Front. Architect. Res. 10 (2), 272—286.

[26]

Li, W., 2025. Enhanced automated art curation using supervised modified CNN for art style classification. Sci. Rep. 15, 7319.

[27]

Llamas, J., Reinoso, D., Rodríguez, A., 2017. Deep learning techniques for the classification of architectural heritage photo-graphs. Appl. Sci. 7 (10), 992.

[28]

Ottoni, A.L.C., Ottoni, L.T.C., 2024a. ImageOP: the image dataset with religious buildings in the world heritage town of Ouro Preto for deep learning classification. Heritage 7 (11), 6499—6525.

[29]

Ottoni, A.L.C., Ottoni, L.T.C., 2024b. Labeling cultural heritage data: between accuracy and interpretation. J. Heritage Inform. 13 (2), 121—137.

[30]

Parisotto, M., Campanaro, D.M., Giuffrida, M.V., 2022. Unsupervised clustering of archaeological images using deep generative models. J. Archaeol. Sci.: Report 42, 103362.

[31]

Reshetnikov, A., Marinescu, M.-C., More Lopez, J., 2023. DEArt: dataset of European art. In: Karlinsky, L., Michaeli, T., Nishino, K. (Eds.), Computer Vision - ECCV 2022 Workshops, Proceedings. Springer, pp. 218—233.

[32]

Simonyan, K., Zisserman, A., 2015. Very deep convolutional networks for large-scale image recognition. In: 3rd International Conference on Learning Representations (ICLR 2015), Conference Track Proceedings.

[33]

Song, H., Kim, M., Park, D., Shin, Y., Lee, J.-G., 2022. Learning from noisy labels with deep neural networks: a survey. IEEE Transact. Neural Networks Learn. Syst. 33 (3), 1175—1195.

[34]

Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z., 2016. Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2818—2826.

[35]

Tan, M., Wang, L., Zhang, Y., 2016. Art classification using convolutional neural networks. Neurocomputing 210, 1—9.

[36]

Tan, M., Zhang, Y., Zhao, Q., 2018. Domain-adversarial training for sketch-based image retrieval. IEEE Trans. Image Process. 27 (9), 4379—4392.

[37]

Wang, H., 2024. Automatic classification of museum artifacts based on unsupervised models. J. Cult. Herit. Inform. 11 (1), 55—67.

[38]

West, M.C., 2017. A Review of Subject Indexing and Social Tagging Projects in Art Museums. University of North Carolina at Chapel Hill. Master's paper.

[39]

Xu, K., Lee, Y.-J., Zitnick, C.L., 2014. 3D scene labeling using voxel-based contextual features. IEEE Trans. Pattern Anal. Mach. Intell. 37 (10), 2027—2041.

[40]

Zhang, Z., Cao, Y., 2025. Annotating degraded mural images from the Mogao Caves: toward deep learning applications in cultural restoration. Dig. Scholar. Human. 40 (1), 76—95.

Rights & permissions

2095-2635/2025 The Authors. Publishing services by Elsevier B.V. on behalf of KeAi Communications Co. Ltd.

PDF (10277KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/

〈 〉