1 Introduction
Murals, as an ancient form of artistic expression, have adorned walls of architectural structures since the dawn of human civilization, narrating historical chapters through their unique artistic language. Chinese mural art, in particular, is rich and varied, having accumulated countless precious works over millennia (
He et al., 2014). These murals are not only vital carriers of Chinese culture but also invaluable resources for studying ancient social customs, ideological concepts, religious beliefs, architectural styles, and folk traditions (
Izzo et al., 2015;
M. Li et al., 2022). Being directly integrated into the built environment―whether in temples, palaces, tombs, or cave dwellings―these murals exhibit a unique architectural artistry and serve as a critical component of historical architectural spaces. Scholars engage in mural studies not only to appreciate their aesthetic and historical value but also to uncover the profound cultural and symbolic meanings embedded within these artworks. By analyzing the composition, colors, and thematic elements, researchers gain valuable insights into the aesthetic concepts, lifestyles, and spiritual worlds of ancient societies.
However, with the passage of time, natural erosion of the attached building, biological deterioration, and human destruction have inflicted severe damage upon these fragile artistic, resulting in cracks, peeling, and mold formation (
Alonso-Villar et al., 2023). These damages may break the integrity of the mural, not only reducing visual perception, but also causing loss of useful information and difficulty for further research. Traditional restoration techniques require extensive expertise, are highly time-consuming, and often impose irreversible impacts on the original artwork. Therefore, utilizing computer technology for digital restoration of murals is gradually becoming a new trend in the protection and restoration of ancient art (
Purkait and Chanda, 2012). Digital restoration techniques offer an efficient and non-invasive means to reconstruct the original appearance of murals while preserving their historical and cultural significance, which is of great significance for the protection of cultural heritage (
Shao et al., 2023).
Digital image restoration has long been an active area of research within computational science, where various methodologies have been proposed. Pioneering works such as (
Ballester et al., 2001;
Elharrouss et al., 2020;
Jam et al., 2021;
Qiang et al., 2017) primarily relied on diffusion techniques to propagate information from neighboring regions to fill in missing areas.
Barnes et al. (2009) advanced this approach with their random nearest neighbor patch matching technique, improving the repair process. Despite these advancements, patch-based methods are inherently limited by the similarity and complexity of the surrounding content.
Bertalmio et al. (2000) anisotropic diffusion improved clarity but was limited to small repairs. With the development of deep learning on computer vision, contemporary deep learning-based methods have gained prominence due to their end-to-end training capabilities and prowess in capturing image structure and semantics. These include progressive restoration, structure-guided, attention-driven, convolutional, and diversified restoration methods. Progressively, approaches like (
Yu et al., 2018;
Zhu et al., 2021) refine damaged regions from rough to fine. Structure guidance (
Yang et al., 2020) uses known structures as guides, while attention-based (
Wang et al., 2020;
Xie et al., 2019;
Yu et al., 2018) leverages human-like attention for efficiency and quality. Convolution-aware restoration (
Yan and Zhang, 2022) marks missing areas with masks, integrating learnable attention modules. Diversified restoration (
Wan et al., 2021;
Zheng et al., 2019) yields visually plausible outcomes.
Compared to conventional image restoration tasks, which are primarily designed for portraits or landscapes, mural restoration presents a set of unique challenges. First, murals often exist in dimly lit environments, such as tombs, caves, or temples, where strict conservation measures limit lighting conditions (
Cao et al., 2021). This results in low-light poor-quality images, posing significant challenges for restoration work. Secondly, murals possess distinct color palettes and textures, making the direct application of pre-trained models from other domains largely ineffective. Building restoration networks specifically tailored for murals requires a large amount of annotated mural data, which are scarce and difficult to obtain. Lastly, the degradation patterns in murals are highly heterogeneous, with varying degrees of cracks, fading, and biological damage across different regions. This non-uniformity necessitates adaptive restoration strategies that can dynamically adjust to localized damage, posing an additional challenge for automated restoration techniques.
Early research on mural restoration primarily relied on traditional image inpainting techniques. For example, various studies (
Chen et al., 2021;
Ren, 2024) have sought to improve the Criminisi algorithm to mitigate block artifacts in traditional inpainting approaches.
Purkait and Chanda (2012) proposed a patch-based anisotropic diffusion technique to enhance lines and brush strokes, while
Giakoumis and Pitas (1998) developed a semi-automatic algorithm focusing on the restoration of cracks in mural paintings. These approaches improved specific aspects of mural repair but were often limited in handling large-scale or heterogeneous degradations.
More recently, machine learning—based methods tailored to mural restoration have emerged.
Song et al. (2020) introduced AGAN, an adversarial generative model that adapts to mural image domains for improved restoration realism.
Zhang et al. (2023) developed SPN, a semantic prior-based network that integrates contextual cues to enhance inpainting performance.
L. Li et al. (2022) proposed a Mural Line Drawing—guided approach, leveraging structural information to preserve line integrity during restoration. While these methods mark significant progress in mural-specific restoration, they mainly focus on structural reconstruction or texture recovery and have not explicitly addressed the challenges of restoration under low-light conditions.
Several efforts have been taken to solve the problems. Existing research efforts have primarily focused on addressing individual challenges in mural restoration, such as low-light enhancement or structural reconstruction, rather than developing an integrated framework that simultaneously considers multiple restoration aspects. On the one hand, progress has been made in low-light image enhancement. Traditional techniques such as histogram equalization (
Arici et al., 2009) adjust brightness and contrast but often amplify noise and fail in complex image settings. Deep learning models (
Lore et al., 2017;
Tao et al., 2017;
Zhang et al., 2019;
Wang et al., 2022) have overcome these limitations through extensive training, offering better generalization and adaptability, handling various image enhancement tasks efficiently. Notable works include Lore et al. (2017), who introduced deep frameworks for low-light enhancement, and
Zhang et al. (2019), who refined enhancement techniques inspired by Retinex theory.
Wang et al. (2019) proposed end-to-end networks with intermediate illumination, and
Yang et al. (2020) utilized recursive networks with adversarial learning for low-light recovery. These enhancements are focused on optimizing visibility under dim-light conditions for mural imaging, thus contributing to overcoming a key challenge in their digital restoration. More recently, Ma et al. (2022) developed a self-calibrated illumination learning framework that enhances image quality efficiently across various scenarios. These advancements have significantly improved the visibility of murals in dimly lit conditions, forming a crucial foundation for digital restoration. On the other hand, research dedicated specifically to mural restoration has primarily focused on structural reconstruction rather than integrating low-light enhancement with restoration.
Thus, developing a comprehensive restoration framework that simultaneously addresses both low-light degradation and mural reconstruction remains an open challenge, and bridging the gap between these two research directions could pave the way for more effective and holistic mural restoration strategies, ensuring the preservation of both visual integrity and historical information with greater accuracy. To tackle this challenge, our work proposes an end-to-end neural network framework designed for the meticulous restoration of low-light murals, particularly in the complex scenario of batch restoration under dim lighting conditions, where conventional methods often struggle. We introduce a two-stage framework, the Multi-level Interactive Siamese Filtering and Restoration Network (MISFR), which is explicitly designed to overcome the unique difficulties associated with mural restoration in poor lighting environments. At the core of our approach lies a dual-phase methodology tailored for murals captured under low-light conditions, as these images frequently suffer from a scarcity of pixel information, making it challenging to extract fine-grained textural and structural details―factors crucial for both neural network comprehension and high-quality restoration. By leveraging multi-level interactive filtering and targeted restoration strategies, our framework is capable of producing enhanced mural reconstructions that not only restore structural integrity but also retain the intricate artistic details essential for historical preservation.
Our neural architecture is structures into primary stages: initial low-light image enhancement and followed by restoration. In the first stage, we employ an adaptive, weight-shared illumination adjustment module that augments the clarity of low-light images, enhancing visibility of texture and structure comprehension for the subsequent restoration process. Once enhanced, the second stage extracts intricate textural and structural features through a unified network that performs multi-layered predictions, simultaneously influencing both the surface image and the deep feature representation. While the surface-level processing focuses on recovering fine details, the deeper layers capture and enrich semantic information, ultimately leading to a restoration of high authenticity. This dual-level approach ensures a perceptually enhanced output that not only refines textures but also preserves structural integrity, producing mural restorations with exceptional fidelity to their original form. To further validate the effectiveness of MISFR, we conducted extensive comparisons across multiple datasets against various existing methods, with the promising results demonstrating its superior performance and robustness in mural restoration.
Summarily, this study’s main contributions encompass the following:
● Inception of MISFR network, a first two-staged restorative framework proficient at countering restoration obstacles in low-light measurement conditions, elevating antique mural restoration standards significantly.
● Deployment of efficacious strategies addressing low-light scenarios and segmentation, amplifying dataset size and optimizing both training outcomes and the restoration efficacy.
● Harmonious integration of prediction at varied depths, nurturing detailed retrieval concurrent with upholding a high level of textural integrity, exemplifying marked improvements in handling murals under minimal light.
2 Materials and methods
2.1 Dataset
This study selects representative and renowned Chinese mural resources from renowned sites, including Hongdong Guangsheng Temple, Fanshi Princess Temple, and Ruicheng Yongle Palace. Given the indoor placement of murals in cave dwellings, tombs, temples, and palaces, with limited lighting to preserve them, many were recorded under poor illumination, making the original images suboptimal for direct detailed observation. To faithfully replicate the prevalent issue of low brightness observed in actual data collection scenarios, we adjust the collected images’ brightness to 55%, 37%, and 12% of their original levels, creating multiple subsets under varying illuminations. To augment diversity and practicality, these dimmed images are further cropped to uniform 256 × 256 pixel tiles (Fig. 1), facilitating subsequent processing and training. Ultimately, our dataset comprises 123,298 low-light images for training and 1527 images applied on three different bright levels for testing, meeting the demands for model training and performance validation.
2.2 Low light enhancement stage
Given the distinctive nature of murals, these artistic masterpieces are commonly discovered and preserved within dimly lit interiors such as rock caves, ancient tombs, and sacred temples. The scarcity of natural light in these settings significantly compromises our visual perception of the murals, particularly when capturing their images, often resulting in severely low-light conditions. This not only reduces overall image visibility but also makes it challenging to discern the invaluable texture details even after restoration, posing significant hurdles for cultural heritage research and preservation.
To tackle this issue, we employ a specialized deep learning network designed for low-light image processing to enhance murals captured under such conditions. This network aims to enhance image quality in low-light environments, revealing the intricate details concealed in shadows, thereby offering clearer representations of the murals’ original textures. This, in turn, furnishes more precise visual information for conservation and restoration efforts, while also enabling researchers to gain a more comprehensive understanding, diving deeper into the stories told by these historical witnesses.
In the classic Retinex theory framework (
Land and McCann, 1971), the relationship between the low-light observation
y and the ideal clear image
z can be expressed as
y =
z⊗
x, where
x denotes the lighting component, which is the pivotal target for optimization in low-light image enhancement. Retinex theory posits that accurate estimation of lighting conditions significantly improves image quality. However, considering the high correlation, potentially linear, between the lighting conditions in mural settings and the observed low-light images, along with limited available datasets, we draw inspiration from (
Ma et al., 2022) to construct a light estimation network
Hθ with shared structure and parameters, where
θ represents learnable parameters. To maintain consistency in the output across multiple optimization rounds, we introduce a self-calibration module
G, built using a network
Kϑ with learnable parameters
ϑ. This self-calibration mechanism ensures that the improvements during progressive optimization are stable and coherent. Consequently, the mathematical formulation of our progressive optimization network is as follows:
where k = t + 1, vt is the transformed input in round t of the asymptotic process, ut is the residual in round t, and xt is the illumination in round t (t = 0, 1, …).
Further, in order to better train the low-light image enhancement network, we refer to the work (
Land and McCann, 1971;
X. Li et al., 2022) which use an unsupervised loss function
LE defined as
where T is the total number of optimization enhancement rounds, which we set to 6 in our experiments. N is the total number of pixels, st-1 is the output of Kϑ, i is the i th pixel, c is the image channel in the YUV color space and σ=0:1 is the standard deviation of the Gaussian kernel. α and ß are the weights which we set to 1.5 and 1 in our experiments.
2.3 Image inpainting stage
The final output from the low-light image enhancement stage serves as the input
Iin for the image restoration network, tasked with restoring damaged areas. Inspired by the work, we adopt a Multi-level Interactive Siamese Filtering (MISF) framework (
X. Li et al., 2022), aiming at achieving high-fidelity mural restoration results across diverse scenes and different sizes.
Since semantic information can be preserved even a large area of the image is lost. Therefore, we extend the filtering from the image level to the deep feature level that contains semantic information. In order to roughly recover the structure of the broken mural, we first input the mural I in which is to be restored with a binary mask M describing the missing regions.
We first employ an encoder-decoder network (
Mao et al., 2016) where the encoder is to extract features of the corrupted image (i.e.,
I) and the decoder is to map the features to the completed image. We have the following formulation for the encoder.
where ϕ(·) is the encoder and FL is the deep feature extracted from the lth layer, i.e., FL = ϕl(FL-1). For example, FL is the output of the last layer of ϕ(·) (i.e., ϕl(·)). The decoder can be formulated as
where ϕ-1(·) is the decoder. Then, we conduct the semantic filtering on extracted features like the image-level filtering
where is the kernel for filtering the pth element of Fl via the neighboring elements, i.e., Np. We use the matrix Kl to include all element-wise kernels (i.e., ). After that, we replace the Fl with in Eq. (4) and conduct the subsequent operations. To let the kernels adapt to different scenes, we also employ a predictive network to predict the kernels like the image-level predictive filtering (i.e., Eq. (3))
where ϕl(·) is the predictive network to produce Kl.
Nevertheless, Semantic filtering fills the missing semantic information at the deep feature level that has a low spatial resolution. Thus, it inevitably loses detailed information. To solve this problem, multi-level interactive Siamese filtering (MISF) that consists of two branches is proposed, which are encoder-decoder networks containing several convolutional blocks, i.e., kernel prediction branch (KPB) and semantic & image filtering branch (SIFB). Kernel Prediction Branch (KPB) is an encoder-decoder network that takes the input corrupted image and features from the Semantic & Image Filtering Branch (SIFB) to predict dynamic convolutional kernels for SIFB at multiple levels. While Semantic & Image Filtering Branch is another encoder-decoder network that performs filtering at both the semantic feature level and image level using the kernels predicted by KPB, allowing for filling of large missing regions while preserving details. The two branches are interactively linked together, performing a cross-attention architecture, with SIFB providing multi-level features for KPB, and KPB predicting dynamic kernels for SIFB’s filtering operations.
Specifically, given a corrupted image I, we feed it to the SIFB that conducts filtering at the image level and semantic level (i.e., filtering at the lth-layer feature) jointly. In other words, SIFB applies initial filters to predict the basic structure of the missing parts. KPB takes the output from SIFB and the original input I to predict dynamic kernels. These kernels are tailored to the specific needs of the damaged mural based on its contextual and textural information. As a result, we can generate the completed image by
here, F1 = ϕL(…ϕL(1)). The kernels for deep feature and image (i.e., Kl and K) are predicted by the KPB
where Fj=ϕj(…ϕj(I)) is the features from the j th layer of SIFB, and Ej=ϕj(…ϕj(I)) is from the j th layer of KPB. We add a convolutional layer (Conv(·)) to adjust the size of ϕL(…ϕj+1([Ej, Fj])) to meet the requirements of kernels. The kernels Kl and K are for the feature-level and image-level filtering, respectively. Moreover, with Eq. (9) and Eq. (10), all predicted kernels for semantic & image filtering are driven by the input image I and the deep feature Fj, which contain all available spatial details and the understanding of the whole scene. As a result, both semantic information and detailed pixels can be properly reconstructed.
Using the kernels, SIFB performs detailed inpainting. This step iteratively refines the inpainted regions by switching between the image and feature levels, this dynamic process (
Chen et al., 2020) further improving both the details and the overall semantic integrity of the mural.
To ensure the production of images with both high visual quality and strong semantic coherence, we adopt the strategy outlined in reference (
Nazeri et al., 2019) and employ a quartet of loss functions for network training. These encompass the
L1 Loss for pixel-wise differences, the GAN Loss to promote realism, Style Loss for maintaining aesthetic consistency, and Perceptual Loss to capture high-level feature similarities. Concretely, given a defective image
I, its predicted restoration
Î, and the ideal uncorrupted image
I*, the composite loss function we utilize is defined as follows:
Here, we choose λ1 = 1, λ2 = λ3 = 0.1, and λ4 = 250.
2.4 Overall loss function and training strategy
We adopt a two-stage strategy to restore ancient murals with low luminance into images with improved observability. Initially, the output from the low-light enhancement phase serves as the input to the image restoration phase, thereby creating a coupling between the two stages. During the training process, we employ an optimizer to adjust network parameters based on the loss function, aiming to minimize the discrepancy between the generated image and the target image. However, simultaneous parameter updates in both stages can lead to instability in the inputs to the image restoration phase due to continuous changes in the parameters of the enhancement phase, thereby increasing the learning difficulty in the image restoration stage. Given the differences in network complexity and the volume of learning samples required for each stage, we have opted for an alternating parameter update strategy. The whole architecture is shown in Fig. 2.
Within a single training epoch, we divide the dataset into several subsets, each containing 60 samples. The first 6 samples in each subset are allocated for optimizing the enhancement network, during which we freeze the parameters of the image restoration network to focus exclusively on updating the low-light image enhancement network. The remaining 54 samples in the subset are used for optimizing the image restoration network, where we freeze the parameters of the low-light image enhancement network and solely optimize the parameters of the image restoration network.
Furthermore, the loss function for the overall MISFR, denoted as LMISFR, which is defined as the combination of enhancement part LE and restoration part LR, which shows as follows:
Here, the parameter λR is set to 1 during the optimization phase of the inpainting network and 0 during the enhancement network optimization phase, while λE is set to 0 during the optimization phase of the inpainting network and 1 during the enhancement network optimization phase.
3 Results
This chapter outlines the achievements made using the MISFR model in the field of mural restoration, and meticulously explains the related experimental designs and outcomes. This section delves into the core aspects of the experiments, specifically covering the datasets employed, experimental configurations, and chosen evaluation criteria. Through multifaceted demonstrations of experimental effects and substantial quantitative metric data, the efficacy and superiority of the MISFR model in the mural restoration task are comprehensively validated.
3.1 Experiment setup
3.1.1 Mask data synthesis and strategy
To address the common damage characteristics of ancient murals, such as stains, corrosion, peeling, mold, and cracks, this study designs a detailed mask generation scheme aimed at realistically simulating these damage forms. First, based on the types of damage, irregular masks symbolizing corrosion damage are generated using a random walk algorithm to simulate their random distribution and shapes (
Gatys et al., 2016). Next, we create “jellylike” masks with blurred edges through a corrosion operation to mimic the blurred effects around corroded areas (
Yu et al., 2022). Peeling damage is simulated using randomly distributed droplet-shaped masks to represent scattered mold spots and splatter marks. Additionally, large area masks represented stains and are constructed from large circular shapes, while elongated masks simulated linear stain damage are generated from random straight lines. Cracks and scratches, being common forms of damage, are specifically simulated in the mask design. To cover the diversity of damage in practical applications, the generated masks vary in coverage from 5% to 50%, comprehensively simulating various degrees of damage. These masks, encompassing various sizes and shapes, are integrated into the dataset to train and validate the model’s ability to restore diverse damage scenarios.
3.1.2 Experimental setup
For the experimental configuration, we utilize a high-performance NVIDIA TITAN RTX GPU. For optimization, we employ the Adam optimizer, setting the batch size to 6 for network training. During the initial training phase of the first 100 epochs, the learning rate is set to start at 0.0001. We then employ a linear decay strategy over the subsequent 100 epochs, gradually reducing the learning rate to zero to ensure a more stable and smooth learning process.
3.1.3 Evaluation metric
Within the scope of this experimentation, we employ three assessment metrics to gauge the restoration model’s efficacy by comparing the original images with those processed by our model under low-light conditions. These metrics include:
Peak Signal-to-Noise Ratio (PSNR): This metric quantifies the fidelity of the reconstruction by measuring the ratio between the maximum possible power of a signal and the power of background noise.
Structural Similarity Index (SSIM): It provides a holistic score of image similarity by considering luminance, contrast, and structural elements, presenting an evaluation that aligns with human perception of structural integrity.
Learned Perceptual Image Patch Similarity (LPIPS): As a deep learning-based assessment, this metric approximates human visual similarity, evaluating the perceptual quality of the reconstructed images by comparing patches with their ground truth.
These metrics collectively offer a comprehensive assessment of the model’s capability to recover details, maintain structural integrity, and mimic the appearance of the original scene under low-light conditions, similar to human perception.
3.2 Comparison
In the qualitative evaluation, Figs. 3—5 showcase the visual comparison of our methodology against several advanced algorithms―RFR (
Yu et al., 2018), PEN (
Zeng et al., 2019), MADF (
Zhu et al., 2021), SCI (
Ma et al., 2022), LG (
Quan et al., 2022), AGAN (
Song et al., 2020), SPN (
Zhang et al., 2023), and Mural (
L. Li et al., 2022) across three varied lighting conditions and three different degrees of occlusion. Specifically, the SCI (
Ma et al., 2022) technique predominantly emphasizes on enhancing low-light image brightness, neglecting the restoration capabilities, thereby highlighting the limitations in low-light recovery and indicating an inability to fully restore low-light contents independently. Additionally, color discrepancies in images treated by SCI (
Ma et al., 2022) from true colors significantly dilute the unique chromatic style and artistic features of murals.
Further analysis reveals that PEN (
Zeng et al., 2019) underperforms inadequately in large occluded areas, exhibiting noticeable repair voids and defects, suggesting its limitations in wide-ranging damage scenarios. While LG (
Quan et al., 2022) exhibits superior overall image quality, it still deviates notably from the original in brightness and color accuracy. Meanwhile, though MADF (
Zhu et al., 2021) surpasses PEN (
Zeng et al., 2019) in overall restoration, it lacks precision in fine textural detailing, reflecting constraints in microstructural recovery. Similarly, RFR (
Yu et al., 2018) performs poorly in regions with prominent pixel variations in murals. In comparison, AGAN (
Song et al., 2020) introduces artifacts in restored regions, SPN (
Zhang et al., 2023) fails to recover texture under severe low-light conditions, and the Mural Line Drawing—guided approach Mural (
L. Li et al., 2022) struggles to maintain structural consistency, further highlighting the necessity of our integrated framework.
It is noteworthy that despite LG (
Quan et al., 2022), MADF (
Zhu et al., 2021), PEN (
Zeng et al., 2019), and RFR (
Yu et al., 2018) strive to repair damaged areas, their low-light produced images fail to adequately display rich texture and structure, especially at the minimal 2% light level, constraining their applicability in real mural restoration contexts. Conversely, our MISFR not only successfully recovers damaged areas, but also brightens images, ensuring visual recognition of texture and structure. At high light intensities, our restorations closely approximate the ground truth images, whereas others’ outputs are blurred. In the dimmest lighting, other networks’ repairs are nearly unidentifiable while MISFR still offers clear contents, highlighting its superiority under low-light conditions.
To evaluate the effectiveness of the proposed MISF architecture for the mural restoration task, we conduct experiments on the 1527 test sets and evaluate the performance against all the other competitors. We used three metrics―LPIPS (lower is better), PSNR (higher is better), and SSIM (higher is better)―to assess the difference between the model’s processed results and the ground truth. The MISF architecture demonstrated a remarkable ability to capture the semantic features of the input scenes (Fig. 5), accurately restoring the damaged regions and outperforming other methods. This versatility makes the MISF architecture a highly efficient and effective solution for mural restoration across diverse scene domains.
3.3 Compatibility test
3.3.1 Mural compatibility test
To examine the compatibility and robustness of the proposed MISFR model on mural restoration, we trained the model using data exclusively from Hongdong Guangsheng Temple and Fanshi Princess Temple, and then tested on data from Ruicheng Yongle Palace. The results, presented in Fig. 6, indicate that the restoration capability experiences only a slight decline compared to standard training conditions, despite the stylistic differences among murals. This suggests that the variations between different murals may exert a lesser impact on the network’s performance than the differences observed across various regions within a single mural, particularly those with diverse patterns.
Furthermore, we conducted a cross-style transferability test to evaluate the generalization capability of our model across different mural styles. Specifically, while the previous experiment focused on murals of the same type, this test examines whether a model trained on Chinese mural datasets can be effectively applied to murals from different cultural backgrounds. To this end, we trained the model on our dataset of Chinese murals and tested it on Michelangelo’s mural data. The results, presented in Fig. 7, demonstrate that despite the significant stylistic differences, our model is still able to achieve satisfactory restoration, particularly in handling dark and damaged regions. These findings further validate the effectiveness and robustness of our MISFR model in mural restoration across diverse artistic traditions.
3.3.2 Mask compatibility test
To better assess the model’s ability to restore murals with various types of damage, we applied different types of masks to simulate diverse forms of mural deterioration. Specifically, we introduced three primary types of masks: dust-like masks, composed of dense small speckles, mimicking mold, stains, or weathering-induced erosion; jelly-like masks, featuring irregular, smooth-edged black regions that resemble ink stains, water damage, or biological corrosion; and block-like masks, characterized by larger, more defined occlusions, used to simulate paint peeling, large-area detachment, or deliberate vandalism. The original mural images and their corresponding restorations are presented in Fig. 8. The results demonstrate that our model maintains consistently strong restoration performance across these different types of damage, effectively reconstructing missing details while preserving the integrity of the original artwork.
3.4 MISF computational complexity analysis
We evaluate the complexity of the model by analyzing its time efficiency in processing low-light-intensity images of damaged murals. Initially, we select images with masks of 25%—35% area size at 55% brightness to compare the processing efficiency of various models during the restoration process. The results are visualized in Fig. 9. The MISFR model demonstrates exceptional robustness and consistently high efficiency in processing time. According to the statistics in Table 1, the model requires only 0.008 s average time to complete the restoration of a single image, highlighting its rapid response capability. Therefore, the moderate computational complexity of the MISFR ensures its efficient application in batch processing of image data at archaeological sites.
3.5 Dual-stage processing strategy effectiveness
To systematically demonstrate the superiority of the dual-stage processing strategy employed by the MISFR network over single-stage processing, this study conducts a series of comprehensive ablation experiments. These experiments contrast the effects of solely applying a single stage (restoration alone) with those of the two-stage process (enhancement followed by restoration), as compared with raw image. Based on the experimental data analysis presented in Fig. 9(a) and Table 2, it was observed that under conditions of 55% brightness and masking ranges of 25%—35%, the isolated implementations of image enhancement or restoration fall short of the desired objectives, with all three performance metrics inferior to those achieved through the integrated dual-stage strategy. The results reveals that images processed through the dual-stage approach closely matched the pixel value distribution of actual data, indicating the superiority and authenticity enhancement of the dual-stage strategy in detail recovery.
3.6 Structure ablation
To validate the effectiveness of the MISFR model structure, we conducted ablation experiments on the network model of post-brightening restoration work. For MISF’s two-filtering structure, we propose three control groups: one is that kernel branch does not obtain data Fj from Filtering branch and is labeled as blind kernel; Second, the feature kernel Kl no longer guides the Filtering branch and is marked as no feature kernel. The third is to make the image kernel K no longer conduct the final Filtering with the Filtering branch. The result is shown in Fig. 10, showing that the existing structure plays a significant role in improving the performance of mural restoration.
3.7 External factor robustness analysis
To further evaluate the robustness of MISFR under realistic imaging conditions, we conducted additional experiments on two common external disturbances: chromatic aberration and image noise, which frequently occur in shadow or low-light mural photography. These disturbances were applied to the test data to simulate practical degradation scenarios.
As shown in Fig. 11, MISFR demonstrates consistent robustness against these perturbations. While both chromatic aberration and noise lead to some degradation in quantitative metrics (LPIPS, PSNR, and SSIM), the overall performance of MISFR remains competitive and reliable. This indicates that MISFR can effectively maintain restoration quality even under adverse conditions, highlighting its resilience to external factors that may arise in practical mural imaging scenarios.
3.8 Parameter comparison
To further evaluate the impact of loss function parameters on model performance, we conducted an ablation study by systematically adjusting the values of λ2, λ3, and λ4. Specifically, we scaled each parameter by a factor of 10 (both increasing and decreasing) from the originally determined values to assess whether the chosen magnitudes were appropriate. The experimental results, shown in Fig. 12, indicate that our selected parameters achieve stable and reliable performance, confirming their suitability for the restoration task. From Fig. 12, it can be observed that the model achieves consistently satisfactory performance under the three preset parameter configurations.
3.9 Web based mural restoration platform
To offer a more explicit and user-friendly interface for researchers and related users to utilize our MISFR algorithm, we have designed a web-based platform using the Streamlit framework.
This platform allows users to upload their mural images (multi-file supported), select a brush tool to draw custom masks (brush-based defect simulation), and execute the image restoration process. The result is an enhanced and inpainted mural image. By integrating these functionalities into a streamlined web application, we aim to facilitate the application of our algorithm in a practical and accessible manner for the restoration of mural art (Fig. 13). Furthermore, to ensure broader accessibility and usability, we have released an open-source GitHub repository that includes the complete codebase of the platform, along with detailed documentation and a step-by-step tutorial. This resource enables researchers and conservation practitioners to conveniently deploy and adapt the platform in real-world heritage restoration practices.
3.10 Comparative evaluation and cultural significance
To further illustrate the practical value of MISFR in mural restoration, we provide a comparative summary between MISFR and traditional mural restoration approaches, as shown in Table 3. Unlike conventional solutions, which are often limited to repairing simple damages or specific textures, MISFR demonstrates broader adaptability. In particular, it can simultaneously enhance and restore low-light mural images, adapt to different mural styles, and be readily deployed via the open-source web-based platform introduced in Section 3.9. On one hand, as a technical advancement that improves restoration quality under low-light and degraded conditions; on the other, as a cultural heritage support tool, offering researchers and practitioners an accessible, deployable, and complementary solution that bridges computational methods with conservation workflows.
4 Discussion and conclusion
In the field of ancient mural preservation, these murals are typically large, irregularly shaped, and subject to unpredictable damage. Scholars often use cameras to capture the valuable content of ancient murals for preservation. However, these images are often in low-light measurement and damaged conditions, making it difficult to directly observe the murals’ textures and structures. To address this issue, we introduce MISFR, a Multi-layer Interactive Dual-filtering Enhancement and Restoration network framework with two-stage strategy, designed to tackle the restoration challenges of mural images captured in low-light measurement conditions, providing superior restoration quality. This system adopts a two-step strategy, specifically tailored for low-light mural images. Additionally, to overcome the limitations of large mural sizes and small datasets, we employ artificial low-light processing and segmentation strategies. To enhance the applicability of MISFR, we utilize various masks simulating real-world scenarios. Extensive experimental results have validated the remarkable effectiveness of our proposed method.
While our work primarily focuses on proposing a restoration method to address the defects in low-light captured mural images, there remains significant room for further improvement. One key challenge in practical applications is the automatic identification and extraction of damaged areas, which could enable more targeted and efficient restoration. Future advancements in this direction could enhance the adaptability of the model to real-world scenarios where murals exhibit complex and varied degradation patterns. Additionally, despite its strong restoration capability, MISFR still struggles with reconstructing highly informative missing content, such as letter symbols, which require semantic understanding beyond texture and structure. Beyond these technical considerations, it is important to reflect on the implications for heritage interpretation and authenticity. While digital restoration enhances visibility and provides valuable support for conservation, it cannot always recover the semantic meaning of murals, such as religious iconography or textual inscriptions. This highlights the necessity of combining digital methods with physical protection measures, expert domain knowledge, and multidisciplinary validation, ensuring that restoration outcomes remain faithful to both the artistic and cultural significance of the heritage. Addressing these limitations through improved feature learning, adaptive restoration strategies, and multimodal semantic integration would further elevate MISFR’s effectiveness in diverse mural restoration tasks.
Furthermore, MISFR should be understood as part of a broader conservation workflow rather than a standalone system. Its potential contributions include (i) supporting digital documentation by creating detailed archives of mural states under different conditions; (ii) assisting conservation planning by enabling rapid pre-restoration previews and prioritization of interventions; (iii) facilitating restoration workflows through non-invasive visualization of possible outcomes; and (iv) fostering education and public engagement via open-source tools and web-based platforms. These applications position MISFR not as an autonomous replacement for human expertise, but as a collaborative tool that enhances efficiency, accessibility, and interdisciplinary dialogue in cultural heritage conservation.
Looking ahead, we plan to collaborate with cultural heritage institutions to collect authentic mural datasets under diverse illumination and degradation conditions. Such real-world data will enable further validation and refinement of MISFR, strengthening its applicability in practical conservation scenarios and ensuring that future improvements are grounded in both computational innovation and heritage needs.
2095-2635/2025 The Authors. Publishing services by Elsevier B.V. on behalf of KeAi Communications Co. Ltd.