1 Introduction
Mars, one of Earth’s closest planetary neighbors, has long been a major target of planetary exploration owing to its diverse geological history, climatic evolution, and potential habitability. Since the 1960s, space agencies worldwide have conducted dozens of successful Mars missions (Fig. 1). The exploration paradigm has gradually evolved from flyby and orbital reconnaissance missions to increasingly sophisticated surface investigations involving landers, rovers, and even aerial platforms. Representative missions include NASA’s Mars Reconnaissance Orbiter (MRO), the Curiosity and Perseverance rovers (with the Ingenuity helicopter), ESA’s Mars Express orbiter and ExoMars programs, and CNSA’s Tianwen-1 probe with the Zhurong rover. Collectively, these missions have returned an unprecedented volume of orbital remote sensing and in situ exploration data. The resulting datasets span multiple spatial scales, sensing modalities, and temporal resolutions. They provide an increasingly comprehensive view of the Martian surface, atmosphere, and environmental evolution.
The rapid growth of Mars exploration data has created an urgent need for efficient and automated analysis methods. Traditional manual interpretation is labor-intensive and time-consuming, whereas rule-based image processing approaches often struggle to scale with continuously expanding datasets. As a result, intelligent data analysis has become increasingly important for both scientific investigations and mission operations.
Recent advances in deep learning and computer vision have provided powerful tools for extracting information from large volumes of planetary imagery. Their potential has been demonstrated across a wide range of Mars exploration applications. In orbital remote sensing, deep learning methods have been successfully applied to crater and dune detection (
Jiang et al., 2021;
Rubanenko et al., 2021;
Qi et al., 2024), surface change detection (
Zhang et al., 2024b;
Wan et al., 2026), large-scale landform classification, and geological mapping (
Xia et al., 2023;
Tiwari et al., 2024). For Mars in situ exploration, related studies have focused on rock target classification (
Li et al., 2020), scene semantic segmentation (
Li et al., 2025b), monocular depth estimation (
Ma et al., 2024), and other perception tasks essential for autonomous rover operations. These studies demonstrate that computer vision is becoming a key enabling technology for intelligent Mars exploration. They also provide the foundation for future autonomous perception, scientific target selection, and onboard decision-making systems.
To characterize the recent growth of this interdisciplinary field, we conducted a bibliometric analysis using the Web of Science Core Collection. For overall Mars research, the query “TS = (Mars OR Martian) AND DT = (Article)” returned 62723 articles as of March 31, 2026. For research linking Mars data with artificial intelligence, the query “TS = (Mars OR Martian) AND TS = (deep learning OR AI) AND DT = (Article)” returned 984 articles on the same date. The annual publication and citation trends are summarized in Fig. 2. Although the absolute number of AI-related Mars studies remains much smaller than that of the broader Mars research field, both publication output and citation impact have increased steadily over the past decade. This trend indicates that computer vision and deep learning are becoming increasingly active components of Mars exploration research.
Despite this rapid growth, existing studies remain scattered across different tasks, datasets, and mission scenarios. Several reviews have addressed related topics, but their scopes are still limited. Specifically, DeLatte et al. presented one of the earliest reviews on convolutional neural network applications to crater detection and highlighted the importance of benchmark datasets for algorithm evaluation (
DeLatte et al., 2019a). Tewari et al. focused more specifically on crater detection and useful performance benchmarks on a unified dataset, thereby offering a valuable engineering reference(
Tewari et al., 2023). Kuang et al. systematically analyzed semantic segmentation methods for rover navigation vision, with emphasis on data acquisition, model architecture, and evaluation metrics (
Kuang et al., 2022). Qu et al. further summarized Martian image segmentation from both orbital remote sensing and rover vision perspective (
Qu et al., 2025). However, these reviews mainly focus on specific tasks such as crater detection or segmentation. A unified review covering datasets, multi-task perception, Mars-specific constraints, and emerging foundation-model approaches is still lacking.
To fill this gap, this review provides a cross-task overview of computer vision for Mars exploration from two complementary perspectives: orbital remote sensing and in situ exploration. It covers major perception tasks, including image classification, object detection, semantic segmentation, depth estimation, three-dimensional (3D) reconstruction, and multi-task learning. It also examines how mission-specific constraints shape the development of computer vision methods for Mars applications. These constraints include limited onboard resources, scarce labeled data, domain shifts, and high reliability requirements. Here, domain shift refers to distribution differences between training and deployment data caused by variations in sensors, illumination conditions, viewing geometries, terrain types, or mission scenarios. As summarized in Table 1, this review differs from previous work by integrating dataset systems, perception tasks, technical bottlenecks, and future research opportunities into a unified framework.
The main contributions of this paper are as follows.
1) A structured review of Mars image datasets. We summarize major benchmark data sets for both orbital remote sensing and in situ exploration, and provide a reference for future algorithm training, validation, and cross-dataset comparison.
2) A cross-task synthesis of computer vision methods for Mars perception. We review the methodological evolution of core tasks, including classification, detection, segmentation, depth estimation, 3D reconstruction, and multi-task perception.
3) An analysis of Mars-specific constraints and future directions. We discuss how the Martian environment and mission requirements drive the development of lightweight models, self-supervised learning, domain adaptation, multimodal perception, and Mars foundation models.
The remainder of this paper is organized as follows: Section 2 reviews Mars image datasets for deep learning research. Section 3 summarizes major computer vision tasks and methods for Mars perception. Section 4 discusses mission-specific requirements, including limited computing resources, scarce labels, domain shifts, and incomplete multimodal information. Section 5 identifies current challenges and technical bottlenecks. Section 6 outlines future directions, including Mars foundation models, embodied autonomous science, uncertainty-aware perception, and multimodal edge intelligence. Section 7 concludes the paper. The overall framework of this structure is illustrated in Fig. 3.
2 Mars image datasets
Mars exploration relies mainly on two types of image data: orbital remote sensing data and in situ exploration data. They differ in acquisition geometry, spatial resolution, observation scale, sensor configuration, and application scenario. These differences lead to distinct visual characteristics. They also determine how computer vision methods should be designed, trained, and evaluated. Orbital data provide broad regional to global coverage, whereas in situ data capture local scenes at much finer scales. Together, they support a complete perception framework for Mars exploration, from global terrain mapping to local rover navigation and scientific target selection.
2.1 Mars orbital remote sensing data
Mars orbital remote sensing data are primarily acquired by imaging systems onboard Mars orbiters, including HiRISE and Context Camera (CTX) onboard MRO, the Thermal Emission Imaging System (THEMIS) onboard Mars Odyssey, and the High-Resolution Stereo Camera (HRSC) onboard Mars Express. As shown in Fig. 4, they provide wide spatial coverage, long temporal baselines, and multi-resolution observations. These properties make them essential for studying Martian landforms, geological structures, atmospheric phenomena, and surface changes.
Impact craters are among the most widely studied targets in Mars orbital image analysis. Their abundance, morphology, and spatial distribution make them important benchmarks for object detection and segmentation algorithms. Early work constructed a THEMIS daytime infrared semantic segmentation dataset to train convolutional neural networks (CNNs) for crater detection within a defined size range (
DeLatte et al., 2019a). To improve model robustness to scale variations, the MDCD dataset with multi-scale annotations (
Yang and Cai, 2021) and a large THEMIS-based global detection data set were later released (
Hsu et al., 2021). The latter contains more than 190000 crater annotations from 92575 image tiles of 256 × 256 pixels. Domain shift has also become an important concern in crater detection. The DACD data set was developed for day-to-night Martian crater detection and contains approximately 1000 images with more than 20000 crater bounding boxes (
Yang and Cai, 2022). This dataset provides a benchmark for evaluating domain adaptation under illumination-induced domain shift. In this context, domain shift refers to differences between training and testing data distributions, such as brightness, shadow patterns, contrast, and crater appearance, which may may degrade cross-domain detection performance.
Beyond crater detection, orbital datasets have expanded toward smaller, more diverse, and more dynamic targets. The MDD-Human dataset built from CTX imagery, provides 8052 training samples for dust devil detection (
Liu et al., 2025b). A more recent multimodal dust devil dataset integrates object detection and semantic segmentation, with annotations for dust devil heads and shadow contours (
Xu et al., 2026). These datasets indicate a clear shift from single-target detection toward detailed, multi-task interpretation of Martian surface and atmospheric phenomena.
Regional semantic segmentation datasets have also been developed for geological-process mapping. The C-PLES dataset targets landslide recognition in Valles Marineris and integrates seven data layers, including RGB imagery, digital elevation model, thermal inertia, and slope. It provides 682 landslide polygon annotations (
Reyes et al., 2023). The MMLSv2 benchmark further introduces a geographically isolated test set to evaluate spatial generalization (
Paheding et al., 2026). For geomorphological classification tasks, DoMars16k dataset contains 16150 CTX image patches covering 15 geomorphic types in five major categories, including aeolian beds, slope features, and impact structures (
Wilhelm et al., 2020). For hyperspectral mineral identification, The HyMars benchmark covers three representative regions, including the Holden crater, and serves as a benchmark for few-shot hyperspectral classification (
Xi et al., 2025).
Despite these advances, many orbital datasets remain fragmented. They differ in annotation format, task definition, sensor type, spatial resolution, and evaluation protocol. This makes cross-dataset comparison difficult. It also limits the assessment of model transferability. Mars-Bench was developed to address this problem by integrating 20 expert-validated orbital and in situ datasets, converting them into standard formats such as COCO (
Lin et al., 2014) and VOC (
Everingham et al., 2010), and defining unified protocols for classification, segmentation, and detection (
Purohit et al., 2026a). This benchmark marks an important transition from isolated dataset construction toward standardized evaluation in Mars computer vision.
Table 2 summarizes representative publicly available Mars orbital remote sensing datasets for deep learning. The datasets are organized by target type and publication year to highlight the evolution from crater-centered detection to multimodal, multi-task, and standardized benchmarks.
2.2 Mars in situ exploration data
Mars in situ exploration data are directly acquired on the Martian surface by scientific payloads onboard landers or rovers. These data are characterized by high spatial resolution, limited spatial coverage, unstructured terrain, and complex illumination conditions. Representative instruments include the Mast camera (Mastcam), Navigation Camera (Navcam), and Navigation and Terrain Camera (NaTeCam). In situ images are central to autonomous rover perception, terrain traversability analysis, obstacle detection, and scientific target selection.
Compared with orbital data, in situ images are more difficult and costly to annotate at the pixel level. Accordingly, existing Mars in situ data sets can be grouped according to their annotation strategies and benchmark purposes, including expert-labeled data sets, crowdsourced annotations, synthetic generation, sparse or weak annotations, and domain adaptation benchmarks. Crowdsourced annotation refers to collecting labels from many non-expert annotators using simplified terrain categories, followed by label aggregation or quality-control procedures. Sparse annotation reduces labeling cost by labeling only selected pixels, regions, or key terrain elements instead of producing dense labels for the entire image. Domain adaptation benchmarks define source and target domains with explicit distribution shifts, such as differences in illumination, sensor type, landing site, mission scenario, or synthetic-to-real appearance, to evaluate cross-domain generalization.
Early
in situ exploration data sets were relatively small and task-specific. A rock classification data set comprising 620 Mastcam images was developed to evaluate transfer learning under few-shot Mars conditions (
Xiao et al., 2021). MarsData data set provided single-rock and multi-rock subsets for rock detection and served as an early supervised learning frameworks (
Li et al., 2020). As rover perception shifted from simple target recognition to scene-level understanding, large-scale semantic segmentation data sets became increasingly important. AI4MARS represented a key milestone (
Swan et al., 2021). It collected more than 35000 Navcam and Mastcam images from Curiosity, Opportunity, and Spirit rovers. Through crowdsourcing, it produced more than 326000 full-image semantic annotations with four simplified categories: soils, bedrocks, sands, and large rocks. AI4MARS demonstrated that non-expert annotation can rapidly generate large-scale traversability data sets. However, its coarse label scheme remains insufficient for detailed geological interpretation.
Mars-Seg was introduced to provide more detailed semantic categories (
Li et al., 2025b). It contains over 5000 finely annotated images, from grayscale images from the Mars Exploration Rover (MER) mission to color images from the Mars Science Laboratory (MSL) mission, and extends the semantic categories to nine classes, including soils, sands, bedrocks, rocks, gravels, rover tracks, and shadows. This label scheme better matches the needs of rover mobility analysis and scientific exploration. It has also become an important benchmark for domain adaptation. MarsScapes further extends
in situ perception to panoramic scenes by stitching 195 Curiosity Mastcam panoramas. The data set has been used to evaluate unsupervised domain adaptation frameworks after sample augmentation (
Liu et al., 2023d).
Synthetic data sets provide another important solution to annotation scarcity. They are especially useful when ground truth is difficult to obtain from real Mars images, such as depth, instance masks, and dense semantic labels. The open-source Outdoor Artificial Intelligent SYstems Simulator (OAISYS) pipeline, built on Blender, can generate photorealistic synthetic images with rich annotations (
Müller et al., 2021). Based on this pipeline, SimMars6K provides 6325 stereo navigation camera images pairs with high-precision semantic, instance, and depth ground truth (
Ma et al., 2024). Its effectiveness has been validated through transfer learning for rock detection. SynMars simulates Zhurong rover perspective and contains up to 60000 images. Together with the MarsData-V2 dataset, it forms a multi-source benchmark for rock segmentation (
Xiong et al., 2023).
Two recent trends are particularly notable. The first is the move from fully supervised learning toward semi-supervised and sparsely annotated settings. S
5Mars data set contains 6000 high-resolution Mastcam images with nine semantic classes sparsely annotated by experts, where approximately 49% of pixels carry high-confidence labels (
Zhang et al., 2024c). This design reduces annotation cost while supporting semi-supervised learning. The second trend is the release of mission-specific data sets from Tianwen-1 and Zhurong. The Zhurong Data set comprises 1376 stereo NaTeCam image pairs annotated with nine terrain classes (
Xu et al., 2025). TWMARS and ZhurongRock data sets focus on binary rock semantic segmentation (
Lv et al., 2022;
Jia et al., 2024). Collectively, these data sets provide valuable benchmarks for stereo scene parsing, rock segmentation, and autonomous obstacle avoidance in the Zhurong traverse region.
Table 3 summarizes representative publicly available Mars in situ exploration data sets for deep learning. The data sets are grouped into synthetic/simulated data sets and real image data sets, reflecting the increasing importance of synthetic-to-real transfer and domain adaptation.
3 Core tasks and methods for mars intelligent perception
With the continued advancement of mars exploration missions, multi-source image data from orbiters, landers, and rovers have increased rapidly in both volume and complexity. These data span global orbital surveys and local in situ investigations. They also cover a wide range of spatial resolutions, viewing geometries, and sensing modalities. Traditional manual interpretation and classical image processing methods are increasingly insufficient for efficient, objective, and large-scale extraction of scientific information. Deep learning, with its capacity for end-to-end feature representation and pattern recognition, has become an important technical pathway for Mars intelligent perception. This section systematically reviews the major computer vision tasks and methodological advances in two representative scenarios: orbital remote sensing and in situ exploration. The discussion follows the macroscopic-to-microscopic observation chain of Mars missions.
3.1 Macroscopic perception from an orbital remote sensing perspective
Cameras and spectrometers onboard Mars orbiters continuously monitor the Martian surface at different spatial and spectral resolutions. These instruments provide the primary data basis for large-scale landform mapping, surface change detection, and environmental monitoring. In this context, deep learning methods are mainly used to identify, delineate, classify, and track geomorphic features and dynamic processes across broad regions. Orbital perception therefore emphasizes large-area coverage, cross-scale feature extraction, and robust generalization across sensors and illumination conditions. Figure 5 illustrates representative perception tasks for orbital remote sensing.
3.1.1 Object detection: from impact craters to automated recognition of multiple landform types
Object detection is one of the most active research directions in Mars orbital image analysis. Its core objective is to locate and classify discrete geomorphic targets in large-scale remote sensing images (
Cao et al., 2024). Impact craters have long been the dominant benchmark for this task. They are abundant, morphologically distinctive, and scientifically important for surface dating and geological interpretation. Their size, morphology, and spatial distribution provide key evidence for inferring regional geological age and surface evolution (
Martinez et al., 2025).
The development of Mars object detection can be broadly divided into three stages: early adaptation of generic CNN detectors, task-specific optimization for multi-scale and small targets, and recent exploration of foundation-model-based detection. Early studies mainly focused on applying general object detection frameworks to Martian imagery. For example, the single-stage detector YOLO was used for rapid detection of impact craters larger than 3 km in thermal infrared images. On a self-constructed crater data set, the study reported mean Average Precision at an Intersection-over-Union threshold of 0.50 (mAP@0.50) values above 82%, including 82.7% for YOLOv5s and 82.5% for YOLOv5l (
Li et al., 2022). Lightweight variants such as YOLOv4-tiny were also tested for fast localization under limited sample conditions, achieving an mAP@0.50 of 88.36% reported on a self-constructed grayscale Martian crater data set (
Barman et al., 2022). These studies established the basic feasibility of deep learning-based crater detection.
As the field developed, scale variation became a central technical challenge. Craters and other Martian landforms vary greatly in size, preservation state, and image contrast. To reduce feature loss for small craters, multi-scale feature fusion has become a major strategy. CS-YOLO improves small-object response through optimized downsampling and cross-layer feature fusion, reporting a reported recall of 82.1% and accuracy of 86.9% for small-crater detection on Mars solar thermal infrared remote-sensing images (
Xiao et al., 2025b). HRFPNet introduces a high-resolution feature pyramid and a balanced regression loss to suppress coordinate regression bias for small targets (
Yang and Cai, 2021). Efficiency has also become important. Mini-SSD reduces parameter count through streamlined architectures and provides a lightweight option for detection tasks (
Jiang et al., 2022). Faster R-CNN and its variants have further demonstrated advantages over handcrafted feature detectors in identifying multi-scale landforms, including volcanic rootless cones and transverse aeolian ridges (
Palafox et al., 2017).
Object detection has gradually expanded beyond crater-centered applications. For aeolian landforms, Mask R-CNN has been used for instance segmentation and contour extraction of barchan dunes, with cross-planet transfer capability validated from Mars to Earth. The study reported an mAP@0.50 of 77% on the Martian test data set (
Rubanenko et al., 2021). For dust devils, the MDD-Human data set and Transformer-based MDT network introduce shape-aware learning to improve sensitivity to elongated structures (
Liu et al., 2025b). An improved Faster R-CNN with a feature pyramid network and Soft-NMS further reduces missed detections in densely distributed dust devil targets, with an average precision of 90.1% and a recall of 96.5% reported on the constructed dust devil data set (
Guo et al., 2024). For volcanic and tectonic landforms, Rotated-SSD uses rotated anchors to localize arbitrarily oriented linear features, such as rock piles and dark slope streaks (
Jiang et al., 2021). Martian skylight identification has also been explored by combining generative adversarial sample synthesis with an improved YOLOv9 model to alleviate positive-sample scarcity (
Li et al., 2025e). Meanwhile, crater-related detection continues to become more refined, including YOLOv7-based intelligent detection, which reported a recall of 86.45% and a precision of 88.40% on a reconstructed crater data set (
Yu, 2024), pitted cone identification using EMHA-BiFPN-SSD (
Yu et al., 2025), and boundary delineation methods that integrate geographic information with deep learning (
Liu et al., 2023a).
More recently, foundation models, that is, large-scale pre-trained models that can be adapted to different downstream tasks with little or no task-specific training, have introduced a new zero-shot paradigm for planetary object detection. The Segment Anything Model (SAM), for instance, can generate candidate crater masks without task-specific fine-tuning. By applying geometric post-processing to SAM outputs, such as filtering based on shape indices, Giannakis et al. showed that foundation models can generalize across images acquired by different sensors (
Giannakis et al., 2024). This reduces the dependence on task-specific labeled data sets. However, foundation-model outputs still require careful filtering and geological validation, especially for degraded, overlapping, or morphologically ambiguous landforms. The emerging trend is therefore not to replace task-specific detectors entirely, but to combine generalization capability of foundation-models with planetary-domain constraints.
3.1.2 Image segmentation: detailed contour extraction and geological unit delineation
Semantic and instance segmentation aim to assign pixel-wise semantic labels and distinguish individual objects. They are essential for extracting crater boundaries, delineating dune and dust storm extents, outlining landslide contours, and supporting automated geological unit mapping (
Kabir et al., 2026;
Li et al., 2026). Compared with object detection, segmentation provides more detailed spatial information. It is therefore particularly important for quantitative geomorphic analysis and fine-scale geological interpretation. With the development of CNNs, Transformers, and more recently foundation models, segmentation has become a core task in Mars orbital remote sensing.
Crater segmentation has served as one of the earliest and most important test cases. Accurate crater boundary extraction is fundamental for measuring crater size, morphology, and degradation state, which are closely related to surface dating and geological evolution. Early work used U-Net for binary segmentation of THEMIS global imagery, demonstrating the feasibility of pixel-level crater recognition. In that study, the segmentation network identified approximately 65%–76% of the craters that were also present in a human-annotated data set (
DeLatte et al., 2019b). SqUNet further improved multi-level feature extraction through embedded encoder–decoder substructures and showed better generalization across cross-body DEM data (
Zhao and Ye, 2023). For instance-level crater extraction, BDMCI combined Mask R-CNN with Mask2Former and incorporated slope-difference geographic information to improve boundary conformity (
Liu et al., 2023a). Foundation models have recently introduced a different route. SAM has been directly applied to HRSC, THEMIS, and HiRISE imagery. With geometric post-processing, its overall recall can approach that of conventional supervised methods (
Giannakis et al., 2024). The EASSA framework further combines SAM-based initial segmentation with Wiener filtering and edge optimization, achieving high-precision contour extraction of rampart craters. On a multiscale rampart crater data set, EASSA reported a detection accuracy of 97.36%, a recall of 93.36%, and an IoU of 0.93 (
Wang et al., 2025a). These studies show a clear shift from fully supervised crater segmentation toward hybrid workflows that combine foundation-model generalization with domain-specific geometric filtering.
Segmentation tasks have also expanded from craters to diverse Martian surface and atmospheric phenomena. For aeolian landforms, Mask R-CNN has been applied to instance segmentation of barchan dunes, enabling the extraction of more than one million individual dunes and revealing hemispheric differences in dune distribution (
Rubanenko et al., 2022). Mask R-CNN with Dice loss has been used for dust storm instance segmentation on daily global maps (
Alshehhi and Gebhardt, 2022a), while U-Net has been applied to textured dust storm regions in MOC imagery (
Ogohara and Gichu, 2022;
Adams et al., 2023) . These applications demonstrate the value of segmentation for both surface-process analysis and atmospheric monitoring.
Multimodal segmentation has become particularly important for complex geological processes. Landslide mapping in Valles Marineris, for example, requires not only optical texture but also topographic and thermophysical information. C-PLES integrates RGB imagery, DEM, thermal inertia, and slope through contextual progressive layer expansion and self-attention, enabling multi-class landslide segmentation (
Reyes et al., 2023). MarsLS-Net employs stacked progressive expansion neuron attention blocks for lightweight and rapid multimodal segmentation (
Paheding et al., 2024). M3LSNet further introduces the Mamba architecture and fuses RGB, DEM, thermal inertia, slope, and grayscale imagery, achieving higher mean Intersection over Union (mIoU) than several mainstream models (
Dai and Cui, 2025b). It shows a clear trend that difficult geological targets increasingly require cross-modal feature fusion rather than image-only segmentation.
For geological mapping and linear structure identification, U-Net and DeepLabV3 + have been validated for automated delineation of planetary geological units using a multimodal Jezero crater data set (
Wilhelm et al., 2022). U-Net has also been applied to preliminary semantic segmentation of linear structures on DEM imagery (
Xiu et al., 2022). For aeolian landform interpretation at the local scale, a cross-scale encoding method jointly using HiRISE imagery, topographic attributes, and rover-derived semantic features improved the boundary smoothness and detail fidelity in Meridiani Planum through visual comparison with ground-truth annotations (
Choromański et al., 2022). For efficient large-scale mapping, MarsMapNet combines superpixel segmentation with multi-view feature fusion, and reduces mapping time by approximately two orders of magnitude compared with pixel-level methods (
Zhao et al., 2024). These studies indicate that deep learning-based automated methods are feasible for assisting Martian geological unit mapping. For instance, U-Net model with an EfficientNet-B0 backbone achieved an mIoU of 0.66, while DeepLabV3 with the same backbone yielded an mIoU of 0.62 (
Wilhelm et al., 2022). These results suggest that automated delineation can provide useful support for expert interpretation, although current methods still require further improvement, cross-region validation, and expert geological verification before being used as operational mapping tools.
Overall, Mars orbital segmentation has evolved from single-class binary segmentation to multi-class, multimodal, and increasingly foundation-model-assisted interpretation. Current methods can delineate craters, dunes, dust storms, landslides, linear structures, and geological units with increasing accuracy. However, several challenges remain. Boundary extraction is still difficult for degraded, overlapping, or weakly expressed landforms. Few-shot classes remain poorly segmented. Cross-region and cross-sensor generalization is still inconsistent. Future progress will likely depend on combining foundation models, self-supervised feature learning, multimodal fusion, and planetary-domain constraints. Such integration will be essential for building global-scale, high-precision semantic maps and for supporting quantitative interpretation of Martian surface processes.
3.1.3 Image classification: large-scale geological mapping
Global-scale geological mapping of Mars requires automatic discrimination of landform types, surface units, and mineral compositions from massive orbital imagery. Image classification provides a direct way to assign semantic labels to image patches or spectral-spatial units. It supports rapid cataloging and spatial mapping of typical targets, including impact craters, dunes, polar deposits, and mineral-bearing terrains. Compared with object detection and segmentation, classification usually provides coarser spatial information, but it is highly efficient for large-area surveys. Deep learning methods have gradually replaced handcrafted features and shallow classifiers, improving the automation and scalability of Mars orbital image interpretation.
Early classification studies mainly relied on transfer learning, i.e., adapting models pre-trained on large external data sets to Martian imagery, to overcome the scarcity of labeled Martian data. HiRISENet applied AlexNet-based transfer learning to HiRISE imagery and constructed a six-class landform classification model. It was deployed on the NASA Planetary Data System and provided the first publicly available content-based retrieval capability for orbital images (
Wagstaff et al., 2018). Subsequent work compared VGG16, InceptionV3, and ResNet50 on seven HiRISE landform classes, confirming the advantages of deeper CNN architectures (
Mohith et al., 2025). MarsDeepNet further modified GoogLeNet by adding convolutional layers and introducing Leaky ReLU, thereby improving feature representation for eight-class terrain classification (
Tiwari et al., 2024). These studies established transfer learning as an effective baseline for Mars image classification.
As classification methods moved toward practical deployment, lightweight design became increasingly important. Onboard and near-real-time applications require efficient inference with limited computing resources. An improved MobileNetV2 incorporating non-local modules and DropBlock regularization achieved a reported classification accuracy of 93.64% in the original study setting and outperformed VGG-16, the original MobileNetV2, ResNet-34, and HiRISENet under the same experimental comparison (
Xia et al., 2023). Knowledge distillation further compresses a high-capacity teacher model by several tens of times while preserving classification accuracy (
Bosowski et al., 2023). These approaches provide feasible routes for edge-oriented Mars classification.
Another important direction is improving classification under rare classes, noisy labels, and distributed data conditions. TerraMorphNet combines expert-designed Martian morphological descriptors with deep embeddings from ConvNeXtV2, improving recall on rare classes (
Roy et al., 2025). Modified ResNet50 has been used for few-shot lineament classification, demonstrating the applicability of deep residual networks to underrepresented structural features (
Yan et al., 2020). A federated learning framework based on Flower enables collaborative training across multiple institutions without directly sharing raw data, offering a potential paradigm for classification on institutionally distributed Mars image data sets (
Chithambaram et al., 2023). For CRISM hyperspectral mineral identification, confident learning dynamically adjusts the loss weights of low-confidence samples and reduces the influence of label noise caused by spectral overlap (
Soor et al., 2025). These methods indicate that classification research is shifting from simple accuracy improvement toward robustness under realistic data limitations.
At the architecture level, new backbone networks continue to expand the representation capacity of Mars classification models. Early shallow CNNs demonstrated that learned features can outperform traditional handcrafted classifiers in dune binary classification. In the original study setting, the CNN model achieved a validation accuracy of 82.11% and a maximum accuracy close to 85%, compared with approximately 80% for traditional classifiers (
Dutta et al., 2023). More recently, MMCLS introduced the Vision Mamba architecture for Mars image classification. By leveraging state space models to capture global context efficiently, this approach outperforms conventional CNN architectures and indicates the potential of state-space modeling for planetary image analysis (
Dai and Cui, 2025a).
Mars orbital image classification has evolved from transfer-learning validation to a broader framework involving lightweight deployment, rare-class recognition, distributed learning, hyperspectral mineral classification, and novel backbone architectures. Current methods have shown strong potential for automated landform cataloging and mineral mapping. However, several challenges remain. Rare classes are still difficult to recognize. Cross-sensor generalization remains unstable. Shadows, dust cover, illumination differences, and spectral overlap can reduce classification reliability. Future progress will likely depend on self-supervised pre-training, multimodal spectral-spatial modeling, uncertainty-aware classification, and hardware-efficient inference. These developments may support the construction of global-scale and high-confidence geological maps of Mars.
3.1.4 Change detection: time-series monitoring of dynamic processes on the Martian surface
Mars is a dynamically evolving planet. Its surface and atmosphere undergo seasonal and interannual changes, including slope streak activity, periglacial collapse, dust storm migration, and aeolian sand movement. Detecting these changes is important for understanding present-day climate processes, surface material transport, and potential geological hazards. Traditional multi-temporal interpretation relies heavily on manual comparison. Such interpretation is inefficient and often fails to capture subtle, sparse, or weakly expressed changes. Deep learning provides a new pathway for automated change detection in bi-temporal and time-series orbital imagery (
Su et al., 2025). To avoid treating overlapping approaches as strictly mutually exclusive categories, this section discusses recent studies according to their primary methodological emphases rather than as a rigid taxonomy. Existing works can be broadly examined from three perspectives: pairwise comparison architectures, efficiency- and adaptation-oriented models, and transferable change representation learning. These perspectives may overlap in practice, as a single method can combine a Siamese architecture with lightweight design, unsupervised adaptation, and transferable representation learning.
Pairwise comparison architectures, particularly Siamese networks, have become a common baseline for Martian change detection. Their dual-branch structure is well suited for comparing image pairs while preserving temporal correspondence. SLCD-Net uses dual inputs to avoid early noise fusion and enhances change-region responses through multi-level complementary fusion and spatial attention. It achieves good sensitivity and robustness in detecting changes in recurring slope lineae and dark slope streaks on HiRISE imagery (
Zhang et al., 2024a). For ice-fragment detachment on steep slopes of the north polar layered deposits, Su et al., introduced an enhanced attention module into the Siamese framework. Dice loss and Focal loss were jointly used to reduce class imbalance and improve the detection of rare periglacial collapse events (
Su et al., 2023).
Efficiency- and adaptation-oriented models address two practical needs: onboard efficiency and cross-sensor generalization. MViT-PCD adopts MobileViT as its backbone and integrates multi-scale feature-difference fusion with unsupervised domain adaptation. This design mitigates sensor-domain shifts while maintaining a low parameter count and high inference efficiency, with reported accuracies of 97.2% and 82.9% on two public Martian change-detection data sets and an inference speed of 43.4 frame per second (FPS) (
Dai et al., 2023). DUSTNet targets dust storm change detection. It uses a noise-resistant module to reduce the effects of registration errors and pseudo-changes, and shows robust cross-sensor performance on a self-constructed data set (
Li et al., 2025f).
Beyond specific landform or atmospheric targets, transferable change representation learning is also important. Convolutional autoencoders combined with transfer learning offer one feasible route for cross-landform and cross-task generalization. Comparative studies show that bottleneck features extracted by convolutional autoencoders can maintain low false-positive rates across different change detection tasks, including recurring slope lineae and new impact craters (
Kerner et al., 2019). That is, transferable representations may reduce the need to design a separate model for every type of Martian change.
Mars orbital change detection is moving from target-specific supervised models toward lightweight, unsupervised, and transferable frameworks. Current methods have improved the monitoring of surface and atmospheric dynamics. However, several challenges remain. Weak change signals are still difficult to detect. Non-ideal image registration can generate false changes. Pixel-level localization accuracy remains limited, especially for small or diffuse targets. Future progress will likely depend on multi-temporal Transformer architectures, physics-aware domain adaptation, uncertainty estimation, and highly compressed models for onboard deployment. These advances will be essential for achieving more real-time, fine-scale, and automated monitoring of active Martian surface processes.
3.2 Detailed environmental understanding in in situ exploration
As exploration missions move from orbit to the surface, cameras onboard landers and rovers provide in situ images at centimeter- to millimeter-scale. The focus of computer vision therefore shifts from large-scale landform recognition to detailed understanding of the local environment along rover traverse. This shift is critical because rovers must recognize terrain, avoid hazards, select scientific targets, and support close-range operations under uncertain surface conditions. In situ perception is thus central to both mission safety and scientific return. Figure 6 presents representative perception tasks for in situ exploration. To make the application relevance explicit, the methods reviewed in this section are mapped to specific rover or lander functions, including traversability analysis, obstacle avoidance, scientific target selection, data prioritization, visual localization, 3D reconstruction, and image-quality restoration.
3.2.1 Image segmentation: unstructured terrain segmentation
Semantic segmentation is a core task for Mars rover environmental perception. It assigns pixel-level labels to surface images and partitions unstructured terrain into traversability-related categories, such as soil, bedrock, sand, and rocks. These outputs provide detailed environmental priors for path planning, obstacle avoidance, and local scientific interpretation (
Fan et al., 2025). Compared with orbital segmentation, in situ segmentation is more sensitive to viewpoint changes, shadows, extreme illumination, ambiguous terrain boundaries, and limited annotations. These constraints have driven five major technical directions, including specialized CNN architectures, Transformer-based segmentation, lightweight real-time models, label-efficient learning, and multimodal fusion.
Early architectural improvements mainly focused on preserving spatial details and reducing boundary confusion. ECRNet introduces an enhanced atrous spatial pyramid pooling module with coordinate attention to preserve spatial information, together with a context refinement module that models inter-class relationships and boundary priors (
Xu et al., 2025). MarsSeg uses multi-level feature extraction, fusion modules, and pixel-level attention to achieve cross-level information and improve rare-class recognition (
Li et al., 2025b). HASS adopts a dual-branch hybrid attention design to model global intra-class consistency and local inter-class relationships, while a local diversity loss constrains adjacent heterogeneous terrains (
Liu et al., 2023b). BIDF separates label interiors from ambiguous boundaries using center and boundary streams and addresses annotation imprecision through multi-task training, achieving an average precision (AP) of 89.1% (
Wang et al., 2022).
Transformer-based architectures have further improved long-range context modeling. SegFormer has been applied to Martian terrain segmentation, demonstrating the effectiveness of hierarchical self-attention for unstructured rover scenes, reporting an accuracy of 90.86% and an mIoU of 83.55% on the AI4Mars data set (
Goutham et al., 2022). Comparative studies of CNNs and Transformers indicate that Transformers are particularly useful for high-resolution segmentation of rare obstacles, while SegFormer offers a favorable balance between accuracy and real-time performance (
Mohammad et al., 2025). U-shaped Transformer networks, including MarsFormer, RockFormer, and EDR-TransUnet, combine global self-attention with local detail enhancement and improve the fineness and robustness of rock segmentation (
Liu et al., 2023c;
Xiong et al., 2023;
Jia et al., 2024). The main advantage of these models lies in their ability to integrate global scene context with local terrain details.
Lightweight and real-time designs are essential for onboard deployment. Rover perception cannot rely on large, high-capacity networks. LBNet achieves efficient fusion of spatial details and semantic information through a low-parameter-count dual-branch design, demonstrating potential for real-time online rock detection. With only 0.37 million parameters, LBNet reported mIoU/FPS values of 93.85%/147.8 on the Perseverance data set and 88.62%/152.5 on the Curiosity data set (
Wei et al., 2024). Mobile-DeepRFB adopts a lightweight backbone and receptive field block to compress model scale and have been validated at practical frame rates on embedded platforms (
Feng et al., 2023). SegMarsViT, based on lightweight MobileViT, captures both local and global contexts, achieving a favorable trade-off between parameter count and inference speed. It reported mIoU values of 68.40%, 78.22%, and 67.28% on the AI4Mars-MSL, MSL-Seg, and S5Mars data sets, respectively, with an inference speed of 69.52 FPS (
Dai et al., 2022). CloverNet further demonstrates high-frame-rate segmentation on edge devices through TensorRT quantization, providing end-to-end validation for onboard autonomous perception (
Gasperini et al., 2023).
To address the issue brought by scarce labeled data, semi-supervised, self-supervised, and domain adaptation methods are widely used to reduce dependence on pixel-level labels. S
5Mars introduces strong augmentation tailored to Martian image characteristics and confidence-guided pseudo-label consistency learning, achieving competitive performance under few-shot conditions (
Zhang et al., 2024c). The MBE network incorporates adaptive local data augmentation and a symmetric cyclic focal loss to improve unlabeled data utilization and mitigating class imbalance. On the S5Mars data set, it reported an Intersection over Union (IoU) of 68.28%, outperforming the second-best method by 1.40% under the 1/4 partition protocol (
Xiao et al., 2025c). A contrastive learning-based semi-supervised method achieves high accuracy with a limited number of labeled samples and substantially improves recall for rare classes (
Goh et al., 2022). For unsupervised domain adaptation, UDAFormer uses a teacher-student framework with improved augmentation regularization to transfer knowledge from synthetic data to real Mars images. On the MarsScapes-A→MarsScapes-B unsupervised domain adaptation (UDA) task, UDAFormer reported the highest mIoU of 57.95%, outperforming the second-best UDAFormer by 2.55% and the best adversarial-learning method, LTIR, by 6.76% (
Liu et al., 2023d). A transfer learning knowledge reuse framework further leverages historical missions and simulated scenarios, demonstrating the segmentation capability can be rapidly established for a rover’s new landing site (
Liu et al., 2022). These approaches are especially important for future missions, where new terrains may be encountered before sufficient labels are available.
Multimodal fusion provides an additional route to robust
in situ segmentation. RGB images alone may be insufficient under shadows, dust, or visually ambiguous terrain. Depth, thermal, and other physical measurements can provide complementary information. OmniUnet jointly models RGB imagery, depth information, and thermal imaging data, using thermal inertia differences among surface materials to distinguish key traversable terrain classes. On a Jetson Orin Nano platform, OmniUnet reported a pixel accuracy of 80.37% and an inference time of 673 ms per multimodal image (
Castilla-Arquillo et al., 2025). The trend is that rover-scene segmentation is moving from single-image semantic labeling toward multimodal and physically informed terrain understanding.
Overall, in situ semantic segmentation has evolved from CNN-based terrain parsing to Transformer-enhanced, lightweight, label-efficient, and multimodal perception. Current methods have improved the recognition of rocks, soil, bedrock, sand, shadows, and traversability-related classes. However, robust deployment remains challenging. Models must handle rare obstacles, uncertain boundaries, shadows, dust, domain shifts, and limited onboard resources at the same time. Future progress will likely depend on combining efficient architectures, synthetic-to-real adaptation, uncertainty-aware segmentation, and multimodal physical constraints.
3.2.2 Object detection: high-value target detection and localization
In Mars in situ exploration, object detection aims to locate mission-relevant targets in complex surface environments. Unlike orbital object detection, which focuses mainly on large-scale landforms, rover-based detection deals with rocks, scientific targets, and engineering objects at centimeter-to-meter scales. The outputs directly support autonomous obstacle avoidance, robotic arm operation, sample handling, and scientific target selection. Accuracy and reliability are therefore critical.
Rock and terrain-object detection forms the foundation of rover perception. Early studies demonstrated the feasibility of lightweight single-stage and two-stage detectors for Martian surface scenes. (
Guo et al., 2023) systematically compared YOLOv5, CenterNet, and Faster R-CNN on a Martian terrain data set and showed that, with hyperparameter tuning and data augmentation, lightweight models can distinguish traversability-related classes such as sands, bedrocks, and rocks at competitive inference speeds. MarsRock-Faster R-CNN integrates a ResNet50 backbone with a feature pyramid network to handle large rock-scale variations. It improves the detection of rocks ranging from 0.5 to 2 m in diameter and has been used to estimate rock abundance in Utopia Planitia, providing quantitative support for landing-site assessment. The method reported precision and recall values of approximately 80% for Mars rock extraction from HiRISE images (
Wang et al., 2025b). Although this study is based on orbital HiRISE imagery rather than rover
in situ images, it is discussed here because orbital rock-abundance estimation provides upstream information for landing-site safety assessment, traverse planning, and surface mission preparation. Furlán et al. proposed an improved U-Net with a simplified encoder-decoder structure, substantially reducing parameter count and achieving rock detection and segmentation within seconds per image in a Mars-like simulated environment. This demonstrates the potential of lightweight models for near real-time obstacle perception (
Furlán et al., 2019). Agarwal et al. further explored the conversion of semantic segmentation masks from AI4Mars into bounding boxes for YOLOv8-based detection of navigation-related terrain classes. In the reported experiment, the best model achieved an mAP@0.50 of 34.33%, an mAP@0.50:0.95 of 24.30%, a precision of 57.31%, and a recall of 33.61% at Epoch 30. However, this strategy can lose spatial detail during mask-to-box conversion, which limits detection accuracy (
Agarwal et al., 2025).
For high-value scientific targets and engineering objects, the main challenge is extreme sample scarcity. This has made physical simulation, synthetic data generation, and domain adaptation particularly important. In the Mars Sample Return context,
Daftry et al. (2021) compared traditional template matching with Mask R-CNN-based instance segmentation for sample tube grasping. Deep models showed stronger robustness to shadows, partial occlusion, and dust coverage, whereas template matching retained engineering interpretability and flight-heritage advantages.
Castilla-Arquillo et al. (2022) proposed a lightweight transfer learning framework to further reduce dependence on real labels. A YOLOv3-tiny detector was first pre-trained on synthetic data generated under diverse illumination conditions. Its backbone was then frozen and fine-tuned using a small set of real Mars-like images. Virtual-to-real transfer has also been explored for rare scientific targets. For shatter cone detection, where positive samples are usually unavailable,
Bechtold et al. (2023) inserted 3D digital models into stereo-reconstructed Martian scenes to generate synthetic training data for a Mask R-CNN. Cross-scene generalization was then validated through field tests in a Mars-analog desert environment.
Mars in situ object detection research is moving from the direct transfer of generic detection models toward task-specific, simulation-assisted, and resource-aware frameworks. The central objective is to localize rocks, hazards, scientific targets, and engineering objects reliably under limited annotations and onboard computing resources. Future progress will likely depend on combining lightweight detectors, synthetic-to-real adaptation, uncertainty estimation, and task-specific physical constraints. Such integration is essential for safe rover operation and efficient scientific exploration.
3.2.3 Image classification: rapid semantic scene assessment
Image classification provides rovers with rapid semantic discrimination of local scenes. It supports scientific target selection, data prioritization, anomaly screening, and autonomous decision-making. Unlike object detection, which localizes individual instances, and semantic segmentation, which assigns pixel-level labels, image classification assigns semantic labels to entire images or image patches. It is therefore useful for fast terrain assessment and scene-level understanding. To handle variable illumination, changing viewpoints, complex backgrounds, and scarce annotations, researchers have explored transfer learning, attention mechanisms, semi-supervised learning, self-supervised learning, clustering, and novel backbone architectures.
Transfer learning provided an early and practical solution to the limited availability of labeled Martian surface images.
Wagstaff et al. (2018) applied an ImageNet pre-trained model to Mars surface image classification and developed the first content-based retrieval system for Curiosity imagery, validating the feasibility of cross-domain knowledge transfer. Li et al. (2020) further used VGG-16 fine-tuning and data augmentation to achieve autonomous rock recognition under few-shot conditions. Attention mechanisms were then introduced to improve robustness in complex rover scenes. MRSCAtt incorporates channel and spatial attention modules, enhancing classification performance under variable illumination and cluttered backgrounds, can achieve a reported accuracy of 81.53% on the MSL Surface Data set (
Chakravarthy et al., 2021).
Reducing annotation dependence has become a major direction for
in situ image classification. A semi-supervised contrastive learning framework jointly optimizes supervised inter-class loss and unsupervised similarity loss, helping to mitigate the failure of contrastive learning caused by redundancy in Martian image data (
Wang et al., 2023). Self-supervised distillation further reduces the need for manual labels. In this framework, a teacher model is pre-trained through contrastive learning, and its representations are transferred to a lightweight student network through knowledge distillation. This provides a promising solution for onboard edge deployment under limited annotation and computing resources (
Goh et al., 2023).
Recent work has also explored new backbone architectures for stronger scene representation. Vision Mamba has been introduced for Mars image classification, where non-causal state-space duality is used to capture global context efficiently and improve generalization on surface imagery (
Dai and Cui, 2025a). A Mamba-based zero-shot scene classification framework further extracts semantic and visual features using Mamba and VMamba, respectively, and applies cross-modal attention to infer unseen classes. Using sentence vectors, the framework reported overall accuracy (OA) values of 41.34%, 59.69%, and 77.77% under the 6/4, 7/3, and 8/2 seen/unseen class ratios, respectively, representing improvements of 2.28%, 4.75%, and 4.15% over the teacher model under the same experimental settings (
Tan et al., 2025). This opens a new direction for autonomous scene understanding when annotated examples are absent or extremely limited
Unsupervised clustering and novelty detection complement supervised classification by discovering structure in large unlabeled rover image archives. Deep clustering jointly optimizes iterative K-means and CNN feature learning, enabling geologically meaningful visual clusters to be identified from massive Mastcam images (
Parente and Panambur, 2020). Deep constrained cluster-ing introduces spatial proximity and depth priors to guide clustering toward geological structures rather than surface appearance differences. On the Curiosity rover data set with 150 clusters, the method increased the proportion of homogeneous clusters by 16.7%, reduced the Davies–Bouldin index from 3.86 to 1.82, and improving retrieval accuracy from 86.71% to 89.86% (
Panambur and Parente, 2025).
Onboard deployment also requires classification models to be compact and interpretable. Binarized neural networks employ extreme quantization for onboard screening of anomalous surface features (
Hollen and John, 2025). An explainability framework based on integrated gradients guides structured pruning by quantifying neuronal contributions, thereby compressing models while preserving decision transparency (
Lundstrom et al., 2022). These studies show that classification performance alone is insufficient; model size, inference efficiency, and interpretability are equally important for rover applications.
Mars in situ image classification has evolved from transfer-learning-based scene recognition to a broader framework involving attention-enhanced robustness, label-efficient learning, unsupervised discovery, novel backbone architectures, and resource-aware deployment. Current methods can support rapid scene assessment and data prioritization. However, challenges remain in rare-class recognition, zero-shot generalization, cross-mission transfer, and reliable onboard inference. Future progress will likely depend on self-supervised pre-training, uncertainty-aware classification, compact foundation models, and interpretable decision mechanisms for autonomous rover operations.
3.2.4 Depth estimation and 3D reconstruction: from 2D perception to 3D scene understanding
Depth estimation and 3D reconstruction enable Mars rovers to recover spatial geometry from two-dimensional images, supporting autonomous environmental perception and precise manipulation. Semantic segmentation and object detection provide information in the image plane. Depth estimation adds geometric distance, while 3D reconstruction integrates multiple observations into spatial representations. Together, these techniques extend rover perception from 2D recognition to 3D scene understanding. Mars in situ 3D perception faces several constraints. Stereo data may be unavailable or incomplete. Surface textures can be weak or repetitive. Illumination can vary strongly with local time, terrain geometry, and dust conditions. Under these conditions, recent studies have explored monocular depth estimation, multi-task joint learning, semantic-guided reconstruction, and visual localization.
Monocular depth estimation provides a practical route when stereo vision is limited.
Ma et al. (2024) proposed a Y-shaped dual-task network that uses monocular depth estimation as an auxiliary branch. The model injects 3D spatial cues into the semantic segmentation backbone through a depth-aware spatial attention module, improving rock boundary and small-target recognition without requiring additional stereo data.
Tian et al. (2024) further developed a lightweight 3D semantic recon-struction framework. It integrates real-time semantic segmentation with a depth generation model to recover semantic 3D point clouds from monocular close-up images. This framework forms an end-to-end perception pipeline from 2D pixels to 3D scenes, with reported performance of 84.0% mIoU for segmentation, an absolute relative depth error of 0.367, and an overall perception speed of 9.5 FPS.
For long-distance traverses, visual localization and reconstruction must remain reliable across sites and terrain types.
Kou et al. (2025) proposed a self-supervised keypoint extraction network with multi-scale deformable structures and self-attention descriptors to improve feature repeatability and discriminability in weak-texture scenes. Compared with the ASIFT-based Zhurong rover localization method, this approach reduced localization error by 12.5% and improved localization robustness by 50.8%.
Li and Wu (2025) incorporated Transformer-based semantic cues into the entire feature matching pipeline, improving reliability in texture-poor regions and introducing semantic edge constraints to preserve terrain discontinuities for high-fidelity 3D reconstruction.
At a higher representation level, 3D semantic mapping links geometric reconstruction with terrain understanding.
Chiodini et al. (2020) fused semantic segmentation with stereo depth estimation and projected pixel-level class labels onto voxel occupancy maps, improving terrain evaluation efficiency. In a simulated environment,
Liu et al. (2025d) developed a semantic SLAM system based on a digital twin platform. The system combines a Siamese Transformer with the Hough transform and couples rock instance segmentation with 3D reconstruction in real time. The system reported an average segmentation accuracy of 87.86% at 57.34 FPS, a sub-0.1 m absolute trajectory error, and a mean relative rock-counting error of 8.63%.
Mars in situ depth estimation and 3D reconstruction are moving from standalone geometric recovery toward semantic–geometric joint modeling. Multi-task learning connects depth and semantics understanding, semantic priors enhance geometric matching in weak-texture scenes, and 3D semantic maps support terrain evaluation and navigation. The central goal is to build rover perception systems that understand not only what is present in a scene, but also where objects are located in 3D space. Future progress will depend on lightweight monocular depth estimation, robust semantic matching, uncertainty-aware reconstruction, and efficient 3D mapping under limited onboard computing resources.
3.2.5 Image enhancement and synthetic data generation: addressing extreme imaging conditions
Mars in situ images are often degraded by dust scattering, strong illumination variations, sensor noise, and communication-related compression. These effects can reduce contrast, blur local details, and introduce blocking or ringing artifacts. At the same time, the scarcity of labeled Martian surface data limits the training of supervised models. Image enhancement and synthetic data generation address these two challenges from complementary directions. Image enhancement improves the usability of real mission images, whereas synthetic data generation expands the training and testing space when labeled Martian data are limited. In this context, enhancement methods can improve the reliability of downstream perception from real mission images, whereas synthetic data generation can support pre-mission training and robustness testing under limited Martian labels.
Image restoration methods have been developed to mitigate dust, scattering, and noise.
Ye et al. (2024) incorporated an atmospheric scattering model into a generative adversarial network. The model estimates depth, scattering coefficients, and transmission maps through dual branches for dust synthesis and removal, while physical consistency constraints improve restoration realism. In the reported evaluation, the method achieved a standard deviation of 42.292, an average gradient of 4.951, and an FID of 215.529; the incorporation of the atmospheric scattering model reducing FID from 228.753 to 215.529 compared with vanilla Cycle-GAN.
Xiang and Ye (2023) further integrated a scattering model into a self-supervised cycle-consistency framework, enabling domain translation between unpaired dust storm and clean images while reducing color distortion and artificial edge enhancement. For photon noise in multispectral images,
Li et al. (2024) generated a quality mask through implicit noise mapping and used a lightweight attention network for adaptive single-image denoising.
Other enhancement studies focus on compression artifacts and contrast improvement.
Liu et al. (2025a) proposed a semantically guided quality enhancement network to reduce blocking artifacts caused by satellite-to-ground compression. The method exploits the high semantic similarity of Martian images through a reference block transfer module and demonstrates cross-scene generalization.
Sathish et al. (2026) developed a contrast enhancement framework that combines visual sensitivity transformation with a slime mold optimizer. It adaptively adjusts parameters using no-reference quality metrics to balance contrast improvement and information fidelity. On 100 low-quality Martian images, the method reported a contrast improvement ratio of 1.21 ± 0.09, a low lightness order error of 2.35 ± 3.86, and a sparse feature fidelity of 0.97 ± 0.01.
Synthetic data generation provides another way to overcome annotation scarcity and test model robustness.
Jiang et al. (2023) developed a modular terrain simulator that generates scenes with semantic and depth ground truth by randomizing height maps, rock distributions, and textures.
Zhang and Li (2024) combined generative AI with a physics-based rendering engine in a UAV-oriented simulation system. A diffusion model is used to generate high-fidelity 3D rocks and materials for closed-loop validation of large-scale scene perception algorithms.
Attaoui and Pastore (2025) further developed a GAN-based testing toolset that automatically discovers realistic failure cases, which can be used for ground-based robustness evaluation and iterative optimization of segmentation models.
In summary, image enhancement and synthetic data generation form an important support layer for Mars rover vision. Enhancement methods address image degradation caused by dust, noise, illumination, and compression. Synthetic data generation mitigates annotation scarcity and enables controlled testing under extreme conditions. The two directions are increasingly connected through physical modeling, generative AI, and digital twin technologies. Together, they provide a stronger data foundation for reliable rover perception in the Martian environment. Beyond in situ scenarios, image enhancement is also important for orbital remote sensing, where super-resolution techniques can recover high-resolution spatial details from lower-resolution observations (
Li et al., 2025d).
4 Special requirements of mars exploration missions for deep learning
Mars exploration missions require intelligent perception and decision-making under severely constrained operating conditions. These constraints distinguish Mars-oriented deep learning from most Earth-based applications. Terrestrial deep learning systems usually benefit from large annotated data sets, high-performance computing, and stable communication infrastructure. Mars missions do not. Deep learning systems for Mars exploration must therefore address six structural constraints.
1) Onboard computing, storage, and communication resources are highly limited, requiring efficient edge inference.
2) High-quality annotated data are scarce, requiring data-efficient learning rather than large-scale supervised training.
3) Expert annotation is expensive and time-consuming, motivating self-supervised and active learning strategies.
4) Exploration data arrive continuously, while the observed environment changes over time; models must therefore support online adaptation and incremental learning.
5) Strong distribution shifts exist between Earth and Mars, and also among different Martian regions and sensors, requiring transfer learning and domain adaptation.
6) Single-sensor perception is incomplete, requiring multimodal fusion of heterogeneous data sources.
These constraints are interconnected. Limited onboard resources drive lightweight model design and compression. Data scarcity and annotation cost motivate few-shot, self-supervised, and active learning. Streaming data require online updating, while domain shifts demand transfer learning and adaptation. Multimodal fusion improves robustness by combining complementary information from different sensors. The following sections discuss these mission-specific requirements and the corresponding technical responses.
4.1 Storage and computing constraints: edge computing and lightweight deployment
Mars rovers and landers must perform perception, planning, and decision-making under strict on-board resource constraints. Although they collect high-resolution local observations, long communication latency and limited downlink bandwidth prevent all raw data from being transmitted to Earth for centralized processing. Edge computing, defined here as onboard processing on resource-constrained devices close to the data sources, is therefore essential for near-real-time interpretation during autonomous traversal, obstacle avoidance, target selection, and in situ measurements. As shown in Fig. 7, the edge computing workflow of Mars in situ explorers during autonomous traversal and in situ measurements can be divided into two stages: edge-based intelligent perception and edge-based real-time decision-making. Both stages must operate reliably under the harsh Martian environment and within limited storage, computational, and energy budgets. Consequently, edge intelligence has become a core technical direction for Mars in situ exploration. In this context, model design must balance accuracy, inference speed, robustness, memory footprint, energy consumption, and computational cost, rather than optimizing prediction accuracy alone.
Current studies mainly follow two technical paths: lightweight network architecture design and model compression. For lightweight design, Light4Mars introduces lightweight Transformer modules to reduce parameter count while maintaining acceptable segmentation accuracy, enabling deployment on rover onboard computing platforms (
Xiong et al., 2024). RockNet employs a cross-dimensional feature interaction mechanism to replace conventional downsampling and upsampling operations, reducing computational overhead while preserving segmentation performance (
Wei et al., 2025). For model compression, knowledge distillation, quantization, and pruning have been explored. For instance, a knowledge distillation framework for zero-shot Mars scene classification transfers knowledge from a large teacher model to a much smaller student model, thereby reducing the onboard inference burden (
Tan et al., 2024).
Mars edge computing differs fundamentally from conventional Internet-of-Things edge scenarios. Terrestrial edge systems can often access cloud resources or receive frequent online updates. In contrast, for Mars, once launched, the hardware and software are largely fixed, and large-scale upgrades are impractical. Lightweight model design must therefore satisfy two requirements at the same time: strong pre-deployment optimization and robust post-deployment performance. Models must be compact, but they must also remain reliable and generalizable during long-term autonomous operation.
Quantization-aware training has recently been introduced into deep-space edge intelligence. It simulates quantization-induced accuracy loss during training, guiding the model to preserve discriminative features under low-bit inference. This allows integer operations to replace floating-point operations during deployment, reducing storage, power consumption, and computational cost (
Paul and Paul, 2025;
Xie et al., 2025a). When combined with pruning, knowledge distillation, and neural architecture search, these techniques can improve the trade-off between model capacity and prediction accuracy. They are especially relevant for tasks such as Martian terrain segmentation and landform classification, where both efficiency and reliability are required.
Therefore, edge deployment for Mars exploration requires hardware–algorithm co-optimization. Model compression alone is not sufficient. Lightweight networks must also be robust to unfamiliar terrain, sensor noise, and environmental variation. End-to-end co-design across task requirements, model architecture, and onboard hardware provides a practical pathway for deploying high-performance deep perception models on resource-constrained Mars platforms.
4.2 Scarcity of annotated data: few-shot learning paradigms
Mars exploration faces severe labeled-data scarcity, often under conditions more constrained than many Earth-based computer vision tasks. Many scientifically valuable geological and geomorphological structures, such as dark dunes, araneiform terrain, and Swiss cheese terrain, exhibit long-tailed occurrence and annotation patterns in available observations and labeled data sets. In this context, “long-tailed” refers to class-frequency imbalance: a small number of common surface or geomorphic classes account for most available observations and annotations, whereas many rare, regionally restricted, seasonally variable, or insufficiently annotated classes contain only limited samples. These tail classes may be globally uncommon, concentrated in specific regions or seasonal environments, or poorly represented in existing annotated data sets. Meanwhile, new missions may encounter unfamiliar landforms, lithologies, or surface textures that were not included during pre-training. Therefore, Mars perception can be regarded as a typical data-limited and open-world problem.
Few-shot learning provides an important response to this challenge. Its goal is to recognize or segment new classes using a small number of labeled examples. For Mars exploration, this capability is especially valuable because new scientific targets cannot always be anticipated before launch, and additional annotations may be difficult to obtain during operations. From a methodological perspective, current few-shot learning approaches mainly involve three paradigms: meta-learning, metric learning, and generative data augmentation.
Meta-learning aims to learn how to adapt. Instead of training a model only for one fixed task, meta-learning trains the model across many related tasks so that it can rapidly adapt to a new task with only a few examples. This is well aligned with Mars exploration, where a rover or orbiter may need to recognize an unfamiliar terrain type from very limited observations. In this setting, the key is not simply memorizing known classes, but learning transferable initialization, optimization rules, or task-level representations. Causal reasoning has recently been introduced into the meta-learning framework to reduce spurious correlations and improve robustness under sparse supervision (
Liu et al., 2025c).
Metric learning provides another practical route for few-shot Mars perception. It learns an embedding space in which similar samples are close to each other and different classes are separated. Classification can then be performed by comparing the distance between query samples and a small support set. This non-parametric decision mechanism is attractive for resource-constrained platforms because new classes can be added without retraining a large classifier. Brownian covariance-based deep metric learning improves embedding discriminability by capturing second-order feature statistics (
Dong et al., 2025). Attention mechanisms further enhance sensitivity to subtle differences among geological or geomorphic targets (
Jia et al., 2025;
Lin et al., 2025).
Generative data augmentation expands the limited training space. Rather than relying only on simple transformations such as rotation or cropping, recent methods use diffusion models to synthesize diverse target appearances and scene contexts. DiffSatSeg uses diffusion priors and parameter-efficient fine-tuning to achieve strong performance with a single target sample (
Li et al., 2025a). Controllable diffusion-based augmentation embeds few-shot targets into diverse remote sensing backgrounds, thereby increasing data coverage and reducing overfitting (
Liu et al., 2025e). These methods are particularly useful when real examples are scarce, but target morphology and environmental context can be simulated or constrained.
In practice, these three paradigms are often complementary. Meta-learning supports rapid task adaptation. Metric learning provides a simple and flexible decision space. Generative augmentation increases sample diversity and improves robustness. Together, they form a multi-layered framework for learning from limited annotated data. The central goal is to extract transferable and discriminative information from very few labeled examples, while avoiding overfitting to rare or mission-specific samples. For Mars exploration, few-shot learning is therefore not only a technical convenience, but also a necessary capability for recognizing unexpected targets in data-limited environments.
4.3 Annotation cost constraints: self-supervised and active learning
Martian images often require expert interpretation by planetary scientists. Target-level labels, geomorphic boundaries, and pixel-level segmentation masks all depend on domain knowledge (Fig. 8). Such annotation is slow, labor-intensive, and difficult to scale (
Wang et al., 2023). Modern orbiters, landers, and rovers continue to produce increasing volumes of observational data. This creates a structural mismatch between data generation and expert annotation capacity. Self-supervised learning and active learning provide two complementary responses to this problem. The former reduces dependence on manual labels, whereas the latter improves the efficiency of limited annotation budgets.
Self-supervised learning constructs supervisory signals directly from unlabeled data. It allows models to learn useful representations before task-specific labels are available. Manual annotation is then required mainly during downstream fine-tuning. Current mainstream paradigms include contrastive learning and masked image modeling. Contrastive learning learns discriminative features by comparing positive and negative sample pairs in an embedding space (
Chen et al., 2020). Masked image modeling trains the model to reconstruct randomly masked image regions and thereby capture structural and semantic relationships (
He et al., 2022). Both paradigms are well suited to Mars exploration, where unlabeled images are far more abundant than expert-labeled samples.
Recent Mars-related studies show the value of self-supervised representation learning across different data types and tasks. A domain-specific Mars model pre-trained on seven million CTX images using self-supervised learning outperforms Earth-domain pre-trained models in downstream classification and retrieval tasks (
Fang et al., 2025a). For noisy CRISM data, self-supervised learning constructs training signals from noisy observations themselves to reconstruct spectra, avoiding the need for expert-provided noise-free labels (
Wang et al., 2025c). In hyperspectral image classification, graph contrastive learning improves performance under few-shot conditions (
Xi et al., 2025). For onboard vision, CLOVER combines contrastive pre-training and self-distillation to reduce model size while preserving representation quality, making it more suitable for resource-constrained deployment (
Vincent et al., 2024).
Active learning addresses the same problem from a different direction. Instead of learning only from unlabeled data, it optimizes which samples should be annotated. During training, the model identifies samples that are most informative, uncertain, diverse, or likely to change the decision boundary, and submits them for expert annotation (
Ma, 2024). The goal is to obtain the largest performance gain from the smallest number of annotated samples. This is particularly relevant for Mars exploration, where both expert time and communication bandwidth are limited.
Active learning has two practical values in Mars missions. First, it can help prioritize which images should be downlinked or reviewed by experts. Under limited communication resources, a rover or orbiter could select images with the highest expected model-improvement value instead of transmitting redundant observations. Second, active learning can synergize with weakly supervised annotation strategies. For Martian landform semantic segmentation,
Ye et al. (2025) replaced dense pixel-level masks with sparse scribble annotations and used pseudo-labeling to obtain high-quality segmentation results with limited expert annotation. The combination of active querying and weak annotation reduces the burden of frame-by-frame and pixel-by-pixel inter-pretation while still preserving useful supervision.
Self-supervised learning and active learning are complementary rather than competing approaches. Self-supervised learning reduces the need for labels by extracting representations from unlabeled data. Active learning increases the value of each labeled sample by selecting the most informative examples. Together, they address annotation cost from two sides: reducing annotation dependence and improving annotation efficiency.
4.4 Streaming data processing: online incremental learning mechanisms
Data acquisition in Mars exploration is characterized by two defining features: streaming data return and dynamic environmental conditions. Rovers continuously collect observations along their traverse and transmit data to Earth in batches over time. Meanwhile, seasonal dust storms, diurnal frost cycles, and newly encountered landforms can shift the data distribution. Under these conditions, the conventional workflow of offline training followed by fixed-parameter deployment may not maintain stable performance over extended mission durations (Fig. 9). Models must be able to incorporate new information while retaining previously learned knowledge.
Online incremental learning, a continual learning paradigm designed for streaming data and distribution shifts, is well-suited to Mars exploration. Their goal is to update models as new data arrive, while avoiding catastrophic forgetting of earlier knowledge. The central challenge is the stability-plasticity dilemma. A model must remain plastic enough to adapt to new terrain, illumination, or sensor conditions, but stable enough to preserve recognition of previously learned targets (
Minhas et al., 2025). This issue is particularly important for Mars exploration. A rover may encounter unfamiliar lithologies or terrain textures in a new region, but it must still recognize common objects and landforms learned earlier. Current continual learning methods fall into three broad categories: regularization-based methods, replay-based methods, and parameter-isolation methods (
Vanaparthi et al., 2026). Regularization-based methods impose importance-weighted constraints during parameter updates to suppress excessive modification of critical weights. Elastic Weight Consolidation (EWC) estimates parameter importance using the Fisher information matrix and penalizes large changes to important weights, thereby reducing forgetting without storing historical samples (
Kirkpatrick et al., 2017). This strategy is memory efficient, but its effectiveness depends on accurate importance estimation.
Replay-based methods preserve past knowledge by revisiting earlier samples during new training stages. They maintain a limited memory buffer and mix previous samples with incoming new data. Replay can be implemented using stored raw samples or generative replay (
Bidaki et al., 2025). A reservoir-sampling-based buffer provides a practical mechanism for selecting representative samples from a data stream under limited memory (
Buzzega et al., 2020). Replay methods are generally effective, but they require additional storage and careful buffer management.
Parameter-isolation methods reduce interference by assigning different parameter subspaces to different tasks or data distributions. This structural separation can help preserve old knowledge while learning new patterns (
Yildirim et al., 2024). However, as tasks accumulate, model size may continue to grow linearly, which can strain onboard memory and computation (
Jiang et al., 2025). This makes direct application to Mars platforms difficult unless combined with compression or architecture-sharing strategies.
In practice, hybrid strategies are commonly adopted. For example, regularization can be combined with a small replay buffer to improve both memory retention and adaptability (
Baysal and Bayılmış, 2025). Few-shot class-incremental learning has also been applied to Mars hyperspectral image classification, demonstrating the feasibility of updating class knowledge under limited annotations (
Zhang et al., 2025).
Notably, deploying online incremental learning in Mars missions must also respect strict onboard computing and storage constraints. Regularization methods have relatively small memory footprints, but their gradient computations still introduce overhead. Replay methods require storage for representative samples. Parameter-isolation methods may increase model size as new tasks accumulate. Therefore, continual learning for Mars exploration cannot be considered separately from lightweight design, model compression, and edge hardware constraints. The key challenge is to balance adaptability, memory retention, and resource efficiency during long-term autonomous operation.
4.5 Data distribution shift: cross-domain transfer and domain adaptation
Deep learning has achieved considerable success in Earth-based applications, but its use in Mars exploration remains limited by the scarcity of high-quality annotated data. Existing Earth-domain models provide a valuable transfer foundation, but they cannot usually be deployed directly in Martian scenarios. The main obstacle is distribution shift. Mars images differ from Earth images in illumination, atmospheric conditions, surface texture, target morphology, color statistics, and observation geometry. These differences can substantially reduce model performance. Figure 10 illustrates the data distribution shift between the Earth research domain and the Mars exploration domain, highlighting the need for cross-domain transfer and domain adaptation.
Cross-domain transfer learning aims to reuse transferable knowledge from a source domain (e.g., Earth) to improve learning efficiency and generalization in a target domain (e.g., Mars) (
Zhao et al., 2025). In Mars exploration, the source domain may be Earth remote sensing, lunar imagery, simulated planetary scenes, or data from previous Mars missions. A common strategy is stepwise transfer. A model is first pre-trained on large-scale Earth remote sensing data to learn generic visual representations. It is then adapted to a medium-scale Mars data set for feature alignment. Finally, it can be transferred to a few-shot target data set for a specific perception or decision-making task. This multi-level reuse strategy can reduce the training burden caused by limited Martian annotations. However, it assumes that source and target feature distributions can be progressively aligned, which may not always hold when the domain gap is large.
The conventional pre-training–fine-tuning paradigm remains useful for Mars vision tasks. In this paradigm, a model is first pre-trained on large-scale source-domain data, such as ImageNet, Earth remote-sensing imagery, or other generic visual data sets, and is then fine-tuned with a limited number of labeled Mars samples for downstream tasks such as crater detection, terrain segmentation, or geological and geomorphological classification. This strategy can improve model initialization and task adaptation when labeled Mars data are scarce, but it remains limited. Fine-tuning updates model parameters using a small number of target-domain labels, but it does not explicitly enforce feature-distribution alignment between the source and target domains. As a result, a fine-tuned model may still generalize poorly when illumination, texture, imaging geometry, sensor characteristics or target morphology differs substantially across domains. Domain adaptation provides a more direct response to this problem. Its core objective is to learn domain-invariant representations so that a classifier trained on a labeled source domain can generalize to an under-annotated or unlabeled target domain. Depending on the availability of target-domain labels, domain adaptation can be supervised, semi-supervised, or unsupervised. For Mars exploration, unsupervised domain adaptation is particularly important because expert annotation is costly and often unavailable.
Current domain adaptation methods mainly follow two technical routes. The first is adversarial feature alignment, in which a feature extractor and a domain discriminator are trained against each other to reduce source-target discrepancy. The second is self-training with pseudo-labels, in which high-confidence target predictions are iteratively used as surrogate labels to refine the decision boundary. Both strategies attempt to use unlabeled target-domain data more effectively, but they remain sensitive to noisy pseudo-labels, unstable alignment, and large semantic gaps between domains.
In Mars exploration, domain adaptation has been tested across several perception tasks. For terrain segmentation, UDAFormer uses a Transformer-based unsupervised framework with teacher-student knowledge distillation and output-guided biased sampling to mitigate domain shifts among Mars data sets (
Liu et al., 2023d). For crater detection, a cross-attention-guided multi-level domain adaptation network transfers knowledge from lunar craters to the Martian domain via image-level and instance-level feature alignment. This shows that cross-planet transfer can help reduce dependence on Martian annotations. The feasibility of transferring knowledge from the Moon to Mars has also been validated for geological landmark detection (
Alshehhi and Gebhardt, 2022b). For Mars image classification, the cross-domain performance of dual-attention networks, knowledge transfer networks, and Vision Transformers has been systematically compared, providing an empirical basis for selecting suitable architectures under domain shift (
Kale et al., 2022).
A further challenge is that the Martian target domain itself is not stationary. The true difficulty is not only transferring models from Earth to Mars, but also transferring them from one Martian environment to another. Different landing sites, seasons, imaging times, sensors, and terrain units can all introduce sub-domain shifts. As discussed in Section 4.4, rover observations also evolve along the traverse. Therefore, cross-domain transfer should be connected with online and continual learning. Consequently, A practical Mars perception model must not only bridge the Earth-Mars or Moon-Mars gap, but also adapt to persistent distribution shifts within Mars. This requirement imposes additional constraints on model robustness, update efficiency, and onboard computational resources.
4.6 Incomplete multi-source information: multimodal collaborative perception
Mars orbiters, landers, and rovers carry diverse payloads, including optical cameras, LiDAR, thermal infrared imagers, ground-penetrating radar, and spectrometers. Each sensor observes only part of the Martian environment. Optical images provide rich texture and color information, but they can be affected by shadows, dust, and weak surface contrast. LiDAR point clouds provide geometric structure, but they may lack semantic detail. Thermal infrared data can distinguish surface units with similar visual appearances but different physical properties. Radar can reveal subsurface structures that are invisible to cameras. Because each modality provides only partial information, multimodal perception is essential for building a more robust and comprehensive representation of the Martian environmental.
Multimodal fusion methods are typically categorized into data-level, feature-level, and decision-level fusion. Data-level fusion preserves the richest original information, but it requires accurate calibration, registration, and temporal synchronization among sensors. Decision-level fusion is flexible and easier to implement, but it may lose detailed inter-modal relationships. Feature-level fusion achieves a favorable balance between information fidelity and system complexity and has become the dominant paradigm in multimodal deep learning. Within feature-level fusion, cross-modal attention is particularly important. Cross-modal attention models dependencies among modalities through self-attention or cross-attention, allowing the network to adaptively select informative sources and determine when and how different modalities should be fused. However, many fusion methods implicitly assume that all modalities are available and synchronized. This assumption may not hold in practical Mars missions, where sensor inputs can be missing, asynchronous, or partially degraded.
In Mars exploration tasks, multimodal fusion has shown clear advantages. For crater identification, a CNN-Transformer cross-modal adaptive feature fusion network learns deep correlations between imagery and digital elevation models through self-attention. This improves multi-type and multi-scale crater detection by adaptively weighting different feature sources (
Yang et al., 2025). For rover depth estimation, M3Depth uses a wavelet transform to address sparse Martian textures and introduces a depth-normal consistency constraint. This iterative dual-modal refinement improves depth reliability in weakly textured regions (
Li et al., 2025c). For terrain traversability prediction, a self-supervised approach fuses camera and LiDAR data to generate bird’s-eye-view cost maps and uses inertial measurement unit signals for training. This enables robust multimodal alignment without manual annotations (
Xie et al., 2025b). At the orbital level, MOMO, the first multi-sensor Mars foundation model, integrates representations from HiRISE, CTX, and THEMIS and outperforms single-sensor models in downstream segmentation tasks (
Purohit et al., 2026b).
In summary, multimodal collaborative perception addresses the incompleteness of single-sensor observations through information complementarity. It strengthens terrain understanding, target recognition, depth estimation, traversability prediction, and large-scale mapping. However, several challenges remain. Different modalities often have different resolutions, noise characteristics, coverage areas, and acquisition geometries. Accurate registration and synchronization are difficult, especially for moving rovers and heterogeneous payloads. In addition, missing modality represent a practical engineering problem rather than merely a data incompleteness issue. Sensor failures, communication interruptions, power-management-induced shutdowns, limited downlink bandwidth, or asynchronous acquisition may make optical, LiDAR, thermal, radar, or spectroscopic data unavailable during certain mission phases. Future Mars perception systems will therefore need robust multimodal fusion strategies that can operate under misalignment, uncertainty, and incomplete sensor input. Promising directions include modality-dropout training, reliability- or uncertainty-aware modality weighting, cross-modal feature completion, and graceful degradation from multimodal fusion to single-modality inference when only partial sensor data are available. Such systems may form a key technical foundation for reliable autonomous perception under the resource and environmental constraints of Mars exploration.
5 Current challenges and technical limitations
Deep learning has advanced considerably in Mars exploration, covering both orbital remote sensing and in situ perception. However, moving these methods from controlled research settings to open, dynamic, and resource-constrained mission environments remains difficult. The main limitations are not isolated. They span data quality, annotation availability, model reliability, computational resources, domain generalization, and physical interpretability. Figure 11 summarizes the major application challenges and technical limitations of deep learning in Mars exploration missions. This section examines these challenges from six perspectives: data heterogeneity, annotation scarcity, model uncertainty, computational constraints, generalization shift, and decoupling from physical causality. Together, they define a problem-oriented framework for evaluating future research directions.
5.1 Data heterogeneity: cross-scale matching across multi-source payloads
Mars exploration has accumulated large volumes of multi-payload and multi-modal data from both orbital and in situ platforms. These data provide complementary information, but they are also highly heterogeneous. Differences in spatial resolution, spectral dimension, viewing geometry, temporal coverage, and sensor noise make it difficult to integrate information across instruments and scales. This challenge is directly reflected in existing Mars mission data sets and is therefore not merely a generic remote-sensing problem. This heterogeneity is one of the fundamental obstacles to reliable multi-source perception in Mars exploration.
Orbital remote sensing data exhibit heterogeneity in several forms. The first is spatial-scale mismatch. HiRISE delivers submeter-scale imagery, whereas CTX and THEMIS provide imagery at ~6 m/pixel and tens-to-hundreds of meters per pixel, respectively (
McEwen et al., 2007). This large resolution gap complicates cross-scale feature alignment. Forced resampling may introduce aliasing, blur small features, or semantic dilution, thereby impeding unified multi-scale landform mapping. The second is spectral-spatial mismatch. CRISM provides hundreds of hyperspectral bands and strong mineral diagnostic capability, but its spatial resolution is much coarser than that of HiRISE. HiRISE captures detailed morphology but lacks broad spectral coverage. This dimensional mismatch means naïve fusion (e.g., channel concatenation) seldom suffices, and the lack of synchronized, co-located acquisitions yields few pixel-level paired labels, severely constraining supervised fusion. The third is temporal inconsistency. Differences in solar elevation, dust opacity, season, and atmospheric conditions can produce strong variations in albedo, shadow, and contrast. Even after precise geometric registration, models must still learn illumination- and time-invariant representations, which remains difficult without large-scale paired annotations.
In situ exploration faces a different but equally complex form of heterogeneity. Rovers carry navigation and hazard-avoidance cameras, LiDAR, thermal infrared imagers, and multispectral cameras, each differing markedly in spatial resolution, field of view, data rate, and spectral dimension. Navigation cameras provide broad scene context but limited spectral information. Multispectral cameras provide compositional clues but usually cover narrower fields and require multiple pointing. LiDAR provides geometric measurements but produces sparse point clouds. These sensing asymmetries make direct feature alignment difficult across rover payloads. Residual misalignment caused by imperfect calibration or temporal synchronization can further affect local tasks such as rock identification, soil classification, traversability analysis, and robotic arm targeting. These limitations are particularly relevant to rover operations, where perception outputs must support navigation, instrument targeting, and tactical planning under restricted onboard computation, limited communication windows, and time-sensitive operational constraints.
Several technical pathways have been explored to address these heterogeneity challenges. For cross-resolution fusion, implicit neural representations can reconstruct features at arbitrary resolutions through continuous field functions, offering a possible framework for multi-resolution alignment (
Chen et al., 2021). For multimodal alignment, cross-modal contrastive learning maps paired samples into a shared embedding space and learns modality-invariant representations without requiring dense pixel-level labels (
Radford et al., 2021). For temporal variability, domain adaptation and style transfer have been used to reduce the influence of illumination, season, and atmospheric differences on landform recognition. These approaches provide useful methodological references for handling heterogeneous Mars data. However, many of them were originally developed or validated in broader computer vision or Earth remote-sensing contexts, and their Mars-specific effectiveness still requires systematic testing on multi-sensor and multi-temporal Mars data sets. Therefore, these methods should be regarded as promising methodological directions rather than fully established solutions for Mars exploration.
Nevertheless, two major limitations remain. First, many fusion and alignment methods assume sufficient paired or co-registered data across modalities. In Mars exploration, such correspondences are often sparse, incomplete, or expensive to construct. Pixel-level or region-level pairing usually requires both accurate geometric registration and expert validation. This assumption is therefore difficult to satisfy in many real mission data sets. Second, advanced fusion models often introduce high computational overhead. Implicit representations, cross-attention modules, and multimodal Transformers can be expensive in memory and computation. This limits their direct use in onboard edge environments. A key challenge is therefore to balance fusion performance, model complexity, and mission-level resource constraints. Future work should more explicitly evaluate whether the performance gains from multimodal or cross-scale fusion on Mars data sets justify the additional costs in annotation, registration, expert validation, and onboard computation.
5.2 Annotation scarcity: massive observations and sparse supervisory signals
As discussed in Sections 4.2 and 4.3, the scarcity of annotated data and the resulting cost constraints pose a structural limitation for deep learning in Mars exploration. At the challenge level, the key issue is the mismatch between rapidly accumulating observations and slowly produced expert labels. Mars orbiters and rovers continuously generate orbital and in situ data, whereas the number of experts capable of reliable planetary annotation remains limited. This creates a data-rich but label-poor regime. Large volumes of images are available, but only a small fraction can be converted into high-quality supervisory signals for model training.
This limitation is not only quantitative but also qualitative. Annotating Mars images requires domain expertise. For semantic segmentation, experts must delineate geological units, landforms, rocks, shadows, and terrain classes at the pixel level. High-resolution images can require substantial manual effort. More importantly, Martian targets often show strong intra-class variability and inter-class ambiguity. The same landform may look different under different illumination, dust coverage, viewing geometries, or degradation states. Conversely, different classes may appear visually similar. As a result, annotation uncertainty is partly irreducible. Different experts may produce inconsistent boundaries for the same region, and even the same expert may annotate ambiguous targets differently at different times. Therefore, the core challenge is not simply the small number of labels, but the scarcity of labels that are accurate, consistent, and geologically meaningful.
Long-tailed class distributions further intensify the problem. Many scientifically important targets occupy only small portions of images or occur in limited regions. Recurring slope linear, fresh impact craters, unusual aeolian features, and rare lithological units may be far less frequent than background terrain. Random sampling therefore tends to allocate annotation effort to common or uninformative regions. Targeted sampling can increase the representation of rare classes, but it may still fail to capture their full morphological diversity. This makes model training vulnerable to overfitting, class imbalance, and poor generalization.
Self-supervised learning and active learning provide useful mitigation strategies, but they do not fully solve the problem. Self-supervised learning depends on the assumption that pre-training data contain transferable visual structures. When Martian images differ strongly from Earth images in illumination, texture, dust effects, and landform morphology, Earth-domain pre-training may not generalize well to Mars-specific tasks (
Yu et al., 2023;
Fang et al., 2026). Active learning can improve annotation efficiency by selecting informative samples, but it is also affected by class imbalance and semantic ambiguity. In highly imbalanced data sets, query strategies may still select background-dominated or visually obvious samples. In geologically ambiguous regions, uncertainty may reflect real interpretive ambiguity rather than model weakness, making it difficult to decide which samples should be annotated (
Venguswamy et al., 2021). Annotation scarcity is therefore not merely a data engineering problem. This challenge is coupled with data heterogeneity, long-tailed distributions, semantic ambiguity, model uncertainty, and computational constraints. Here, long-tailed class distribution refers to situations in which a few common classes dominate the available observations and annotations, where many rare classes contain only limited samples. Future solutions will need more than simply increasing the number of labels. They should combine expert-in-the-loop annotation, uncertainty-aware sampling, weak and sparse supervision, self-supervised pre-training on Mars-domain data, and label-quality assessment. Such integrated strategies are necessary for building reliable Mars perception models under sparse and imperfect supervision.
5.3 Inherent model deficiencies: uncertainty and hallucination in open environments
The application of deep learning models in Mars exploration is constrained not only by data limitations, but also by the behavior of the models themselves in open and unstructured environments. Mars missions operate under conditions that cannot be fully represented in training data sets. Rovers may encounter unfamiliar landforms, unusual illumination, dust-degraded images, rare surface events, or sensor artifacts, yet must make safe decisions without real-time ground intervention. In such settings, two model-level risks become particularly important: prediction uncertainty and hallucination-like outputs. These risks can affect both mission safety and scientific reliability.
Prediction uncertainty refers to the mismatch between model confidence and actual correctness. It becomes especially serious for out-of-distribution, noisy, or ambiguous inputs. Uncertainty can arise from two sources. Aleatoric uncertainty is caused by inherent observation noise, such as dust scattering, shadows, low contrast, or compression artifacts. Epistemic uncertainty reflects the model’s lack of knowledge about samples outside its training distribution. Standard deep networks often do not recognize their own knowledge boundaries. They may therefore produce confident but incorrect predictions for unseen terrain types or degraded images. For rover operations, such overconfident errors are risky. A hazardous region may be classified as safe, or a traversable surface may be incorrectly treated as an obstacle, directly affecting navigation and target selection.
Hallucination is most relevant to multimodal large models and generative models. Vision–language models may introduce Earth-specific concepts, such as “grassland” or “pavement”, into descriptions of Martian scenes, leading to cross-planet semantic confusion (
Li et al., 2023). Generative models used for depth completion, image restoration, or 3D reconstruction may fill poorly observed regions with visually plausible but physically unrealistic structures. In these cases, the output may appear coherent but lack geological or physical validity. This is different from ordinary classification error because the model can generate additional content that is not supported by the observation.
Several methods can mitigate these risks. Bayesian deep learning, ensemble inference, confidence calibration, and out-of-distribution detection can help quantify epistemic uncertainty and reduce overconfident predictions. Retrieval-augmented generation can constrain vision–language outputs by grounding them in external databases or mission-specific knowledge. However, these methods also introduce practical costs. Ensemble models increase storage and inference time. A key challenge is therefore to make uncertainty estimation and hallucination suppression both reliable and lightweight. Mars perception systems need to know when they are uncertain, when to defer decisions, and when to request human review or additional observations. For safety-critical tasks, model confidence should not be treated as equivalent to correctness. Only with these safeguards can deep learning models be used more reliably in open Martian environments.
5.4 Computational constraints: the accuracy-power trade-off in onboard inference
As discussed in Section 4.1, Mars rovers operate under strict constraints in storage, computing capacity, and energy supply. These constraints are intrinsic to deep-space robotic missions. Historically, radiation-hardened onboard processors have provided substantially lower computational performance than contemporary commercial processors because radiation tolerance, long-term reliability, and qualification cycles are prioritized over peak throughput. For instance, the RAD6000 on the Opportunity rover delivered roughly two orders of magnitude less performance than 2013 ground workstations, making onboard deep learning infeasible at that time (
Burl et al., 2013). Although the RAD750 on the Perseverance rover provides higher performance than the RAD6000, its capacity remains limited for modern deep neural networks used in object detection, semantic segmentation, and 3D perception.
However, the onboard computing landscape is evolving. Recent developments include NASA’s HPSC / Microchip PIC64-HPSC processors for higher-performance and fault-tolerant spaceflight computing (
Powell, 2018;
Wakelin and Ganry, 2025), ISS-based tests of commercial-off-the-shelf (COTS) AI processors such as Qualcomm Snapdragon and Intel Movidius (
Swope et al., 2023), radiation-tolerant FPGA and SoC FPGA platforms such as Microchip RT PolarFire® (
Toguchi et al., 2024), and ESA onboard-AI activities such as Φsat-2 and FPG-AI (
Melega et al., 2023; Department of Information Engineering, 2024). These advances raise the feasible ceiling of onboard intelligence, but their use in Mars missions still depends on radiation qualification, fault mitigation, thermal stability, power-aware scheduling, software maturity, and mission-level reliability assessment.
This computational gap is therefore not simply an engineering inconvenience. It reflects a structural mismatch between the deep-space operating environment and the rapidly increasing complexity of deep learning models. Radiation-hardened processors require long development cycles and conservative designs. Their computational density is limited by radiation-tolerant architecture and reliability requirements. At the same time, rovers operate under tightly constrained power budgets, whether powered by solar panels or radioisotope systems. Dedicated neural network accelerators are being explored, but many remain in validation or early deployment stages for space applications (
Pavel et al., 2026). As a result, Mars deep learning systems must balance three competing goals: accuracy, inference speed, and power consumption, while also satisfying mission reliability requirements.
1) Model compression has diminishing returns. Pruning, quantization, knowledge distillation, and low-rank decomposition can reduce model size and computation. However, accuracy often degrades rapidly beyond a critical compression threshold. Excessive pruning may remove features needed for detailed landform recognition. Low-bit quantization can introduce errors in low-contrast or shadowed regions. Distillation may fail when the capacity gap between teacher and student models is too large (
Tan et al., 2024;
Song et al., 2025). Mission-critical perception requires a minimum level of reliability, leaving residual computational overhead that must still fit within the onboard budget.
2) Real-time operation does not always match hardware-efficient computation. Rover navigation and obstacle avoidance require low-latency inference, especially for semantic segmentation and hazard detection. In practice, this often means single-image or small-batch inference. However, many accelerators achieve their highest efficiency under large-batch processing. This mismatch can reduce actual hardware utilization far below nominal peak performance. The Martian thermal environment adds another constraint. Large temperature variations limit aggressive cooling strategies and complicate dynamic frequency scaling, which may affect stable long-duration inference.
3) Computational reuse across multiple tasks remains difficult. An autonomous rover perception system may need to perform terrain segmentation, obstacle detection, target recognition, depth estimation, and anomaly screening at the same time. Deploying independent models for each task can lead to a near-linear increase in memory and computation. Multi-task learning with a shared backbone offers a possible solution, but it can suffer from gradient conflicts and capacity competition among tasks, reducing the performance of individual outputs (
Vandenhende et al., 2021). Efficient multi-task reuse within a compact network is therefore a major system-level challenge.
Onboard computation is one of the key limiting factors for deploying deep learning in Mars exploration. This limitation cannot be solved by hardware improvement or algorithm compression alone. Although recent onboard computing platforms provide more promising options than legacy rover processors, they do not remove the fundamental accuracy-power-reliability trade-off in deep-space missions. A practical solution requires co-design across task requirements, model architecture, compression strategy, and onboard hardware capability. The central question is therefore not simply how to make models smaller, but how to preserve mission-critical reliability under severe resource constraints. Future systems will require hardware-aware neural architecture search, adaptive inference, shared multi-task backbones, uncertainty-aware early exiting, and joint optimization of accuracy, latency, memory, and power consumption.
5.5 Generalization bottleneck: geological and instrumental domain drift
Beyond the Earth-to-Mars distribution shift discussed in Section 4.5, deep learning models also face substantial generalization challenges within the Martian domain itself. These challenges arise from both geological domain drift and instrument-level domain drift. Geological domain drift refers to distribution changes caused by differences in landing sites, surface materials, geomorphic units, and local environmental conditions. Instrumental domain drift refers to distribution changes introduced by differences in sensor design, calibration procedures, spatial resolution, spectral response, compression, and imaging geometry. Martian landing sites differ strongly in geological setting, surface texture, mineral composition, and geomorphic evolution. Jezero crater, for example, contains deltaic deposits and sedimentary units related to ancient fluvial activity (
Mangold and Caravaca, 2025), whereas Gale crater preserves lacustrine strata with distinct stratigraphic and mineralogical characteristics (
Hurowitz et al., 2017). Even within a single landing site, a rover may move from flat sandy terrain to fractured bedrock, layered outcrops, or rugged foothills. Each traverse segment may therefore introduce new visual, textural, and semantic distributions that differ from the data used for model training or validation.
Instrument-level drift adds another layer of complexity. Different missions use cameras with different optical designs, radiometric calibration procedures, compression strategies, fields of view, and spatial resolutions. As a result, the same type of landform may appear differently across missions or instruments. A model trained on one region, rover, or payload may therefore experience reduced reliability when applied to another. This limitation is especially important for cross-mission learning, where data from multiple rovers or orbiters are combined to improve model training.
The core difficulty is that the direction and magnitude of domain drift are difficult to predict before deployment. Before landing, only orbital reconnaissance and limited prior knowledge are available. These data cannot fully capture the local texture, rock abundance, illumination conditions, surface roughness, or small-scale geomorphic diversity encountered during rover operations. When a rover enters a new region that differs from the training distribution, the visual cues learned by the model may no longer be reliable. Texture, shape, color, shadow patterns, and contextual priors can all shift at the same time. Without ground-truth labels or sufficient in situ validation data, the model may not recognize that it has crossed its generalization boundary and may continue to produce confident but unreliable predictions. This concern is consistent with the general out-of-distribution problem in machine learning, but its practical manifestation in Mars exploration is driven by site-specific geology, limited labeled samples, and restricted opportunities for in situ model validation.
Instrument-level domain drift arises from inherent differences among the imaging systems of different payloads. Orbiter cameras such as HiRISE, CTX, and the Color and Stereo Surface Imaging System (CaSSIS) differ in modulation transfer function, noise characteristics, viewing geometry, and compression level. These instrumental differences can produce apparent variations that are comparable to, or even larger than, natural surface variations. When a feature extractor trained on one instrument is transferred to another, performance may decline because of frequency-response differences, receptive-field mismatch, and changes in image statistics. However, the magnitude of this effect remains task- and data set-dependent and should be evaluated through explicit cross-instrument benchmarks. Current domain adaptation and style transfer methods often assume paired images, shared classes, or sufficient target-domain samples. Such assumptions are difficult to satisfy across different Mars missions, instruments, and landing sites.
This problem reveals a central paradox in applying deep learning to Mars exploration. High model performance usually assumes that training and testing data follow similar distributions. Mars exploration, however, is designed to enter unknown and potentially out-of-distribution environments. The real challenge is therefore not only to improve average accuracy on benchmark data sets, but also to detect when a model is leaving its reliable operating domain. Instead of assuming perfect generalization, Mars perception models should be able to recognize domain shift, report uncertainty, and trigger adaptation when new geological or instrumental conditions are encountered. Future studies should therefore report not only within-domain accuracy, but also cross-site, cross-instrument, and low-label adaptation performance to clarify the Mars-specific evidence for generalization.
5.6 Weak coupling with physical causality: the semantic gap between data-driven models and physical laws
Deep learning methods used in Mars exploration mainly rely on statistical association mining and pattern recognition. They learn correlations between input data and output labels, but they do not necessarily learn the physical mechanisms that generate the observed features. In contrast, a central goal of planetary science is to understand the origin and evolution of geological phenomena. Planetary interpretation asks not only “what is this feature,” but also “why/how did it form.” This difference creates a semantic gap between data-driven prediction and physical causal interpretation. In this review, this gap should be considered as a limitation of current data-driven perception models for scientific interpretation, rather than evidence that deep learning is generally unreliable for Mars applications. Model outputs may be accurate at the label level but still weakly connected to geological processes, material properties, and environmental evolution.
This gap first appears in the limited capacity of current models for genetic reasoning. Deep neural networks can identify geomorphic units, such as craters and dunes, with increasing accuracy (
DeLatte et al., 2019a). However, their discriminative basis may not follow the scientific reasoning chain that involves formation mechanisms and material composition. Geological interpretation often depends on morphology, stratigraphic context, material composition, spatial relationships, and formation processes (
Jaumann et al., 2024;
McNeil et al., 2025). A neural network may instead rely on statistically convenient cues such as color, brightness contrast, texture, or annotation bias. For instance, illumination-induced shadows and shadow–highlight boundaries may be misidentified as crater rims under varying illumination or viewing geometries (
Urbach and Stepinski, 2009;
DeLatte et al., 2019a;
Emami et al., 2019). Models may also group features with different origins when those features share superficial visual textures. Such statistical shortcuts can fail when illumination, viewing geometry, dust cover, or weathering state changes. For planetary science, models that output labels without causal or genetic context have limited explanatory value. This issue is particularly important when model outputs are used to support hypotheses about landform formation, surface processes, or habitability-related environments, rather than only to produce perception labels.
A deeper limitation is the lack of explicit physical constraints during model training and inference (
Karniadakis et al., 2021;
Meng et al., 2025). Martian geological features are controlled by physical processes. Aeolian dune morphology is related to wind transport and sediment supply. Impact crater morphology reflects excavation, ejecta emplacement, degradation, and target properties. Dike or fracture patterns may be influenced by stress fields, magma migration, and mechanical layering. Purely data-driven models do not directly encode these process constraints. As a result, they may learn surface correlations without being able to assess whether a predicted pattern is physically plausible. When encountering rare or out-of-distribution landforms, such models cannot reason from first principles or provide a mechanistic explanation for their predictions. The concern is therefore not that data-driven models are ineffective for perception, but that their predictions may be insufficient for standalone scientific explanation without additional physical or geological constraints.
Weak coupling with causality also affects representation learning and sample selection. Self-supervised learning often relies on visual consistency between pre-training and target data (
Fang et al., 2026). However, common data augmentation operations may not always be physically meaningful for Mars imagery (
Koßmann et al., 2022). Inappropriate geometric transformations can distort diagnostic morphology, while arbitrary color perturbations may violate radiometric or spectral relationship of the Martian surface. Without physical constraints, self-supervised learning may learn features that are statistically stable but geologically uninformative. Active learning faces a related problem. If uncertainty is measured only statistically, the model may prioritize samples that are visual outliers but not scientifically important, or overlook samples that are geologically significant but visually subtle (
Chen et al., 2024). Physical priors are therefore needed to guide both representation learning and annotation strategies. At present, however, this remains more of a research direction than a mature Mars-specific solution, because systematic evaluations of physically constrained learning on large Mars benchmarks are still limited.
Bridging this gap requires moving from purely data-driven recognition toward physically informed learning (
Azari et al., 2020). Possible directions include incorporating geomorphic rules, radiative transfer constraints, shape priors, stratigraphic relationships, mineralogical knowledge, and process-based simulations into model design or evaluation. Physics-informed loss functions, causal representation learning, hybrid simulation–learning frameworks, and expert-in-the-loop validation may help constrain model predictions within scientifically plausible ranges (
Kumari et al., 2025). The goal is not to replace data-driven learning, but to make data-driven-learning more consistent with planetary processes and more useful for scientific reasoning.
Therefore, the weak coupling between deep learning and physical causality is a fundamental limitation for Mars exploration. It prevents models from becoming fully trustworthy scientific instruments, even when their benchmark accuracy is high. Until statistical learning is connected with physical reasoning, deep learning will remain primarily an effective perception tool rather than a reliable mechanism for scientific interpretation. A more realistic near-term goal is to develop perception models that can report uncertainty, expose physically relevant evidence, and support expert geological validation.
6 Future research trends and open problems
The preceding sections reviewed the core tasks, mission-specific requirements, and current limitations of deep learning for Mars intelligent perception. Most existing applications still remain at the stage of visual recognition, classification, segmentation, or detection. These methods can extract useful patterns from observational data, but they have not yet fully supported higher-level scientific reasoning, such as autonomously discovering, explaining, and testing hypotheses about Martian surface processes and planetary evolution. The challenges discussed in Section 5, including computational constraints, annotation scarcity, model uncertainty, limited generalization, and weak coupling with physical causality, share a common root. Current methods still rely mainly on statistical correlations in data, while physical priors from planetary geology, geomorphology, radiative transfer, and surface processes are not yet fully incorporated as constraints or guidance.
A central future direction is therefore to move from data-driven perception toward knowledge-guided understanding. Recent advances in multimodal large models, self-supervised representation learning, and embodied intelligence provide new opportunities for this transition (
Dobrea et al., 2025;
Caldas et al., 2026;
Fang et al., 2026;
Heldmann et al., 2026). For Mars exploration, these developments should not simply be imported from Earth-based computer vision. They need to be redesigned for planetary data, mission constraints, and scientific reasoning. Key future directions include Mars-specific foundation models, embodied autonomous scientific exploration, physical causal modeling, uncertainty-aware inference, and multimodal collaborative perception with onboard edge intelligence. Collectively, these directions may provide important building blocks for future intelligent Mars exploration systems and may help connect visual perception with scientific interpretation.
6.1 Mars foundation models: from general pre-training to domain-specific representations
Foundation models have made rapid progress in Earth observation (
Cong et al., 2022;
Sun et al., 2022;
Xiao et al., 2025a), but their application to Mars remote sensing and in situ perception remains at an early stage. A major lesson from recent studies is that general pre-training is not sufficient for Mars. Models specifically designed for, or adapted to, Martian data have been reported to achieve better performance than models pre-trained solely on ImageNet or Earth observation data sets in several downstream tasks, including crater detection, terrain segmentation, and landform classification (
Purohit et al., 2024;
Fang et al., 2025b;
Fang et al., 2025a). The multi-sensor foundation model MOMO, which integrates HiRISE, CTX, and THEMIS data, further demonstrates the value of sensor-specific and domain-specific representation learning, achieving stronger performance than general pre-trained models across multiple tasks (
Purohit et al., 2026b). Mars-Bench also shows that the lack of unified benchmarks has limited objective comparison and systematic progress in this field (
Purohit et al., 2026a).
Future Mars foundation models should be more than large visual backbones. MarsRetrieval, the first cross-modal geospatial retrieval benchmark for Mars, represents an important step in this direction. It systematically assesses the cross-modal alignment of dual-tower encoders and multimodal large language models through paired image-text retrieval, landform retrieval, and global geolocalization. Its results show that general foundation models often fail to capture subtle Martian geomorphic differences, highlighting the need for domain-specific fine-tuning and Mars-aware representation learning (
Wang et al., 2026). Similarly, a ViT model self-supervised on seven million CTX images outperforms Earth-centric models in image retrieval and geological pattern discovery (
Fang et al., 2025b).
Future research should therefore focus on three directions. 1) Developing unified representation learning frameworks for multi-source Martian data; 2) expanding Mars-domain self-supervised pre-training so that massive unlabeled observations can be converted into robust and transferable representations; and 3) establishing standardized benchmarks and evaluation protocols that test not only accuracy, but also cross-sensor transfer, cross-region generalization, uncertainty calibration, and scientific retrieval capability. Collectively, these advances may support the development of Mars foundation models from generic visual pre-training toward domain-specific scientific analysis and provide a possible pathway toward future scientific discovery.
6.2 Embodied intelligence and autonomous scientific discovery: from passive perception to active exploration
Mars rover operations still rely heavily on ground-based planning and teleoperation. Most high-level commands are generated on Earth, uplinked to the rover, and then executed under strict safety constraints. This workflow is reliable, but it limits exploration efficiency because communication delays and limited downlink capacity slow the perception-decision-action cycle. A major future direction is therefore to move from passive perception to active exploration, where onboard systems can perceive, reason, plan, and act with greater autonomy.
Embodied intelligence provides a framework for this transition. In this context, the rover is not only a data-collection platform, but also an active agent that interacts with the Martian environment. It must connect visual perception, terrain understanding, scientific target evaluation, path planning, and robotic manipulation into a closed-loop system. Recent demonstrations suggest that large models may support parts of this workflow. NASA has already employed the Claude large language model to plan surface driving routes for the Perseverance rover, providing preliminary validation of large language models for Mars autonomous driving (
Varley, 2026). In autonomous scientific discovery, the ARTPS system integrates monocular depth estimation, anomaly detection, and a learnable curiosity score, enabling the rover to autonomously identify and prioritize high-value scientific targets without human intervention (
Baydemir, 2025). Multi-agent reinforcement learning has been applied to collaborative exploration in a simulated Jezero crater environment, demonstrating notable potential for improving regional coverage and mission robustness (
Swinton et al., 2025).
Future research should therefore focus on three directions. 1) Developing Mars-oriented embodied foundation models that unify representation and co-optimization of perception, reasoning, planning, and control modules; 2) embedding curiosity-driven active learning into onboard decision-making, establishing an autonomous exploration cycle of “learning while exploring and discovering while learning”; and 3) developing multi-agent collaborative exploration architectures that improve regional coverage, efficiency, and mission robustness. The central challenge is to balance autonomy with safety under deep-space latency, limited computing resources, and uncertain environmental conditions.
6.3 Physical causal modeling and uncertainty quantification: from correlation to causation
As discussed in Section 5.6, deep learning models are effective at capturing statistical correlations, but they often struggle to distinguish physical causality from observational artifacts. Illumination, viewing geometry, atmospheric dust, sensor noise, and compression artifacts can all become entangled with genuine geological signals. Recent Mars-related studies have shown that atmospheric optical depth, photometric geometry, dust redistribution, spectral mixing, and sensor noise can affect surface albedo, reflectance, spectral features, and mineralogical interpretation (
Hess et al., 2022;
Vicente‐Retortillo et al., 2023;
Wang et al., 2025c). Although this specific failure mode has not yet been systematically demonstrated for a particular deep-learning model in Mars applications, it represents a plausible risk: a model trained only on statistical associations may misattribute albedo variations caused by solar elevation or viewing geometry to mineralogical differences, or contrast changes caused by atmospheric dust or dust redistribution to surface texture variations. Such potential errors can introduce systematic bias into geological interpretation.
A key future direction is to make Mars deep learning models more physically informed. Causally informed pre-training provides one possible route. By explicitly modeling relationships among environmental variables during representation learning, such methods may encourage models to learn features that are less sensitive to external factors such as illumination, season, and observation geometry. Although current evidence mainly comes from Earth remote sensing, these methods may provide useful methodological insights for Mars applications, including crater detection, terrain and geomorphological feature segmentation, landform classification, and rover traversability analysis under limited labeled data. Prior work in remote sensing has shown that causally informed pre-training can improve downstream performance and few-shot generalization (
Ravirathinam et al., 2024).
Future research can therefore focus on three directions. 1) Extending physics-informed neural networks and physically constrained learning to Mars perception tasks by incorporating radiative transfer, geomorphic evolution, and surface-process constraints (
Raissi et al., 2019;
Azari et al., 2020); 2) developing causal discovery methods that infer potential links between landform features, material properties, and geological processes from multi-source observations (
Ravirathinam et al., 2024); and 3) establishing uncertainty quantification standards for Mars intelligent perception, including calibrated confidence, reliability diagrams, out-of-distribution detection, and uncertainty-aware decision thresholds. The central open problem remains how to effectively integrate sparse physical priors with massive observational data within deep models, so that model outputs remain consistent with fundamental physical principles while preserving representational flexibility and predictive performance.
6.4 Multimodal collaborative perception and edge intelligence: from data silos to holistic fusion
Mars exploration data are inherently multi-source, including orbital high-resolution imagery, surface navigation images, multispectral observations, radar sounding data, and in situ measurements. At present, however, many processing pipelines remain separated by sensor type, mission platform, or task objective. This creates information silos and limits comprehensive environmental understanding. Multimodal collaborative perception aims to overcome this limitation by fusing complementary information across sensors, platforms, and scales (
Yang et al., 2025, 2026).
Future research should move from task-specific fusion toward integrated multimodal representation and onboard edge intelligence. Key directions include: 1) developing unified multimodal fusion architectures for joint representation and collaborative reasoning across orbital and surface domains, thereby removing inter-modality barriers; 2) compressing multimodal models through quantization, knowledge distillation, pruning, and neural architecture search so that they can be deployed onboard without substantial loss of reliability; and 3) exploring hardware-algorithm co-design paradigms for edge intelligence, enabling perception, uncertainty estimation, online learning, and decision support under strict limits in energy, computation, storage, and communication. The progression from data silos to integrated fusion, and from centralized ground processing to autonomous edge intelligence, will be a central trajectory of next-generation intelligent Mars exploration.
7 Conclusions
Deep learning is becoming an increasingly important technical component for Mars exploration, as mission data continue to grow in volume, diversity, and complexity. This review has summarized recent progress along the logical chain of data, tasks, mission requirements, current challenges, and future research trends. From orbital remote sensing to rover-based in situ exploration, deep learning methods have been applied to object detection, image segmentation, classification, change detection, depth estimation, 3D reconstruction, image enhancement, and synthetic data generation. These applications have improved the efficiency of crater detection, geological unit mapping, terrain assessment, high-value target localization, and dynamic surface monitoring. Deep learning also provides new technical support for linking large-scale orbital reconnaissance with local rover perception.
Mars exploration differs fundamentally from conventional Earth-based computer vision. Mars perception systems must operate under limited onboard storage, computing capacity, power supply, and communication bandwidth. Mars perception systems must also handle scarce annotations, streaming data return, domain shifts across planets, regions, instruments, and missions, and incomplete information from individual sensors. In response to these mission constraints, lightweight architectures, model compression, few-shot learning, self-supervised learning, active learning, domain adaptation, continual learning, and multimodal fusion have been explored. Together, these approaches provide feasible pathways for improving both ground-based data interpretation and future onboard intelligent processing.
Despite substantial progress, several limitations remain unresolved. Mars data are highly heterogeneous across spatial scales, spectral dimensions, imaging geometries, and temporal conditions. High-quality annotations remain sparse, costly, and sometimes inconsistent because of geological ambiguity. Deep learning models may produce overconfident errors under out-of-distribution conditions, while generative and multimodal models may generate outputs that are visually plausible but lack physical support. Onboard deployment remains constrained by trade-offs among accuracy, latency, memory, and power consumption. Model generalization is also vulnerable to geological and instrumental domain drift when models are transferred across landing sites, sensors, missions, or rover platforms. A deeper limitation lies in the weak coupling between data-driven prediction and physical causality. Current models can often recognize what is present in an image, but reliable explanation of why a geological feature formed or whether a prediction is physically plausible remains difficult.
Future research should therefore move beyond task-specific perception toward Mars-oriented intelligent understanding. Domain-specific Mars foundation models trained on large-scale unlabeled orbital and in situ data may provide transferable representations across sensors, regions, and tasks. Embodied intelligence may enable rovers to shift from passive data collection to active scientific exploration, with onboard systems capable of perception, reasoning, planning, and adaptive observation. Physically informed learning, causal modeling, and uncertainty quantification will be essential for making model outputs more interpretable and scientifically reliable. Multimodal collaborative perception and edge intelligence will further support integrated reasoning across optical, spectral, topographic, radar, and in situ measurements under mission-level resource constraints.
In summary, deep learning has shown increasing potential to support Mars remote sensing and in situ image analysis by reducing reliance on purely manual and delayed interpretation and enabling more automated, adaptive, and knowledge-guided workflows. The future value of deep learning in Mars exploration will depend not only on benchmark accuracy, but also on robustness, computational efficiency, uncertainty awareness, physical consistency, and scientific interpretability. Integrating computer vision, onboard computing, planetary science knowledge, and mission operations will be critical for building trustworthy intelligent perception systems for future Mars exploration.