The development of artificial intelligence (AI) techniques has brought revolutionary changes across various realms. In particular, the use of AI-assisted methods to accelerate chemical research has become a popular and rapidly growing trend, leading to numerous groundbreaking works. In this paper, we provide a comprehensive review of current AI techniques in chemistry from a computational perspective, considering various aspects in the design of methods. We begin by discussing the characteristics of data from diverse sources, followed by an overview of various representation methods. Next, we review existing models for several topical tasks in the field, and conclude by highlighting some key challenges that warrant further attention.
Scientific research faces high costs and inefficiencies with traditional methods, but the rise of deep learning and large language models (LLMs) offers innovative solutions. This survey reviews transformer-based LLM applications across scientific fields such as biology, medicine, chemistry, and meteorology, underscoring their role in advancing research. However, the continuous expansion of model size has led to significant memory demands, hindering further development and application of LLMs for science. This survey systematically reviews and categorizes memory-efficient pre-training techniques for large-scale transformers, including algorithm-level, system-level, and hardware-software co-optimization. Taking AlphaFold 2 as an example, we demonstrate how tailored memory optimization methods reduce storage needs while preserving prediction accuracy. By bridging model efficiency and scientific application needs, we hope to provide insights for scalable and cost-effective LLM training in AI for science.
With the development of deep neural networks and differentiable rendering techniques, neural rendering methods, represented by Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have made significant progress. NeRF represents a 3D scene by encoding the appearance and geometry of the scene through neural networks, which are conditioned on both position and viewpoint. In contrast, 3DGS models the scene with a set of Gaussian ellipsoids, allowing for efficient rendering through the rasterization of these ellipsoids into images. However, both two methods are limited to representing static scenes. The rendering and reconstruction of dynamic scenes are critical in virtual reality and computer graphics. As such, extending neural rendering methods from static to dynamic scenes has become an important area of research. This survey organizes dynamic scene rendering methods based on NeRF and 3DGS and categorizes them according to different motion representations. Furthermore, it highlights the relevant applications of dynamic scene rendering, such as autonomous driving, digital humans, and 4D generation. Finally, we summarize the development of dynamic scene rendering and discuss the remaining limitations and open challenges.
Time series forecasting plays a critical role in numerous real-world applications, such as finance, healthcare, transportation, and scientific computing. In recent years, deep learning has become a powerful tool for modeling complex temporal patterns and improving forecasting accuracy. This survey provides an overview of recent deep learning approaches for time series forecasting, involving various architectures including RNNs, CNNs, GNNs, transformers, large language models, MLP-based models, and diffusion models. We first identify key challenges in the field, such as temporal dependency, efficiency, and cross-variable dependency, which drive the development of forecasting techniques. Then, the general advantages and limitations of each architecture are discussed to contextualize their adaptation in time series forecasting. Furthermore, we highlight promising design trends like multi-scale modeling, decomposition, and frequency-domain techniques, which are shaping the future of the field. This paper serves as a compact reference for researchers and practitioners seeking to understand the current landscape and future trajectory of deep learning in time series forecasting.
With the expansion of data availability, machine learning (ML) has achieved remarkable breakthroughs in both academia and industry. However, imbalanced data distributions are prevalent in various types of raw data and severely hinder the performance of ML by biasing the decision-making processes. To deepen the understanding of imbalanced data and facilitate the related research and applications, this survey systematically analyzes various real-world data formats and concludes existing researches for different data formats into four distinct categories: data re-balancing, feature representation, training strategy, and ensemble learning. This structured analysis helps researchers comprehensively understand the pervasive nature of imbalance across diverse data formats, thereby paving a clearer path toward achieving specific research goals. We provide an overview of relevant open-source libraries, spotlight current challenges, and offer novel insights aimed at fostering future advancements in this critical area of study.
Recently, the reasoning capabilities of Large Reasoning Models (LRMs), such as DeepSeek-R1, have witnessed significant advancements through computationally intensive “slow thinking” processes. These models have demonstrated impressive performance across a variety of complex reasoning tasks. However, despite their remarkable success, LRMs come with substantial computational demands that pose considerable challenges in terms of resource consumption, scalability, and accessibility. In contrast, Small Reasoning Models (SRMs), which are often distilled from larger models, offer a more efficient alternative while still achieving competitive performance. Beyond their efficiency, SRMs frequently exhibit distinct capabilities and cognitive trajectories compared with their larger counterparts, making them particularly interesting from both practical and theoretical perspectives. In this work, we provide a timely and comprehensive survey of recently published research focused on SRMs. We first review the current landscape of SRMs. Then, we analyze diverse training paradigms and inference techniques tailored to enhance the reasoning capabilities of SRMs. Furthermore, we offer an extensive review of domain-specific applications where SRMs have been effectively leveraged. Finally, we discuss promising future research directions that aim to bridge existing gaps. By consolidating recent advances, this survey serves as an essential reference for researchers and practitioners interested in leveraging or developing SRMs to unlock advanced reasoning functionalities with improved efficiency.
With the growing prevalence of data-driven decision-making, Text-to-SQL has emerged as a promising solution to lower the barrier to data access by translating natural language queries into executable SQL statements, thereby enhancing user interaction with databases. Despite notable progress driven by deep learning and large language models, significant challenges persist in handling complex queries. This paper presents a comprehensive review of the Text-to-SQL task, structured around two core stages: natural language understanding and natural language translation. Methods are categorized along the technical evolution trajectory into four types: rule-based, machine learning-based, pre-trained language model-based, and large language model-based approaches. Unlike previous surveys, which focus on specific techniques or partial aspects of Text-to-SQL, our work offers a two-stage analytical framework, highlights the impact of large models, and provides a comparative analysis of limitations and trade-offs. Through detailed examination of accuracy, generalization, expressiveness, and computational cost, this survey presents insights into the advantages and disadvantages of each paradigm. Furthermore, the paper summarizes key benchmark datasets and evaluation metrics, and discusses directions to improve the robustness, security, and effectiveness of the existing Text-to-SQL systems.
Generating high-quality 3D assets is a fundamental challenge in computer vision and graphics. While the field has progressed significantly from early VAE/GAN approaches through diffusion models and large reconstruction models, persistent limitations hinder widespread application. Specifically, achieving high geometric and appearance fidelity, intuitive user control, versatile multi-modal conditioning, and directly usable outputs (e.g., structured meshes) remains challenging for established paradigms. This paper surveys the evolution of deep generative models for 3D content creation, with a primary focus on emerging paradigms: autoregressive (AR) generation and Agent-driven approaches, poised to address aforementioned shortcomings. AR models generate assets sequentially (e.g., token-by-token or part-by-part), offering inherent potential for finer control, structured outputs, and integrating user guidance during the step-by-step process. Agent-driven methods, conversely, leverage the reasoning and linguistic capabilities of Large Language Models (LLMs), enabling intuitive and flexible 3D creation by decomposing complex tasks and utilizing external tools through multi-agent systems. We provide a comprehensive overview of these novel techniques, discuss their potential advantages over current methods, and outline key challenges and future directions towards more capable and intelligent 3D generation systems.
While large language models (LLMs) like ChatGPT have shown impressive capabilities in Natural Language Processing (NLP) tasks, a systematic investigation of their potential in this field remains largely unexplored. This study aims to address this gap by exploring the following questions. (1) How are LLMs currently applied to NLP tasks in the literature? (2) Have traditional NLP tasks already been solved with LLMs? (3) What is the future of the LLMs for NLP? To answer these questions, we take the first step to provide a comprehensive overview of LLMs in NLP. Specifically, we first introduce a unified taxonomy including (1) parameter-frozen paradigm and (2) parameter-tuning paradigm to offer a unified perspective for understanding the current progress of LLMs in NLP. Furthermore, we summarize the new frontiers and the corresponding challenges, aiming to inspire further groundbreaking advancements. We hope this work offers valuable insights into {the potential and limitations} of LLMs, while also serving as a practical guide for building effective LLMs in NLP.
Quantum Software Engineering (QSE) has emerged as a research direction practiced by tech joints. Quantum developers face challenges in optimizing quantum computing and QSE concepts. Developers use Stack Overflow to discuss quantum challenges using specialized tags to label posts. These tags often refer to technical quantum aspects. Categorizing quantum practitioners’ questions by concept can help identify common QSE challenges. We conducted studies to classify quantum developers’ questions into various challenges. We extracted 2,829 developers questions from Q&A platforms using quantum-related tags. The posts were analyzed to identify frequent quantum-related challenges and develop a novel grounded theory. The challenges identified include Tooling, Theoretical, Learning, Conceptual, Errors, and API Usage. Through content analysis and grounded theory, the developers’ discussions were annotated with commonly reported quantum challenges to develop a ground truth dataset. ChatGPT was used to validate human annotations and resolve disagreements. Various fine-tuned transformer algorithms, including BERT, DistilBERT, and RoBERTa, were used to classify developer discussions into commonly reported quantum challenges. We achieved an average accuracy of 95% with BERT DistilBERT algorithms, compared to fine-tuned Deep and Machine Learning (D&ML) classifiers, including Feedforward Neural Networks (FNN), Convolutional Neural Networks (CNN), and Long Short-Term Memory networks (LSTM), which achieved accuracies of 89%, 86%, and 84%, respectively. The proposed Transformer-based approach outperforms the previous D&ML-based approach with a 6% increase in accuracy by processing actual developer discussions, i.e., without data augmentation. Furthermore, we applied SHAP (SHapley Additive exPlanations) to provide model interpretability, revealing how specific linguistic features drive predictions and enhancing transparency in the classification process. These improved research findings can help quantum vendors and developers’ discussion forums to better organize developers’ discussions for improved access and readability.
Accurate and efficient multivariate time series (MTS) analysis is increasingly critical for a wide range of intelligent applications, including traffic forecasting, anomaly detection for industrial maintenance, and trajectory classification for health monitoring. Within this realm, Transformers have emerged as the predominant architecture due to their strong ability to capture pairwise dependencies. However, Transformer-based models suffer from quadratic computational complexity and high memory overhead, limiting their scalability and practical deployment for long-term, large-scale MTS modeling. Recently, Mamba has emerged as a promising linear-time alternative with high expressiveness. Nevertheless, directly applying vanilla Mamba to MTS remains suboptimal due to three key limitations: (i) the lack of explicit cross-variate modeling, (ii) difficulty in disentangling the entangled intra-series temporal dynamics and inter-series interactions, and (iii) insufficient modeling of latent time-lag interaction effects. These issues constrain its effectiveness across diverse MTS tasks. To address these challenges, we propose DeMa, a dual-path Delay-Aware Mamba backbone for efficient and effective MTS analysis. DeMa preserves Mamba’s linear-complexity advantage while substantially improving its suitability for multivariate settings. Specifically, DeMa introduces three key innovations: (i) it decomposes the MTS context into intra-series temporal dynamics and inter-series interactions and learns them via two dedicated paths; (ii) it develops a temporal path with a module to capture long-range dynamics within each series, accommodating variable-length inputs and enabling series-independent, parallel computation while maintaining linear complexity; and (iii) it designs a variate path with a module that integrates delay-aware linear attention to model cross-variate dependencies, enhancing fine-grained, delay-sensitive dependency learning. Extensive experiments on five representative tasks, long- and short-term forecasting, data imputation, anomaly detection, and series classification, demonstrate that DeMa achieves state-of-the-art performance while delivering remarkable computational efficiency.