1. School of Computer Science, Northwestern Polytechnical University, Xi’an 710129, China
2. Department of Computing, The Hong Kong Polytechnic University, Hong Kong 999077, China
wenqifan03@gmail.com
shang@nwpu.edu.cn
Show less
History+
Received
Accepted
Published Online
2025-12-31
2026-04-22
2026-07-07
PDF
(7555KB)
Abstract
Accurate and efficient multivariate time series (MTS) analysis is increasingly critical for a wide range of intelligent applications, including traffic forecasting, anomaly detection for industrial maintenance, and trajectory classification for health monitoring. Within this realm, Transformers have emerged as the predominant architecture due to their strong ability to capture pairwise dependencies. However, Transformer-based models suffer from quadratic computational complexity and high memory overhead, limiting their scalability and practical deployment for long-term, large-scale MTS modeling. Recently, Mamba has emerged as a promising linear-time alternative with high expressiveness. Nevertheless, directly applying vanilla Mamba to MTS remains suboptimal due to three key limitations: (i) the lack of explicit cross-variate modeling, (ii) difficulty in disentangling the entangled intra-series temporal dynamics and inter-series interactions, and (iii) insufficient modeling of latent time-lag interaction effects. These issues constrain its effectiveness across diverse MTS tasks. To address these challenges, we propose DeMa, a dual-path Delay-Aware Mamba backbone for efficient and effective MTS analysis. DeMa preserves Mamba’s linear-complexity advantage while substantially improving its suitability for multivariate settings. Specifically, DeMa introduces three key innovations: (i) it decomposes the MTS context into intra-series temporal dynamics and inter-series interactions and learns them via two dedicated paths; (ii) it develops a temporal path with a module to capture long-range dynamics within each series, accommodating variable-length inputs and enabling series-independent, parallel computation while maintaining linear complexity; and (iii) it designs a variate path with a module that integrates delay-aware linear attention to model cross-variate dependencies, enhancing fine-grained, delay-sensitive dependency learning. Extensive experiments on five representative tasks, long- and short-term forecasting, data imputation, anomaly detection, and series classification, demonstrate that DeMa achieves state-of-the-art performance while delivering remarkable computational efficiency.
Multivariate time series (MTS) are a crucial and fundamental data modality in both industrial and daily scenarios. They are collected at scale by physical and virtual sensors, which continuously record system measurements and encapsulate valuable information about the evolving dynamics of real-world systems [1,2]. Consequently, time series analysis serves as a cornerstone for understanding and predicting the behaviors of complex systems, enabling a wide range of intelligent applications, such as traffic flow forecasting for transportation scheduling [3,4], missing data imputation for web stream processing [5], anomaly detection for maintenance [6,7], and trajectory classification for health monitoring [8].
Due to the complex and non-stationary nature of real-world systems, observed MTS often exhibit intricate and entangled patterns. Along the temporal axis, multiple variations, including long-term trends, seasonal cycles, and irregular fluctuations, are typically mixed and overlapped. Along the variate axis, correlated series interact through latent dependencies, where the evolution of one variable can influence, or be influenced by, others with time delays. Accordingly, a series of deep models have been proposed to capture temporal dependencies and variate dependencies. From the perspective of backbone architectures, existing methods can be broadly categorized into RNN-, CNN-, MLP-, and Transformer-based methods [9–14]. Typically, RNN-based models [15–17] leverage recurrent structures to model temporal state transitions, while CNN-based models [18–20] employ temporal convolutional networks (TCNs) to extract local variation patterns. However, both RNN- and CNN-based methods are constrained by limited effective receptive fields, which hinders their ability to capture long-term dependencies. Lightweight MLP-based models [21,22] utilize stacked fully connected layers to model temporal patterns, where dense connections implicitly capture measurement-free relationships among time points within each variable [13]. These linear forecasters are primarily designed for forecasting tasks and are computationally efficient. Nevertheless, due to their architectural simplicity and limited representational capacity, MLP-based models suffer from an information bottleneck, making it difficult to represent long-range and complex temporal dependencies. Transformer-based models, which adopt attention mechanisms to capture pairwise dependencies with a global receptive field, have become mainstream for MTS modeling [10,13,23–27]. These methods typically employ different tokenization strategies [28–31], such as point-wise tokens [23–26], patch-wise tokens [10,27], and series-wise tokens [13], and then apply attention to model temporal dependencies, variate dependencies, or both, as illustrated in Fig. 1(a)(i–iv). Despite their remarkable performance, Transformer-based models face growing concerns about their [21,32], leading to substantial computational demands and memory overhead. Specifically, the cost of self-attention scales quadratically with the token length, as listed in Fig. 1(b)(i–iv). For an MTS instance with variates, each represented by tokens embedded in dimensions, the complexity of fully self-attention models can reach unacceptably [11,12]. As both and grow, quadratic scaling becomes a major bottleneck for long-term, large-scale MTS modeling, motivating the development of more computationally efficient architectures for MTS analysis.
Recently, Mamba has emerged as a promising backbone for capturing complex dependencies in sequential data [33–35]. By leveraging selective mechanisms and hardware-aware computational designs [34], Mamba attains modeling performance comparable to Transformers while maintaining linear complexity with respect to sequence length. Furthermore, Mamba-2 [34] further improves efficiency by exploiting the structured State Space Duality (SSD) property to reformulate computations into highly parallelizable matrix operations. With high expressiveness enabled by the selective mechanism, efficient training and inference supported by parallel computation, and linear scalability with respect to context length, Mamba offers a compelling alternative to attention-based architectures for sequential modeling. Consequently, an increasing number of Mamba-based models have been proposed across various domains, such as Jamba [36] in natural language processing, Vision Mamba [37] in computer vision, Caduceus [38] in genomics, and SSD4Rec [39] for recommender systems.
Despite the recent success of Mamba-style techniques in enabling efficient MTS analysis, directly applying standard Mamba to multivariate time series often yields suboptimal performance compared with state-of-the-art methods [40]. This gap can be mainly attributed to the following limitations. (1) Difficulty in disentangling entangled temporal and variate contexts: Unlike language tokens, which typically lie in relatively discrete contexts, MTS observations are jointly governed by intra-series temporal dynamics and inter-series interactions. These factors often overlap and intertwine, making it difficult for vanilla Mamba to disentangle meaningful and structured representation. (2) Lack of explicit cross-variate modeling: Mamba was originally developed for language-like, essentially univariate sequences, and therefore lacks explicit mechanisms for capturing dependencies among multiple correlated variates. Prior studies have shown that modeling inter-variable interactions is crucial for effective MTS representation learning [13,27,41]. (3) Insufficient explicit modeling of lag dependencies: Mamba relies on recurrent hidden-state updates, in which each state depends primarily on the immediately preceding timestep. This design may limit its ability to explicitly capture lagged effects that are pervasive in real-world MTS, where changes in one variate may influence another only after a non-negligible delay [3,41–43]. For example, in traffic forecasting, an accident in one region may not immediately affect adjacent areas; instead, its impact often propagates through the network over several minutes. Precisely capturing such propagation delays is therefore crucial for accurate dependency modeling and reliable prediction.
Motivated by the above limitations and the urgent practical demands of long-term, large-scale MTS analysis, we aim to develop a Mamba-based time-series representation backbone that simultaneously captures intricate temporal dependencies and delay-aware cross-variate interactions while preserving linear-time efficiency. Toward this goal, we explicitly decompose the time-series context into (i) intra-series temporal dynamics, which can be learned independently and in parallel for each variate, and (ii) inter-series interactions, which are modeled in a delay-aware manner to capture cross-variate dependencies, as illustrated in Fig. 1(a)(v).
To this end, we propose a dual-path Delay-aware Mamba for efficient multivariate time series analysis, namely DeMa. Specifically, DeMa first adaptively selects task-relevant spectra via an Adaptive Fourier Filter and decomposes the input into a Cross-Time Component and a Cross-Variate Component. These two components are then fed into stacked DuoMNet blocks. Each DuoMNet block contains two parallel paths, each coupled with a scan operator and a lightweight Mamba-based module. Concretely, the Cross-Time Scan serializes each variate along the temporal axis into a 1D sequence, which is processed by Mamba-SSD. This design supports MTS-independent parallel computation across variates and efficiently captures long-range temporal dependencies within each series, with computation scaling linearly with the series length. In parallel, the Cross-Variate Scan reorganizes tokens to emphasize inter-variable interactions at each token step. The scanned sequence is processed by Mamba-DALA, which integrates Delay-Aware Linear Attention (DALA) to explicitly model cross-variate dependencies with both a global correlation delay and a token-level relative delay. As a result, DeMa achieves delay-aware variate interaction modeling while preserving linear-time computation. Finally, we fuse the temporal-path and variate-path representations through a weighted fusion layer and project to task-specific outputs via lightweight heads. To sum up, our major contributions include:
● We propose DeMa, a dual-path Mamba-based backbone for general MTS analysis that jointly captures intra-series temporal dynamics via a temporal path and delay-aware cross-variate dependencies via a variate path. Benefiting from a linear-time state-space backbone, DeMa achieves a favorable trade-off between accuracy and efficiency for long-term and large-scale MTS modeling.
● We introduce Mamba-SSD to model intra-series temporal dependencies within each variate. It accommodates variable-length inputs and enables series-independent, parallel modeling via block-based matrix multiplication, thereby improving scalability and throughput.
● We design Mamba-DALA, which integrates Delay-Aware Linear Attention to explicitly model cross-variate interactions with both global correlation delays and token-level relative delays, thereby enhancing fine-grained, delay-sensitive dependency learning.
● We conduct extensive experiments on five mainstream tasks, including long- and short-term forecasting, imputation, classification, and anomaly detection. The results demonstrate that DeMa achieves consistently strong performance across tasks while significantly reducing training time and GPU memory usage, suggesting its practicality and scalability for real-world large-scale MTS analysis.
The remainder of this paper is structured as follows. Section 2 introduces the preliminaries. Section 3 introduces the proposed model, which is evaluated and discussed in Section 4. Then, Section 5 summarizes the recent development of time series analysis. Finally, conclusions are drawn in Section 6.
2 Preliminaries
2.1 Notations and definitions
We denote a multivariate time series (MTS) within a lookback window as , where denotes the number of variates and denotes the number of time steps. The th variate is represented as , while the multivariate observation at time step is denoted by .
The primary objective of this work is to learn a task-agnostic representation , which can be adapted to diverse downstream tasks through task-specific heads. These tasks include point-level tasks, such as forecasting the future time steps with output , as well as anomaly detection and missing-value imputation on the input window with output . In addition, we consider the sequence-level classification task, which assigns the input MTS to one of categories, with output . Table 1 summarizes the main notations used throughout the paper.
2.2 Mamba
2.2.1 State space model (SSM)
The classical State Space Models (SSMs) describe the state evolution of a linear time-invariant system, which map input signal through implicit latent state , where , , , and indicate the time step, sequence length, channel number of the signal and state size, respectively. These models can be formulated as the following linear ordinary differential equations:
where is the state transition matrix that describes how states change over time, is the input matrix that controls how inputs affect state changes, denotes the output matrix that indicates how outputs are generated based on current states and represents the command coefficient that determines how inputs affect outputs directly. Most SSMs exclude the second term in the observation equation, i.e., set , which corresponds to a skip connection in deep learning models. The time-continuous nature poses challenges for integration into deep learning architectures. To alleviate this issue, most methods utilize the Zero-Order Hold rule [33] to discretize continuous time into intervals, which assumes that the function value remains constant over the interval . Equation (1) can be reformulated as
where and . Discrete SSMs can be interpreted as a combination of CNNs and RNNs. Typically, the model employs a convolutional mode for efficient, parallelizable training and switches to a recurrent mode for efficient autoregressive inference. The formulations in Eq. (2) are equivalent to the following convolution [44]:
Thus, the overall process can be represented as
2.2.2 Selective state space model (Mamba)
The discrete SSMs are based on data-independent parameters, meaning that parameters , , and are time-invariant and the same for any input, limiting their effectiveness in compressing context into a smaller state [33]. Mamba introduces a selective mechanism to selectively retain, propagate, or suppress information according to the current token [33,35]. This directly addresses a major weakness of earlier subquadratic sequence models, namely their limited ability for content-based reasoning. As a result, Mamba preserves a form of context-aware sequence modeling that is much closer in spirit to attention, while avoiding the dense token-to-token interactions of standard self-attention. Specifically, it utilizes linear projection to parameterize the weight matrices , and according to model input , improving the context-aware ability, i.e.,
Then the output sequence can be computed with those input-adaptive discretized parameters as follows:
During training, both computation and memory scale linearly with sequence length, rather than quadratically as in standard Transformers. During autoregressive inference, Mamba only needs to update a recurrent hidden state rather than maintaining an ever-growing KV cache over the full context. Moreover, Mamba introduces a hardware-aware parallel scan algorithm [45], enabling selective and time-varying SSMs to be practically efficient on modern accelerators.
2.2.3 Mamba-2
Mamba-2 [34] introduces a comprehensive framework, Structured State-Space Duality (SSD). Leveraging SSD, it reformulates SSMs as semi-separable matrices via matrix transformation and develops a more hardware-efficient computation method based on block-decomposed matrix multiplication, i.e.,
where denotes the matrix form of SSMs that uses the sequentially semiseparable representation, and , and represent the selective space state matrices associated with input tokens and , respectively. denotes the selective matrix of hidden states corresponding to the input tokens ranging from to . Mamba-2 achieves a 2-8 faster training process than Mamba-1’s parallel associative scan while remaining competitive with Transformers.
3 The proposed method: DeMa
3.1 Structure overview
As shown in Fig. 2 (left), DeMa is designed in a challenge-driven manner to address three key limitations in multivariate time-series (MTS) modeling: the entanglement of temporal and variate contexts, the lack of explicit cross-variate modeling, and the inadequate modeling of lag-aware interactions. Accordingly, the framework consists of three stages: the Adaptive Fourier Filter, stacked DuoMNet Blocks, and Task-specific Projection. The Adaptive Fourier Filter first decomposes the raw MTS into an intra-series temporal component and an inter-series variate component, thereby reducing interference between temporal dynamics and cross-variate interactions before representation learning. Built upon this decomposition, the stacked DuoMNet Blocks address the second challenge through a dual-path design. As shown in Fig. 2 (right), each block contains a temporal path for modeling intra-series dependencies and a variate path for capturing inter-series dependencies, allowing the model to explicitly and separately learn these two types of structure. Furthermore, the variate path incorporates delay-aware interaction modeling, enabling DeMa to capture not only whether variables are correlated, but also how their interactions evolve. Finally, the learned shared representation is passed to lightweight Task-specific Projection layers for downstream tasks such as forecasting, imputation, classification, and anomaly detection.
3.1.1 Adaptive fourier filter (AFF)
Real-world multivariate time series often exhibit overlapping and entangled patterns arising from (i) intra-series temporal dynamics within each variate and (ii) inter-series interactions among correlated variates. In the frequency domain, slowly varying structures such as trends and seasonal cycles are typically concentrated in low-frequency bands. In contrast, short-term variations, including abrupt changes and time-varying cross-variate effects, tend to appear at higher frequencies [46,47]. Meanwhile, different tasks emphasize different spectral contents: numerical prediction tasks (such as forecasting, imputation, and anomaly detection) are usually sensitive to local dynamics, whereas semantic tasks (e.g., classification) rely more on global and long-range patterns [48]. These characteristics motivate a task-adaptive spectral decomposition that separates and highlights task-relevant components.
To facilitate modeling of intra-series dynamics and inter-series interactions, we split the input into two complementary frequency subsets. The globally dominant frequencies capture broadly shared, long-term structures and mainly reflect per-variate temporal dynamics, while the residual frequencies encode more window-specific variations that are often indicative of time-varying cross-variate interactions. Concretely, given an input window with length , we first apply the Fast Fourier Transform (FFT) along the temporal axis for each variate and obtain the averaged amplitude for each frequency index in . We rank by amplitude, select the top fraction as , and denote the complement by . The Adaptive Fourier Filter then decomposes as
where denotes the inverse FFT and retains only the coefficients indexed by the specified set. The Cross-Time Component is dominated by global, low-frequency spectral components that primarily reflect intra-series temporal dynamics, and Cross-Variate Component preserves the remaining variations, typically higher-frequency and window-specific, which are related to inter-series interactions.
This decomposition is particularly valuable when a state-space model serves as the backbone for MTS modeling. In Mamba, the HiPPO-based state-space formulation has been shown to induce approximately orthogonal transformations [49]. However, the orthogonality of the learned transform does not imply that the mixed components in the raw MTS can be cleanly separated. If two underlying components are not orthogonal in the input space, their projections onto any orthogonal basis must share some basis directions. We formalize this limitation below.
Theorem 1 (Non-orthogonality implies basis overlap). Given the input series and let be an orthogonal basis. Suppose
and define the supports and . If and are not orthogonal, i.e., , then . Equivalently, two non-orthogonal components cannot be represented on disjoint subsets of an orthogonal basis.
Proof Since is an orthogonal basis, we have for and for all . Thus,
If , then for every at least one of or is zero, implying for all , which implies
This contradicts the assumption . Therefore, there exists at least one index such that and , i.e., . □
Theorem 1 implies that if the Cross-Time component and the Cross-Variate component are non-orthogonal, their representations on any orthogonal basis must overlap, making a complete separation by an orthogonal transform impossible. In practice, this non-orthogonality often arises because different generative factors are not confined to disjoint frequency bands: for example, trend, seasonality, and interaction-driven fluctuations can partially overlap in the same spectral range and thus superimpose in both the time and frequency domains. Therefore, although HiPPO-based SSMs tend to learn orthogonal transformations [49], applying a single Mamba model directly to raw MTS may still entangle these factors. In contrast, our proposed explicitly decomposes into complementary spectral parts, allowing subsequent modules to model them separately and thereby reducing representational entanglement and avoiding redundant modeling capacity.
3.2 DuoMNet block
In this section, we detail the design of the proposed DuoMNet block, which comprises a temporal path to capture cross-time dependencies and a variate path to learn cross-variate dependencies.
3.2.1 Cross-time scan and cross-variate scan
Given the decomposed Cross-Time Component and Cross-Variate Component , we tokenize each component into patch tokens to reduce point-wise redundancy and encode local temporal semantics [10,27]. Concretely, for each variate , we apply reversible instance normalization (RevIN) [50] and partition the normalized series into contiguous patches of length . A lightweight 1-D convolutional encoder maps each patch to a -dimensional embedding. To enable complementary dependency modeling, we organize the resulting embeddings with two scan orders:
(a) We preserve the temporal order of patch tokens in the Cross-Time Component and stack all variates in parallel:
where contains the patch embeddings of the th variate. This scan order enables the temporal module to model intra-series dependencies along the temporal axis independently for each variate.
(b) We instead group tokens in the Cross-Variate Component by the same patch position and scan across variates:
where collects the embeddings of all variates at the th token step. This scan order preserves the chronological order over token indices and facilitates modeling time-varying cross-variate interactions at different token steps.
These two scanned representations, and , are then fed into subsequent modules to learn cross-time and cross-variate dependencies, respectively.
3.2.2 Cross-time dependencies modeling via Mamba-SSD
Given the temporal scanned embeddings of the Cross-Time component, our goal is to capture intra-series temporal dependencies within each variate. We therefore adopt an strategy: each variate is modeled without cross-variate coupling, enabling efficient parallel computation across variates. The proposed Mamba-SSD module consists of four stages: (a) two-branch gating and local mixing, (b) selective SSM parameterization, (c) Structured State Space Duality (SSD) based parallel scan, and (d) gated output projection.
(a) Two-branch gating and local mixing. We first apply linear projections to obtain a content branch and a gate branch . The gate branch serves as a residual controller that adaptively filters and reweights the content features. We then perform local temporal mixing by applying a lightweight 1-D convolution along the token axis on the content branch (for each variate in parallel), which injects an explicit short-range inductive bias to model local continuity that is common in real-world time series, and complements the long-range selective SSM by providing locally aggregated features that stabilize token-dependent parameter generation.
(b) Selective SSM parameterization. Following Mamba-style selective SSMs, we generate token-dependent parameters from . Let denote the state dimension. The step size and the output projections are computed as
where . To ensure stable dynamics, we parameterize the continuous-time transition with a negative diagonal vector and discretize it using :
where denotes element-wise multiplication. Here, controls the effective memory scale at each token, while and modulate how content is written into and read from the latent state, yielding content-adaptive dynamics.
(c) SSD-based parallel computation. Given the input token sequence , the selective SSM at step can be formulated as follows:
which yields . Equivalently, the induced linear operator is block-diagonal across variates by leveraging Structured State-Space Duality (SSD) properties in Mamba-2 [34]. Specifically, the selective SSM model presented in Eq. (15) can be reformulated as
where denotes the matrix form of SSMs that uses the sequentially semiseparable representation, and represents the mapping from the th input token to the th output token. Here, and are the selective state-space parameters associated with the th and th tokens, respectively, and denotes the product of the transition terms from step to .
By representing state-space models as semiseparable matrices via matrix transformations, the structured state-space model enables block-based matrix multiplication, allowing content-aware modeling analogous to attention while achieving subquadratic-time computation. Building on this insight, we perform the SSD operator in parallel over all variates:
where each is a semiseparable operator acting only on the th variate sequence. Notably, during this process, we prevent any cross-series information flow by re-initializing (i.e., zeroing) the hidden state at the beginning of each series, thereby forcing the Mamba-SSD path to focus exclusively on temporal modeling. This explicitly enforces that the Cross-Time branch captures temporal dynamics within each variate while remaining fully parallelizable across variates.
(d) Gated output projection. Finally, we combine the SSD output from the content branch with the gate branch to obtain the final output:
where is the sigmoid function and denotes element-wise multiplication. The purpose of this gated projection is to provide an explicit, token-adaptive modulation on the propagated state features, selectively enhancing informative temporal responses while attenuating irrelevant or noisy activations, thereby improving robustness and expressiveness without introducing cross-variational coupling. Thus, the overall process in the Mamba-SSD module is
3.2.3 Cross-variate dependencies modeling via mamba-DALA
A critical yet often overlooked factor in multivariate time series modeling is the delay dependency: variations in one variate may influence another only after a time lag. For instance, in traffic forecasting, congestion at an intersection typically propagates to nearby intersections with a non-negligible delay rather than instantaneously. This lagged propagation property makes forecasting more challenging. Accurately capturing such delayed interactions is essential for both predictive performance and practical utility, yet this effect is often overlooked in existing time series models.
To efficiently capture cross-variate dependencies in the Cross-Variate component while explicitly accounting for delay effect, we develop a Mamba-style mixing block that replaces the recurrent SSM in the Mamba backbone with a Delay-Aware Linear Attention (DALA) operator, termed Mamba-DALA. Built upon Mamba’s efficient block design [51], Mamba-DALA performs token-wise pairwise interaction modeling across variates, yielding delay-aware cross-variate aggregation among multiple correlated latent series.
(a) Mamba-style gating and parameterization. Following the two-branch design in Mamba, we apply linear projections to obtain a content branch and a gate branch . We further construct the linear-attention operator by , where .
(b) Delay-aware linear attention (DALA). To robustly encode propagation delays, we consider two delay signals: a global correlation delay and a token-level relative delay. The key idea is to first correct coarse cross-series misalignment with a global offset, and then explicitly model the relative temporal distance between interacting tokens. In this way, the model can both identify the most plausible cross-series alignment and preserve ordered, causal interactions at the token level. Accordingly, we first estimate a correlation-based delay by maximizing the cross-correlation between variates, which serves as an explicit prior for cross-series alignment. Importantly, this correlation-based delay is not imposed as a hard alignment constraint. Instead, it provides a coarse global prior. At the same time, the final delay-aware interaction is further refined by token-level relative delay encoding via Rotary Position Embedding (RoPE) [52], which captures time lags between query-key token pairs.
Global correlation delay. we first estimate the global propagation delay by maximizing the cross-correlation [53] between latent correlated series and via temporal shifting [42]. Meanwhile, we quantify the correlation strength as the peak correlation value:
where denotes a -step shift applied to , and represents the Pearson correlation function. The strength is used to weight cross-variate aggregation, so that more strongly correlated pairs contribute more. In contrast, noisier or less reliable pairs are naturally downweighted, reducing the influence of noisy or weakly correlated variate pairs. Since is defined in the original time point scale, we map it to the token scale as an integer shift:
where is the number of time points per token. Here, indicates that series lags behind series on the token scale.
where is the time points per token, indicates that series lags behind series on the token scale.
Token-level relative delay. For a query token at position and a key token at position , RoPE parameterizes their token-wise relative delay as
where and denote rotary position embeddings for the th and th tokens, respectively, and is the shared predefined RoPE parameterization [52]. To incorporate the global token shift, we align the key token from series to the effective position when interacting with series . Consequently, the query-key interaction is governed by a unified effective delay:
which jointly accounts for the global propagation offset and the local token-wise lag.
Delay-aware attention aggregation. Given a query token at position in series , we allow it to attend to key tokens from all variates , while enforcing the causal constraint to avoid information leakage. We inject the effective delay by applying RoPE to the query at position and to the key at position , yielding and , respectively. The delay-aware linear attention output is computed as
where is the kernel function in [54], i.e., with , and denotes the element-wise power . This kernelization enhances the expressiveness of linear attention by amplifying informative query-key similarity patterns, while further emphasizes reliable cross-variate interactions. Applying Eq. (25) to all tokens yields the DALA output , which captures delay-aware cross-variate dependencies among all patched series tokens. Overall, the robustness of DALA stems from three aspects: (i) the global delay provides coarse alignment, (ii) token-level relative delay encoding further refines interactions locally, and (iii) correlation-strength weighting suppresses unreliable variate pairs.
(c) Gated output projection. Following Mamba’s gated mixing mechanism, we modulate with the gate branch through an element-wise multiplicative gate, and then project it back to the model dimension:
where denotes element-wise multiplication and is the gating activation. Therefore, the overall Mamba-DALA module can be summarized as
3.2.4 DuoMNet block design
As illustrated in Fig. 2, each DuoMNet block consists of two complementary paths: (i) a temporal path implemented by Mamba-SSD to capture cross-time dependencies, and (ii) a variate path implemented by Mamba-DALA to model delay-aware cross-variate dependencies. We update the input and of both paths via residual connections and pass them to the next block:
To integrate the two complementary outputs, we first apply LayerNorm to each path and then perform a weighted fusion:
where and are mixing hyperparameters. We then apply a feed-forward network (FFN) to the mixed representation and use a residual connection followed by LN to obtain the block output:
Here, is the output of the -th DuoMNet block. By stacking DuoMNet blocks, we obtain . Following a layer-wise aggregation strategy, the final representation is computed as
3.3 Task-specific projection
To adapt DeMa to diverse downstream tasks (e.g., forecasting, imputation, anomaly detection, and classification), we attach a lightweight task head on top of the final task-agnostic representation to learn task-specific targets. For regression-oriented tasks such as forecasting, imputation, and anomaly detection, we employ an MLP as the task head . For forecasting, the prediction is , where is the prediction horizon. For imputation and anomaly detection, the head produces point-wise outputs over the lookback window, i.e., , where is the lookback length. For series-level classification, we first aggregate token-wise representations into a global feature, and then apply an MLP classifier with softmax , where and is the number of classes.
3.4 Training process and complexity analysis
The overall training procedure is summarized in Algorithm 1. The computational cost of DeMa is dominated by two parallel modules, Mamba-SSD and Mamba-DALA. Let be the number of patches and the model dimension. For Mamba-SSD, the Cross-Time scan yields , which can be viewed as a scanned sequence of length for SSD-based selective state-space computation. The resulting FLOPs scale linearly with the scanned length and quadratically with the feature dimension, giving a per-block complexity of . Similarly, Mamba-DALA operates on , which can equivalently be viewed as a sequence of length . Its core delay-aware linear mixing is implemented via linear attention and thus also requires FLOPs per block. Since the two paths are computed in parallel, the overall complexity per DuoMNet block remains . Consequently, DeMa achieves linear-time complexity with respect to the effective scanned length , enabling efficient modeling for long-horizon and large-scale multivariate time series in regimes where the effective token length dominates the feature dimension, i.e., . Therefore, although DeMa contains multiple modules, its dominant computation remains concentrated in two linear-time paths, making the overall architecture computationally practical.
4 Experiments
In this section, we investigate the following research questions to validate the effectiveness, efficiency, and scalability of DeMa:
● RQ1: How does DeMa perform across five representative time series tasks?
● RQ2: Does DeMa achieve a favorable accuracy-efficiency trade-off as a general backbone for multivariate time-series analysis?
● RQ3: How does DeMa scale with increasing input length in terms of per-iteration runtime and GPU memory, compared with Transformer-based baselines?
● RQ4: How does each key component contribute to the overall performance of DeMa?
● RQ5: How sensitive is DeMa to major hyperparameters?
4.1 Experimental setup
4.1.1 Datasets
To validate the effectiveness of DeMa, we conduct extensive experiments across five mainstream time-series analysis tasks: long-term forecasting, short-term forecasting, imputation, classification, and anomaly detection. All datasets are drawn from the benchmark collections in the Time Series Library (TSLib) [40]. Table 2 summarizes the datasets used for each task.
● Forecasting. We evaluate both long-term and short-term forecasting. For long-term forecasting, we adopt widely used benchmarks including Electricity [55], ETT with four subsets (ETTh1, ETTh2, ETTm1, ETTm2) [23], Exchange [56], Traffic [25], Weather [25], and Solar-Energy [56]. For short-term forecasting, we use four traffic datasets from the PEMS family [57]: PEMS03, PEMS04, PEMS07, and PEMS08.
● Imputation. Time-series imputation aims to recover missing values from contextual observations. We evaluate on ETTh1, ETTh2, ETTm1, ETTm2, Electricity, and Weather. Following TimesNet [18], we randomly mask time points with ratios in to assess robustness under varying missing rates.
● Anomaly detection. Anomaly detection aims to identify abnormal patterns in time series. We adopt five widely used industrial benchmarks: Server Machine Dataset (SMD) [58], Mars Science Laboratory rover (MSL) [59], Soil Moisture Active Passive satellite (SMAP) [59], Secure Water Treatment (SWaT) [60], and Pooled Server Metrics (PSM) [61].
● Classification. For time-series classification, we use ten multivariate datasets from the UEA Time Series Classification Archive [62], covering diverse domains such as gesture recognition, action recognition, audio recognition, and medical diagnosis.
4.1.2 Baselines
To ensure a comprehensive comparison, we include a broad set of strong baselines that are representative of the latest advances in the time series community. Specifically, our baselines span four families: (1) Mamba-based models: Affirm [63], S-Mamba [64], CMamba [65], and SAMBA [66]; (2) Transformer-based models: iTransformer [13], Crossformer [27], and PatchTST [10]; (3) MLP-based models: TimeMixer [22], DLinear [21], and RLinear [67]; and (4) TCN-based models: ModernTCN [20] and TimesNet [18].
4.2 Metrics and implementation details
To ensure a fair and comprehensive comparison, we follow the experimental protocol of the well-established Time Series Library [40]. For long-term forecasting and imputation, we report mean squared error (MSE) and mean absolute error (MAE). For time-series classification, we report accuracy. For anomaly detection, we report the F1-score, which balances precision and recall and is well-suited to the highly imbalanced nature of anomalous events. In each table, the best and second-best results are highlighted in red and underlined, respectively.
We use the published optimal hyperparameter settings for all baselines. All experiments are implemented in PyTorch and conducted on three NVIDIA RTX A6000 GPUs (48 GB). We train all models with the Adam optimizer [68] under an loss, and tune the initial learning rate in . The number of DuoMNet blocks is selected from via hyperparameter search, and the representation dimension is chosen from . We set the batch size to 32 and train for 50 epochs. For the Adaptive Fourier Filter, the top-ratio is selected from a limited candidate set, and we use in the final configuration. We further conduct a grid search over the fusion weights , and use for forecasting and classification, and for imputation and anomaly detection. For tokenization, we fix the patch size and stride to . DeMa follows a modular pipeline. Once the temporal path (Mamba-SSD) and the variate path (Mamba-DALA) are instantiated, the model mainly repeats the same DuoMNet block, which helps keep the implementation structured and reproducible.
4.3 Main results across five time-series tasks (RQ1)
4.3.1 Long-term forecasting results
Time-series forecasting is a fundamental task in time-series analysis, aiming to predict future values from historical observations. Table 3 reports the long-term forecasting results with a fixed lookback length and prediction horizons . Overall, DeMa delivers top-tier accuracy consistently across all datasets and horizons, ranking within the top two in every setting. In particular, DeMa achieves the best results in 34/36 cases for MSE and 30/36 cases for MAE, demonstrating strong robustness and generalization across diverse temporal dynamics.
Compared with Transformer-based models such as iTransformer, PatchTST, and Crossformer, DeMa consistently performs better across all settings, indicating that explicitly disentangling and modeling temporal and variate dependencies is more effective for long-horizon forecasting than relying primarily on global attention. DeMa also consistently surpasses Mamba variants and non-attention baselines. For example, on Traffic (Avg), DeMa reduces the MSE from 0.414 to 0.372, achieving a 10.1% improvement over the strongest Mamba baseline, S-Mamba, and also outperforms representative MLP- and TCN-based models, including TimeMixer (0.484) and ModernTCN (0.546). These gains can be attributed to DeMa’s comprehensive dependency modeling: Mamba-SSD efficiently captures long-range temporal dependencies through parallel selective state-space computation, while Mamba-DALA enhances cross-variate interaction modeling. Their synergy enables DeMa to exploit long-range temporal structures and inter-series correlations jointly, yielding consistently strong performance in long-term forecasting. We note that DeMa remains highly competitive even on datasets with only a few variates. As summarized in Table 2, ETTh1, ETTh2, ETTm1, and ETTm2 each contain 7 variates, while Exchange contains 8 variates. On these small-scale benchmarks, DeMa still achieves best or near-best averaged results in Table 3. These results indicate that the benefit of DeMa persists in small-scale scenarios, though its relative advantage becomes more pronounced as scalability pressure increases.
We further examine DeMa under relatively weak inter-variable dependencies, as analyzed on ETTh2, ETTm1, and ETTm2. Even in these settings, DeMa remains highly competitive, achieving the best average MSE on ETTm1 (0.367) and ETTm2 (0.270), while maintaining comparable performance on ETTh2 (0.370). This observation suggests that when the correlation-based delay prior becomes less informative, the temporal path can still capture the dominant intra-series dynamics, whereas the variate path provides complementary gains.
4.3.2 Short-term forecasting results
We evaluate short-term forecasting on four PEMS traffic benchmarks, where accurate prediction is challenging due to intricate spatiotemporal dependencies among sensors and strong inter-series interactions that drive citywide traffic dynamics. Table 4 summarizes the results under a fixed lookback length and prediction horizons . Overall, DeMa achieves the best performance in all 16 settings (both MSE and MAE), demonstrating its strong effectiveness for short-term forecasting. In terms of averaged MSE, DeMa improves over the strongest competitor by a clear margin, reducing errors by 25.8% on PEMS03, 15.7% on PEMS04, 12.7% on PEMS07, and 16.4% on PEMS08. Notably, variate-independent architectures exhibit pronounced performance degradation on these datasets. For example, PatchTST, which largely models each variate independently, is consistently inferior to DeMa across all horizons; a similar trend is observed for MLP-based methods, whose channel-independent mixing fails to capture dynamic cross-sensor interactions. We attribute the consistent gains of DeMa to its explicit and effective cross-variate modeling: Mamba-DALA strengthens cross-variate interactions, including delay-aware dependencies among traffic sensors, and complements the cross-time temporal modeling of Mamba-SSD. Their synergy enables DeMa to jointly exploit temporal dynamics and delay-aware inter-sensor dependencies, resulting in substantially improved short-term forecasting performance on the PEMS benchmarks.
4.3.3 Imputation results
Time series imputation aims to recover missing values from partially observed series, which is crucial in real-world scenarios where data acquisition is imperfect due to sensor malfunctions or unexpected transmission glitches. As reported in Table 5, DeMa consistently achieves the best performance across all datasets and masking ratios, surpassing both Transformer-based and SSM-based competitors by clear margins. A notable trend is that the advantage of DeMa becomes more pronounced as the missing rate increases, where accurate recovery increasingly relies on fine-grained modeling of local fluctuations and precise temporal alignment between observed contexts and missing entries. For instance, on ETTh1 with masking, DeMa attains a low error of , whereas Transformer baselines (e.g., PatchTST: ; iTransformer: ) and SSM-based methods (e.g., S-Mamba: ) degrade substantially, indicating limited capability in capturing the local variations required for reliable completion. These results suggest that imputation is not merely a global dependency modeling problem, but a fine-grained completion task that hinges on (i) local context continuity around missing positions and (ii) accurate relative lag relations between correlated variates. Concretely, Mamba-SSD provides efficient temporal context aggregation to supply informative local cues. At the same time, Mamba-DALA enables delay-aware cross-variate interactions, enabling observed variates to contribute aligned evidence for reconstructing missing values. Together, their synergy enables more precise recovery of local dynamics and inter-series signals, resulting in consistently superior imputation accuracy.
4.4 Classification results
Time series classification aims to assign a semantic label to a multivariate series by recognizing both global patterns and discriminative local cues. We evaluate on 10 multivariate datasets from the UEA Time Series Archive [62] and report the average classification accuracy of each method across the 10 benchmarks. As shown in Fig. 3(a), DeMa achieves the highest average accuracy of , consistently outperforming strong baselines from all major families, including ModernTCN (), TimesNet (), Crossformer (), and TimeMixer (). A key observation is that TCN-based methods achieve competitive classification performance. This is because temporal convolutions provide a strong inductive bias for extracting discriminative local motifs, e.g., shape-like patterns, that are often sufficient to determine class labels. With hierarchical convolutions, TCNs can further aggregate multi-scale evidence while remaining robust to small phase shifts and noise through parameter sharing and localized receptive fields, making them particularly effective for semantic recognition. In contrast, MLP-based approaches remain inferior (e.g., DLinear: , RLinear: ), since their predominantly linear mixing is less capable of building high-level semantic features and modeling complex inter-series interactions. The superior performance of DeMa stems from its fine-grained and hierarchical dependency encoding: stacking DuoMNet blocks progressively aggregates multi-scale temporal evidence, while the dual-path design explicitly captures both temporal variations and cross-variate cues. In particular, Mamba-SSD strengthens long-range temporal structure modeling, and Mamba-DALA enhances cross-variate interactions, enabling DeMa to form more informative and semantically aligned representations for classification beyond purely convolutional local feature extraction.
4.4.1 Anomaly detection results
Anomaly detection aims to identify rare and abnormal patterns in time series, which often correspond to faults, critical events, or outliers requiring timely intervention. Following prior work [40], we evaluate on five widely used anomaly detection benchmarks and report the average F1-score across datasets [69]. For fair comparison, we adopt reconstruction error [18] as the anomaly criterion for all methods. As shown in Fig. 3(b), DeMa achieves the best average F1-score of , outperforming strong competitors such as ModernTCN () and TimesNet (). This gain highlights the advantage of our fine-grained modeling. The proposed SSDDALA design jointly strengthens (i) local temporal continuity and long-range temporal context via Mamba-SSD, and (ii) delay-aware cross-variate interactions via Mamba-DALA, enabling more faithful reconstruction of normal dynamics. As a result, abnormal deviations yield sharper and more separable reconstruction residuals, leading to improved detection accuracy. In contrast, iTransformer [13] and S-Mamba [64] perform notably worse, likely because prevalent normal patterns can dominate their series-wise global similarity modeling by attention; consequently, the subtle abnormal segments are more easily diluted when aggregating pairwise correlations, leading to inferior detection. Finally, we observe that methods benefiting from explicit decomposition cues (e.g., TimesNet and DeMa) tend to achieve stronger overall detection performance, underscoring the importance of separating stable periodic structures from irregular fluctuations so that violations of normal regularities become more detectable.
4.5 Performance–efficiency trade-off (RQ2)
Figure 4 summarizes the overall effectiveness and efficiency of DeMa. Figure 4(a) provides an overall comparison of baselines across five representative time-series tasks. Overall, DeMa achieves consistently strong results on all tasks, demonstrating robust task generality rather than tuning to a single forecasting setting. Figure 4(b) further evaluates the accuracy-efficiency trade-off on a large-scale traffic dataset with variates under long-term forecasting, where the y-axis denotes MSE, the x-axis indicates training time, and the bubble size reflects GPU memory footprint. As shown, DeMa attains the lowest MSE () with moderate training time ( ms/iter) and manageable memory usage ( GB), offering a favorable balance between accuracy and computational cost. These results suggest that DeMa is a practical and scalable solution for large-scale multivariate time-series analysis.
DeMa exhibits clear efficiency advantages over Transformer-based baselines while remaining more accurate, benefiting from the linear-time state-space backbone. For instance, Crossformer and PatchTST require substantially higher training time and memory footprint (e.g., ms/iter and GB for Crossformer), yet still yield higher MSE than DeMa in Fig. 4(b). Moreover, DeMa consistently outperforms TCN-based models in the multi-task comparison (Fig. 4(a)). While TCNs are efficient and effective at capturing local motifs, their limited receptive field and forecasting-oriented inductive bias may limit transferability to tasks that require long-range aggregation. In contrast, DeMa learns more transferable representations by jointly encoding long-range temporal dynamics and delay-aware inter-series interactions. Finally, DeMa improves the trade-off between accuracy and efficiency compared to recent Mamba variants. As shown in Fig. 4(b), DeMa is faster and lighter than Affirm and CMamba (e.g., ms/iter and GB for DeMa versus ms/iter and GB for Affirm, and ms/iter and GB for CMamba), while achieving lower MSE.
In summary, DeMa combines state-of-the-art accuracy with practical training speed and memory efficiency, making it a strong candidate for scalable, general time-series analysis.
4.6 Scalability vs. input length (RQ3)
Semantic information in time series is often formed by aggregating evidence over long horizons rather than isolated timestamps. Although longer lookback windows provide richer context, they also place stricter demands on training efficiency and GPU memory. To evaluate scalability with respect to the input length, Fig. 5 reports the per-iteration running time and GPU memory footprint of DeMa and representative Transformer-based baselines on ETTh1. We increase the lookback length from while fixing the embedding dimension to for a fair comparison.
As shown, DeMa consistently achieves the lowest running time and memory usage across all lengths, and its growth remains close to linear as the series length increases. By contrast, Transformer-based methods (e.g., Transformer, PatchTST, and Crossformer) exhibit rapidly increasing computation and memory overhead with longer inputs, which substantially limits their practicality in long-term settings. Notably, iTransformer is more efficient than token-wise Transformers because it computes series-wise attention, avoiding quadratic growth with respect to the token length; however, it remains consistently slower and more memory-intensive than DeMa at large input lengths. These results further explain why the advantages of DeMa become especially pronounced when the input context grows: the dual path preserves near-linear scaling with series length, whereas attention-based baselines incur substantially steeper computational and memory growth.
These observations align with our complexity analysis in Fig. 1(b). Leveraging the parallelizable Mamba-SSD and Mamba-DALA blocks, DeMa scales linearly with the token length, with time complexity , where is the number of variates, is the number of tokens after tokenization, and is the embedding dimension. In typical long-term MTS scenarios, , making DeMa a more scalable backbone for long-term and large-scale multivariate time-series modeling.
4.7 Ablation study (RQ4)
To quantify the contribution of each key component in DeMa, including the Adaptive Fourier Filter, Mamba-SSD, and Mamba-DALA, we conduct ablation studies on Traffic (long-term forecasting) and PEMS04 (short-term forecasting). As reported in Table 6, we consider two types of ablations: replacement (Repl.), which substitutes a target module with an alternative design, and removal (w/o), which disables the module entirely. Overall, each component consistently improves forecasting accuracy, and the full model yields the best performance on both datasets.
● Adaptive fourier filter. Comparing Rows 1 and 5, removing the decomposition module causes clear degradation on both datasets (Traffic MSE: ; PEMS04 MSE: ). This verifies that the Adaptive Fourier Filter provides a beneficial decomposition prior by separating slowly varying structures (e.g., trend and seasonality) from short-term fluctuations, thereby reducing spectral interference and facilitating more reliable dependency modeling in both the temporal and variate paths, consistent with decomposition-based forecasting literature [25,47].
● Mamba-SSD for temporal modeling. Rows 2–3 replace the temporal Mamba-SSD with attention-style alternatives, leading to substantial performance drops (e.g., Row 3: Traffic MSE , a increase). This indicates that Mamba-SSD is critical for efficiently capturing long-range temporal dynamics. Moreover, removing the temporal path entirely (Row 7) results in the largest degradation (Traffic MSE , ; PEMS04 MSE ), confirming that strong cross-time modeling is indispensable for accurate forecasting.
● Mamba-DALA for cross-variate interaction modeling. Rows 4 and 6 highlight the necessity of Mamba-DALA for modeling cross-variate dependencies. Replacing Mamba-DALA with a temporal-style Mamba-SSD interaction (Row 4) already hurts performance (Traffic MSE ; PEMS04 MSE ), suggesting that directly reusing temporal mechanisms is insufficient for variate interactions. More importantly, removing Mamba-DALA (Row 6) causes a pronounced drop (Traffic MSE , ; PEMS04 MSE ), verifying that delay-aware cross-variate modeling provides critical complementary cues beyond temporal dynamics.
Overall, the ablation results show that the main components of DeMa are functionally complementary rather than redundant. This supports that the structured design of DeMa is challenge-driven and empirically justified, rather than an unnecessarily complicated composition.
4.8 Hyperparameter sensitivity analysis (RQ5)
In DeMa, we adopt a weighted fusion layer to combine the temporal-path representation and the variate-path representation, controlled by the fusion weights and in Eq. (29). To examine the sensitivity of this fusion trade-off, we perform a grid search over on forecasting and imputation on Weather [25], classification on Heartbeat [62], and anomaly detection on MSL [59]. The results are summarized in Figs. 6(a)−6(d), where higher is better for Accuracy, F1-score, and .
Across all tasks, DeMa is most reliable when both paths remain active (i.e., neither nor is overly small), confirming that temporal modeling and cross-variate interaction are complementary rather than substitutable. Forecasting is relatively less sensitive to and (Fig. 6(a)), showing only mild variation across the grid. An explanation is that long-horizon prediction is often dominated by stable structures such as trend and seasonality; thus, once the temporal path is sufficiently emphasized, the model can still perform well even under moderate fusion bias, and the variate path provides a relatively smaller marginal gain unless the dataset is strongly coupled. Similarly, classification remains stable across a wide range of (Fig. 6(b)). This is because classification primarily relies on global semantic discrimination: the stacked DuoMNet blocks can progressively aggregate high-level evidence and compensate for moderate fusion imbalance, as long as neither path is completely suppressed and both temporal and cross-variate cues are retained.
In contrast, imputation is more sensitive to the fusion weights (Fig. 6(c)), because recovering missing points typically requires both local temporal continuity (temporal path) and cross-variate complementarity (variate path). Over-weighting either path can lead to incomplete recovery, e.g., temporal-only fusion may miss informative cross-variate cues, while variate-only fusion may fail to preserve fine-grained local dynamics, especially under higher missing ratios. Similarly, anomaly detection exhibits strong sensitivity (Fig. 6(d)), since anomalies are rare, local patterns that demand temporal modeling to capture abrupt deviations, while cross-variate modeling helps validate abnormality via multi-sensor consistency; improper weighting may either dilute anomalies under dominant normal patterns or amplify spurious fluctuations, increasing false alarms. These observations suggest that and should be selected task-awarely, with balanced fusion particularly important for fine-grained, locality-critical tasks such as imputation and anomaly detection.
5 Related work
5.1 Transformer-based time series models
Transformers have become a dominant paradigm for multivariate time series (MTS) modeling due to their strong ability to capture pairwise dependencies with a global receptive field. Representative approaches include Informer [23], Fedformer [24], Autoformer [25], Pyraformer [26], PatchTST [10], Crossformer [27], and UniTime [70]. A fundamental limitation of these models lies in the quadratic complexity of self-attention with respect to the token length , which hinders scalability to long-horizon settings. To alleviate this issue, sparse or structured attention has been explored to reduce the computational burden (e.g., to ) [23,55,71]. Yet, it may overlook informative dependencies when only a subset of interactions is retained. More recently, iTransformer [13] reformulates MTS tokens by treating each variate as a token and applying attention over variates, shifting the quadratic term from to the number of variates . While efficient when , this design may sacrifice fine-grained local temporal interactions that are crucial for locality-sensitive tasks.
5.2 MLP-based time series models
MLP-style architectures offer an attractive alternative for efficient time-series modeling by replacing attention with lightweight mixing operations. Classic representatives include DLinear [21] and its variants such as RLinear [67], which adopt decomposition-aware linear mapping strategies to achieve strong forecasting performance with low computational cost. TimeMixer [22] further enhances expressiveness via MLP-based mixing across temporal and channel dimensions. Despite their efficiency, purely MLP-based designs often rely on relatively limited inductive biases for explicitly modeling complex cross-variate interactions and fine-grained dependency structures, which can restrict their generality across diverse MTS tasks beyond forecasting.
5.3 TCN-based time series models
Temporal convolutional networks (TCNs) model temporal dependencies via convolutional receptive fields and have been widely used in time series forecasting and representation learning. ModernTCN [20] strengthens classical TCNs with modern architectural practices to improve both accuracy and efficiency, while TimesNet [18] leverages convolutional modeling together with period-aware representations to capture multi-period temporal patterns. Although convolution provides strong locality and parallelism, TCN-based models typically require carefully designed receptive fields to capture long-range dependencies, and explicitly modeling cross-variate interactions remains non-trivial when scaling to high-dimensional MTS.
5.4 Mamba-based time series models
State space models (SSMs), particularly Mamba-style selective SSMs, have recently emerged as a compelling linear-time alternative to Transformers, demonstrating strong performance across diverse domains [35]. This progress has spurred increasing interest in adapting Mamba to time series analysis [63–66,72,73]. For example, CMamba [65] augments Mamba-style temporal modeling with extra modules to capture cross-variate interactions, while S-Mamba [64], MambaMixer [72], MambaTS [73], and more recent variants such as Affirm [63] and SAMBA [66] explore architectural adaptations to improve effectiveness on time series benchmarks. Despite these advances, existing Mamba-based designs still face notable challenges for general MTS analysis. First, dependency entanglement is common: temporal dynamics and cross-variate interactions frequently overlap and are mixed, which can obscure task-relevant cues. Second, fine-grained cross-variate modeling remains insufficient: many methods primarily emphasize global interactions and may under-exploit local, delay-sensitive cross-variate effects. Motivated by these gaps, we propose DeMa, which disentangles cross-time and cross-variate dependency encodings via a dual-path architecture, enabling efficient yet effective modeling for general time-series analysis.
6 Conclusion
In this work, we propose DeMa, a dual-path delay-aware Mamba backbone for efficient multivariate time-series analysis. DeMa explicitly decomposes the time-series context into Cross-Time and Cross-Variate components via an Adaptive Fourier Filter. These components are processed by stacked DuoMNet blocks with two parallel, scan-coupled pathways: (i) a temporal path that applies Cross-Time Scan and Mamba-SSD to capture long-range intra-series dependencies in a series-independent and parallel manner, and (ii) a variate path that applies Cross-Variate Scan and Mamba-DALA to model delay-aware inter-series interactions using DALA with both global correlation delays and token-level relative delays. The resulting representations are fused through a lightweight weighted fusion layer and projected to task-specific outputs. Extensive experiments across five mainstream tasks demonstrate that DeMa achieves consistently strong accuracy while substantially reducing training time and GPU memory usage, suggesting a favorable accuracy-efficiency trade-off and practical scalability for long-horizon and large-scale MTS modeling.
Kong X, Chen Z, Liu W, Ning K, Zhang L, Marier S M, Liu Y, Chen Y, Xia F. Deep learning for time series forecasting: a survey. International Journal of Machine Learning and Cybernetics, 2025, 16(7−8): 5079−5112
[2]
Liang Z, An R, Fan W, Rao Y, Liang Y. iTFKAN: interpretable time series forecasting with kolmogorov-arnold network. 2025, arXiv preprint arXiv: 2504.16432
[3]
An R, Zhang Y, Liang Z, Fan W, Liang Y, Shang X, Li Q. Damba-ST: domain-adaptive mamba for efficient urban spatio-temporal prediction. 2025, arXiv preprint arXiv: 2506.18939
[4]
Huo G, Zhang Y, Wang B, Gao J, Hu Y, Yin B. Hierarchical spatio-temporal graph convolutional networks and transformer network for traffic flow forecasting. IEEE Transactions on Intelligent Transportation Systems, 2023, 24( 4): 3855–3867
[5]
Fang C, Wang C. Time series data imputation: a survey on deep learning approaches. 2020, arXiv preprint arXiv: 2011.11347
[6]
Chen F, Qin Z, Zhou M, Zhang Y, Deng S, Fan L, Pang G, Wen Q. LARA: a light and anti-overfitting retraining approach for unsupervised time series anomaly detection. In: Proceedings of the ACM Web Conference 2024. 2024, 4138−4149
[7]
Blázquez-García A, Conde A, Mori U, Lozano J A. A review on outlier/anomaly detection in time series data. ACM Computing Surveys, 2021, 54(3): Article 56
[8]
Qin Z, Zhang Y, Meng S, Qin Z, Choo K K R. Imaging and fusing time series for wearable sensor-based human activity recognition. Information Fusion, 2020, 53: 80−87
[9]
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 6000−6010
[10]
Nie Y, Nguyen N H, Sinthong P, Kalagnanam J. A time series is worth 64 words: long-term forecasting with transformers. In: Proceedings of the 11th International Conference on Learning Representations. 2023
[11]
Liu Y, Zhang H, Li C, Huang X, Wang J, Long M. Timer: generative pre-trained transformers are large time series models. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 32369−32399
[12]
Woo G, Liu C, Kumar A, Xiong C, Savarese S, Sahoo D. Unified training of universal time series forecasting transformers. In: Proceedings of the 41st International Conference on Machine Learning. Proceedings of Machine Learning Research, 2024, 235: 53140−53164
[13]
Liu Y, Hu T, Zhang H, Wu H, Wang S, Ma L, Long M. iTransformer: inverted transformers are effective for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024
[14]
Han L, Ye H J, Zhan D C. The capacity and robustness trade-off: revisiting the channel independent strategy for multivariate time series forecasting. IEEE Transactions on Knowledge and Data Engineering, 2024, 36( 11): 7129–7142
[15]
Lin S, Lin W, Wu W, Zhao F, Mo R, Zhang H. SegRNN: segment recurrent neural network for long-term time-series forecasting. IEEE Internet of Things Journal, 2026, 13( 5): 9861–9871
[16]
Hewamalage H, Bergmeir C, Bandara K. Recurrent neural networks for time series forecasting: current status and future directions. International Journal of Forecasting, 2021, 37( 1): 388–427
[17]
Rangapuram S S, Seeger M, Gasthaus J, Stella L, Wang Y, Januschowski T. Deep state space models for time series forecasting. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. 2018, 7796−7805
[18]
Wu H, Hu T, Liu Y, Zhou H, Wang J, Long M. TimesNet: temporal 2D-variation modeling for general time series analysis. In: Proceedings of the 11th International Conference on Learning Representations. 2023
[19]
Liu M, Zeng A, Chen M, Xu Z, Lai Q, Ma L, Xu Q. SCINet: time series modeling and forecasting with sample convolution and interaction. Advances in Neural Information Processing Systems, 2022, 35: 5816−5828
[20]
Luo D, Wang X. ModernTCN: a modern pure convolution structure for general time series analysis. In: Proceedings of the 12th International Conference on Learning Representations. 2024
[21]
Zeng A, Chen M, Zhang L, Xu Q. Are transformers effective for time series forecasting? In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. 2023, 11121−11128
[22]
Wang S, Wu H, Shi X, Hu T, Luo H, Ma L, Zhang J Y, Zhou J. TimeMixer: decomposable multiscale mixing for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024
[23]
Zhou H, Zhang S, Peng J, Zhang S, Li J, Xiong H, Zhang W. Informer: beyond efficient transformer for long sequence time-series forecasting. In: Proceedings of the 35th AAAI Conference on Artificial Intelligence. 2021, 11106−11115
[24]
Zhou T, Ma Z, Wen Q, Wang X, Sun L, Jin R. FEDformer: frequency enhanced decomposed transformer for long-term series forecasting. In: Proceedings of the 39th International Conference on Machine Learning. 2022, 27268−27286
[25]
Wu H, Xu J, Wang J, Long M. Autoformer: decomposition transformers with auto-correlation for long-term series forecasting. In: Proceedings of the 35th Conference on Neural Information Processing Systems. 2021, 22419−22430
[26]
Liu S, Yu H, Liao C, Li J, Lin W, Liu A X, Dustdar S. Pyraformer: low-complexity pyramidal attention for long-range time series modeling and forecasting. In: Proceedings of the 10th International Conference on Learning Representations. 2022
[27]
Zhang Y, Yan J. Crossformer: transformer utilizing cross-dimension dependency for multivariate time series forecasting. In: Proceedings of the 11th International Conference on Learning Representations. 2023
[28]
Jia J, Gao J, Xue B, Wang J, Cai Q, Chen Q, Zhao X, Jiang P, Gai K. From principles to applications: a comprehensive survey of discrete tokenizers in generation, comprehension, recommendation, and information retrieval. 2025, arXiv preprint arXiv: 2502.12448
[29]
Qu H, Fan W, Zhao Z, Li Q. TokenRec: learning to tokenize ID for LLM-based generative recommendations. IEEE Transactions on Knowledge and Data Engineering, 2025, 37( 10): 6216–6231
[30]
Qu H, Lin S, Ding Y, Wang Y, Fan W. Diffusion generative recommendation with continuous tokens. In: Proceedings of the ACM Web Conference 2026. 2026, 7259−7270
[31]
Zhou Y, Qu H, Liu Y, Lin S, Song L, Fan W. HD-Prot: a protein language model for joint sequence-structure modeling with continuous structure tokens. 2025, arXiv preprint arXiv: 2512.15133
[32]
Ekambaram V, Jati A, Nguyen N, Sinthong P, Kalagnanam J. TSMixer: lightweight MLP-mixer model for multivariate time series forecasting. In: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2023, 459−469
[33]
Gu A, Dao T. Mamba: linear-time sequence modeling with selective state spaces. 2023, arXiv preprint arXiv: 2312.00752
[34]
Dao T, Gu A. Transformers are SSMs: generalized models and efficient algorithms through structured state space duality. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 10041−10071
[35]
Qu H, Ning L, An R, Fan W, Derr T, Liu H, Xu X, Li Q. A survey of Mamba. ACM Transactions on Intelligent Systems and Technology, 2026. DOI: 10.1145/3819582
[36]
Lenz B, Lieber O, Arazi A, Bergman A, Manevich A, , et al. Jamba: hybrid transformer-mamba language models. In: Proceedings of the 13th International Conference on Learning Representations. 2025
[37]
Zhu L, Liao B, Zhang Q, Wang X, Liu W, Wang X. Vision Mamba: efficient visual representation learning with bidirectional state space model. In: Proceedings of the 41st International Conference on Machine Learning. Proceedings of Machine Learning Research, 2024, 235: 62429−62442
[38]
Schiff Y, Kao C H, Gokaslan A, Dao T, Gu A, Kuleshov V. Caduceus: bi-directional equivariant long-range DNA sequence modeling. In: Proceedings of the 41st International Conference on Machine Learning. Proceedings of Machine Learning Research, 2024, 235: 43632−43648
[39]
Zhang Y, Qu H, Ning L, Fan W, Li Q. SSD4Rec: a structured state space duality model for efficient sequential recommendation. ACM Transactions on Information Systems, 2026, 44( 2): 29
[40]
Wang Y, Wu H, Dong J, Liu Y, Wang C, Long M, Wang J. Deep time series models: a comprehensive survey and benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, volume(issue): 1−20
[41]
Zhao L, Shen Y. Rethinking channel dependence for multivariate time series forecasting: learning from leading indicators. In: Proceedings of the 12th International Conference on Learning Representations. 2024
[42]
Long Q, Fang Z, Fang C, Chen C, Wang P, Zhou Y. Unveiling delay effects in traffic forecasting: a perspective from spatial-temporal delay differential equations. In: Proceedings of the ACM Web Conference 2024. 2024, 1035−1044
[43]
Jiang J, Han C, Zhao W X, Wang J. PDFormer: propagation delay-aware dynamic long-range transformer for traffic flow prediction. In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. 2023, 4365−4373
[44]
Gu A, Dao T, Ermon S, Rudra A, Ré C. HiPPO: recurrent memory with optimal polynomial projections. Advances in Neural Information Processing Systems, 2020, 33: 1474−1487
[45]
Harris M, Sengupta S, Owens J D. Parallel prefix sum (scan) with CUDA. In: Nguyen H, ed. GPU Gems 3. Chapter 39. Addison-Wesley Professional, 2007, 851−876
[46]
Yi K, Zhang Q, Fan W, Wang S, Wang P, He H, Lian D, An N, Cao L, Niu Z. Frequency-domain MLPs are more effective learners in time series forecasting. Advances in Neural Information ProcessingSystems, 2023, 36: 76656−76679
[47]
Liu Y, Li C, Wang J, Long M. Koopa: learning non-stationary time series dynamics with Koopman predictors. Advances in Neural Information Processing Systems, 2023, 36: 12271−12290
[48]
Yue Z, Wang Y, Duan J, Yang T, Huang C, Tong Y, Xu B. TS2Vec: towards universal representation of time series. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence.2022, 8980−8987
[49]
Gu A, Johnson I, Timalsina A, Rudra A, Ré C. How to train your HIPPO: state space models with generalized orthogonal basis projections. In: Proceedings of the 11th International Conference on Learning Representations. 2023
[50]
Kim T, Kim J, Tae Y, Park C, Choi J H, Choo J. Reversible instance normalization for accurate time-series forecasting against distribution shift. In: Proceedings of the 10th International Conference on Learning Representations. 2022
[51]
Han D, Wang Z, Xia Z, Han Y, Pu Y, Ge C, Song J, Song S, Zheng B, Huang G. Demystify Mamba in vision: a linear attention perspective. Advances in Neural Information Processing Systems, 2024, 37: 127181−127203
[52]
Su J, Ahmed M, Lu Y, Pan S, Bo W, Liu Y. RoFormer: enhanced transformer with rotary position embedding. Neurocomputing, 2024, 568: 127063
[53]
Hertz D. Time delay estimation by combining efficient algorithms and generalized cross-correlation methods. IEEE Transactions on Acoustics, Speech, and Signal Processing, 1986, 34( 1): 1–7
[54]
Han D, Pan X, Han Y, Song S, Huang G. Flatten transformer: vision transformer using focused linear attention. In: Proceedings of 2023 IEEE/CVF International Conference on Computer Vision. 2023, 5938−5948
[55]
Li S, Jin X, Xuan Y, Zhou X, Chen W, Wang Y X, Yan X. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. 2019, 5243−5253
[56]
Lai G, Chang W C, Yang Y, Liu H. Modeling long-and short-term temporal patterns with deep neural networks. In: Proceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 2018, 95−104
[57]
Chen C, Petty K, Skabardonis A, Varaiya P, Jia Z. Freeway performance measurement system: mining loop detector data. Transportation Research Record: Journal of the Transportation Research Board, 2001, 1748( 1): 96–102
[58]
Su Y, Zhao Y, Niu C, Liu R, Sun W, Pei D. Robust anomaly detection for multivariate time series through stochastic recurrent neural network. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2019, 2828−2837
[59]
Hundman K, Constantinou V, Laporte C, Colwell I, Soderstrom T. Detecting spacecraft anomalies using LSTMs and nonparametric dynamic thresholding. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2018, 387−395
[60]
Mathur A P, Tippenhauer N O. SWaT: a water treatment testbed for research and training on ICS security. In: Proceedings of the 2016 International Workshop on Cyber-Physical Systems for Smart Water Networks. 2016, 31−36
[61]
Abdulaal A, Liu Z, Lancewicki T. Practical approach to asynchronous multivariate time series anomaly detection and localization. In: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. 2021, 2485−2494
[62]
Bagnall A, Dau H A, Lines J, Flynn M, Large J, Bostrom A, Southam P, Keogh E. The UEA multivariate time series classification archive, 2018. 2018, arXiv preprint arXiv: 1811.00075
[63]
Wu Y, Meng X, Hu H, Zhang J, Dong Y, Lu D. Affirm: interactive mamba with adaptive Fourier filters for long-term time series forecasting. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. 2025, 21599−21607
[64]
Wang Z, Kong F, Feng S, Wang M, Yang X, Zhao H, Wang D, Zhang Y. Is Mamba effective for time series forecasting? Neurocomputing, 2025, 619: 129178
[65]
Zeng C, Liu Z, Zheng G, Kong L. CMamba: channel correlation enhanced state space models for multivariate time series forecasting. 2024, arXiv preprint arXiv: 2406.05316
[66]
Weng Z, Han J, Jiang W, Liu H. Simplified mamba with disentangled dependency encoding for long-term time series forecasting. 2024, arXiv preprint arXiv: 2408.12068v2
[67]
Li Z, Qi S, Li Y, Xu Z. Revisiting long-term time series forecasting: an investigation on linear mapping. 2023, arXiv preprint arXiv: 2305.10721
[68]
Kingma D P, Ba J. Adam: a method for stochastic optimization. In: Proceedings of the 3rd International Conference on Learning Representations. 2015
[69]
Kim S, Choi K, Choi H S, Lee B, Yoon S. Towards a rigorous evaluation of time-series anomaly detection. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. 2022, 7194−7201
[70]
Liu X, Hu J, Li Y, Diao S, Liang Y, Hooi B, Zimmermann R. UniTime: a language-empowered unified model for cross-domain time series forecasting. In: Proceedings of the ACM Web Conference 2024. 2024, 4095−4106
[71]
Kitaev N, Kaiser L, Levskaya A. Reformer: the efficient transformer. In: Proceedings of the 8th International Conference on Learning Representations. 2020
[72]
Behrouz A, Santacatterina M, Zabih R. MambaMixer: efficient selective state space models with dual token and channel selection. 2024, arXiv preprint arXiv: 2403.19888
[73]
Cai X, Zhu Y, Wang X, Yao Y. MambaTS: improved selective state space models for long-term time series forecasting. 2024, arXiv preprint arXiv: 2405.16440
Rights & permissions
The Author(s) 2026. This article is published with open access at link.springer.com and journal.hep.com.cn