NeurIPS 2026
1Seoul National University
†Corresponding author
Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging
due to the complex temporal dynamics and noise of real-world data. Existing approaches, however,
exhibit two limitations: missing and observed values are embedded within the same representation
space without explicit structural separation, and continuous diffusion-based methods are trained to
predict added noise rather than the original signal. To address these, we propose the
Masked Diffusion Time-series Imputation Model (MDTIM), which leverages the training paradigm of
masked diffusion models for imputation tasks. The [MASK] token is structurally orthogonal
to valid observations, and the model directly predicts the original values, naturally aligning both the
representation and the learning objective with the imputation task. To bridge the gap between discrete
masked diffusion and the continuous, ordinal nature of time series, we further introduce
Stochastic Discretization, which maps continuous values to ordinal-aware tokens while preserving
continuous dynamics. Our experiments on diverse benchmarks confirm that MDTIM achieves superior
robustness and scalability, consistently outperforming state-of-the-art deterministic and generative
baselines across various missing scenarios.
The absorbing [MASK] token is orthogonal to every observed value — the model always
knows what is missing, with no indicator heuristics.
MDTIM predicts the original values directly instead of added noise, matching the training objective to the imputation task.
Stochastic Discretization and soft labels preserve the continuous, ordinal structure of the signal inside a discrete token space.
From continuous values to ordinal-aware tokens. Observed values are
normalized and perturbed with bounded uniform noise, discretized into token bins, and supervised with
ordinal-aware soft labels centered on the true bin — keeping the quantization differentiable in effect
while the [MASK] state stays strictly separate. Training combines a soft-label
cross-entropy with a spectral (FFT) consistency loss, and continuous values are reconstructed as the
expectation over the predicted bin distribution.
Qualitative example. MDTIM reconstructions with calibrated uncertainty bands (1σ, 2σ) against ground truth on ETTh.
Main results. MAE on missing positions across four benchmarks under uniform and geometric masking at 30% and 70% missing rates. Best results in bold, second best underlined.
@inproceedings{kim2026discretizing,
title={Discretizing Continuous Time Series for Imputation with Masked Diffusion Training},
author={Kim, Dongbin and Lee, Seungyun and Shin, Geonwoo and Lee, Jaewook},
booktitle={Advances in Neural Information Processing Systems},
volume={39},
year={2026}
}