A preprint posted on arXiv on Oct. 6 reports that choosing a time-series classifier on validation data with one missing-data pattern and deploying it on another reduced test balanced accuracy in a controlled study.

The result matters for model selection pipelines because the paper tests a common practical shortcut: validating on one form of incomplete data and assuming the chosen method will transfer when missing observations arrive in a different pattern at deployment, according to arXiv CS.LG.

According to the preprint, the study used a controlled 2x2 design on 64 univariate UCR datasets. Validation and test sets were masked with either random point missingness or circular block missingness at six rates from 5% to 30%, then imputed by linear interpolation to select among three prespecified classifiers: 1NN-DTW, MiniRocket with a Ridge classifier, and a statistical-feature Random Forest.

The paper says training data remained complete and compared selections made under matched and mismatched validation patterns on the same masked test sets. Across the 64 datasets, mismatched validation reduced the selected classifier's test balanced accuracy by 1.14 percentage points on average, with a 95% confidence interval of 0.79 to 1.51, and losses on 49 datasets, according to arXiv CS.LG.

The reported effect rose with missingness. The loss was negligible at 5% missingness and increased to 2.46 percentage points at 30%, the preprint says. It was concentrated in point-masked deployment, where the paper reports a 1.84 percentage-point loss; for block-masked deployment, the effect was described as small and not significant in the abstract on arXiv.

The preprint also says mismatch changed the selected classifier in 35.5% of paired comparisons, though a changed selection did not always reduce performance. A supplementary analysis using non-wrapping linear blocks reproduced the finding with a larger effect of 1.67 percentage points, according to the same arXiv listing.

The paper is a preprint and not peer reviewed, and the arXiv abstract provides no independent replication results.