fix(automl): include latest observation in time-series CV - #1616
Open
林SO (Linxiushen) wants to merge 1 commit into
Open
林SO (Linxiushen) wants to merge 1 commit into
林SO (Linxiushen) wants to merge 1 commit into
Conversation
Li Jiang (thinkall)
approved these changes
Sep 27, 2026
Li Jiang (thinkall)
left a comment
Collaborator
There was a problem hiding this comment.
Overall review of the complete PR: approved.
The exclusive CV endpoint now correctly includes the latest observation while preserving strict train-before-validation ordering, expanding-window behavior, custom step spacing, and parity with sklearn TimeSeriesSplit. Boundary and end-to-end coverage exercises small datasets, horizons, gaps/overlaps, fold counts, and leakage constraints. No blocking issues found.
Posted by thinkall-agent-auto-reviewer
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why are these changes needed?
Time-series cross-validation currently places every fold one row too early and excludes the final observation from validation.
cv_train_val_setsstarts fromlen(train_data) - 1, although its validation slice has an exclusive end. Use the sample count as the exclusive endpoint so the last fold reaches the latest observation. This agrees with theTimeSeriesSplitconfigured byTimeSeriesTaskand preserves customcv_step_sizespacing.A public
AutoML.fitreproduction with 59 zero-valued observations followed by a value of 100, the average forecaster, three folds and a five-step horizon reports MAE 0.0 before this change. It correctly reports 100 / 15 = 6.6666666667 afterward. The omitted observation can therefore affect model-selection scores.The regressions compare folds against scikit-learn for two horizons and three index types, exercise overlapping and separated validation windows, and run the real AutoML/Statsmodels forecast-and-score path. All nine new cases fail on unchanged
main.Related issue number
No existing issue found for this boundary error.
Validation
main: 9 failed, 3 existing tests passed.python -m pytest test/automl/test_ts_data.py -q: 12 passed on Windows with both Python 3.12 / NumPy 2.5.3 / pandas 3.0.6 / scikit-learn 1.9.1 / Statsmodels 0.15.0 and Python 3.10 / NumPy 1.26.4 / pandas 2.3.3 / scikit-learn 1.7.2 / Statsmodels 0.14.6.python -m pytest test/automl/test_forecast.py -k 'statsmodels or average_forecasters or simple_forecaster or log_training_metric' -q: 9 passed, 9 deselected.python -m pre_commit run --all-files --show-diff-on-failure: all hooks passed.Checks
Prepared with AI coding assistance; the fix and regressions were independently reviewed and executed locally.