feat: add SEFR and SEFR-boost as opt-in classification estimators - #1613
Hamidreza Keshavarz (hamidkm9) wants to merge 5 commits into
Conversation
SEFR derives one weight per feature plus a bias from class-conditional
feature means, so fitting is a single O(n_samples * n_features) pass with
no iterative optimization and the model is n_features + 1 floats.
Keshavarz, Saniee Abadeh, Rawassizadeh (2020), arXiv:2006.04620
This makes it the cheapest learner in the portfolio and the one most
likely to return a model under a very small time_budget on large data,
where a single un-interruptible lrl1 fit can overrun the budget.
Implemented in NumPy rather than added as a dependency: the algorithm is
a handful of array operations and FLAML's only required dependency is
NumPy. Two estimators are registered:
- "sefr": SEFR itself. Feature scaling is part of the estimator, since
SEFR's weight formula assumes non-negative features and FLAML does
not scale anywhere in its pipeline. Margins are calibrated (Platt by
default) because FLAML's default multiclass metric is log_loss.
- "sefr_boost": AdaBoost over SEFR base learners, which gives the
search a real but cheap-to-traverse space.
Both are opt-in and deliberately left out of default_estimator_list:
they help under short budgets on small and medium data, but dilute the
budget away from lgbm on large data.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@microsoft-github-policy-service agree |
Li Jiang (thinkall)
left a comment
There was a problem hiding this comment.
Overall review of the complete PR: changes are required.
flaml/automl/contrib/sefr.py:107,234: signed sparse inputs silently use max-abs scaling, violating SEFR's required[0,1]domain and reversing predictions. Remove unsafe modes or implement/reject with guaranteed non-negative transforms.flaml/automl/contrib/sefr.py:137-158: sample weights affect the SEFR head but are ignored by probability calibration; zero-weight classes can also divide by zero. Validate weights and pass effective weights into calibration.flaml/automl/contrib/sefr.py:314:AdaBoostClassifier(estimator=...)is incompatible with FLAML's supported scikit-learn 1.0/1.1 floor. Supportbase_estimatorfor older APIs or raise the minimum consistently.flaml/automl/contrib/sefr.py:204: fitted-state, input, feature-count, sample-weight, NaN/inf, and sparse capability validation do not satisfy basic sklearn estimator contracts. Use sklearn validation helpers and accurate tags.flaml/automl/contrib/sefr.py:157and related docs: default Platt calibration is iterative and multiclass/scaler/calibrator state exceeds the advertised single-pass andn_features + 1model-size guarantees. Correct the implementation default or the claims.
Posted by thinkall-agent-auto-reviewer
- Keep SEFR inputs in its non-negative domain. The unsafe "maxabs" mode is removed; dense input is min-max scaled and clipped to [0, 1] at predict time, sparse input is scaled by its column max and rejected if it has negative entries, and scaling="none" rejects negative features. - Validate sample_weight (shape, finite, non-negative), reject classes whose total weight is zero, and pass the effective weights into calibration and the scaler's feature range. - Pass the SEFR base learner to AdaBoostClassifier as base_estimator on scikit-learn < 1.2, matching FLAML's >= 1.0 floor. - Use sklearn validation (validate_data / _validate_data, check_is_fitted, check_classification_targets), accept dict class_weight, and declare sparse and positive-only tags. Passes check_estimator on scikit-learn 1.0.2 and 1.9.1 except check_class_weight_classifiers, which SEFR cannot meet by construction and is listed as an expected failure. - Calibration now uses one sigmoid scale shared by all heads, so predict, decision_function and predict_proba agree. The closed-form "sigmoid" mode is the default; "platt" is documented as iterative. - Correct the single-pass and n_features + 1 claims in the docstrings and docs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The shared-scale calibration from the previous commit kept predict, decision_function and predict_proba consistent but dropped the per-class intercepts of the old per-head Platt, which cost a lot of multiclass log_loss (connect-4 0.656 -> 0.781, Dionis 3.07 -> 3.65). Multiclass margins now map to calibrated logits a * m + b with one shared scale and one bias per class, and to softmax(a * m + b); decision_function returns these logits, so all three methods still agree. "sigmoid" stays closed form (a = 1/std, b = log weighted class prior); "platt" fits a and b by weighted maximum likelihood (L-BFGS), on a deterministic row subsample that keeps every class once the margins exceed 2e7 entries. Binary behaviour is unchanged. sefr_boost now tunes the base learner's calibration, starting at "platt", since AdaBoost's SAMME.R uses the base learner's probabilities. On the budget sweep (10s budget) sefr's Dionis log loss goes from 3.069 to 2.163 and its max overrun from 18.5x to 8.6x; sefr_boost's goes from 3.049 to 2.185 and 81.9x to 47.2x. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Thanks Li Jiang (@thinkall) for the review. All five points are addressed in the new commits:
|
Li Jiang (thinkall)
left a comment
There was a problem hiding this comment.
Overall review of the complete current PR: changes are still required.
flaml/automl/contrib/sefr.py:291: binary Platt calibration disables the intercept, so it cannot learn imbalanced or weighted class priors. Fit and retain a binary intercept and derive prediction consistently from the calibrated score.flaml/automl/contrib/sefr.py:305: calibration subsampling can preserve a zero-weight class row while omitting every positive-weight row for that class. Preserve positive-weight coverage and revalidate effective class totals after subsampling.flaml/automl/contrib/sefr.py:322: advertisedclass_weightdoes not materially affect the decision rule and the standard sklearn contract check is suppressed. Implement it correctly or remove it from the public API/search space.flaml/automl/contrib/sefr.py:179: sparse support is not declared through legacy_more_tags()on supported sklearn 1.0-1.5, so meta-estimators treat SEFR as dense-only.flaml/automl/contrib/sefr.py:286and documentation: multiclass fitting materializes multiplen_samples × n_classesarrays and performs a full pass per class, contradicting the constant-pass claim and creating multi-gigabyte allocations on documented datasets. Subsample/stream before materialization and document actual complexity.
Posted by thinkall-agent-auto-reviewer
…att intercept Addresses the second review round: - Multiclass fitting no longer copies X per class or materializes n_samples x n_classes arrays. Per-class feature sums come from one sparse one-hot matrix product, every one-vs-rest head and its Eq. 9 bias follow from them in closed form, and score spreads are accumulated in row chunks. Fitting takes three passes over the data whatever the number of classes and reproduces the previous model to ~1e-14. On 355-class Dionis a fit drops from 13.3s / 4.27 GB peak to 1.2s / 0.42 GB (sigmoid) and from 17.7s to 4.5s / 0.64 GB (platt). - class_weight now acts as misclassification costs on the calibrated decision (log(c) added to each logit in closed form; cost-weighted likelihood for platt) instead of cancelling in the class means. check_class_weight_classifiers now passes, so check_estimator runs without exemptions. - Binary platt fits an intercept as well as a slope, and predict follows the calibrated score. Platt uses a small fixed penalty on the slope only, which keeps separable data finite and is invariant to rescaling the weights (AdaBoost normalizes them). - The calibration subsample keeps a positive-weight row of every class and revalidates class totals on the subsample. - _more_tags declares sparse input for scikit-learn < 1.6. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Thanks Li Jiang (@thinkall). All five points are addressed in 7fdd46f:
|
Li Jiang (thinkall)
left a comment
There was a problem hiding this comment.
Overall review of the complete current PR: changes are still required.
flaml/automl/contrib/sefr.py:304: a model fitted on dense data learns a scaler offset, but sparse prediction input skips that offset. Identical values can produce different scores and labels solely from ndarray versus CSR representation. Apply the fitted transformation consistently or reject unsupported cross-format input.flaml/automl/contrib/sefr.py:239:epsandthreshold_shiftare not validated. Zero/NaN values can fit successfully with non-finite coefficients or probabilities. Require finite positiveepsand finitethreshold_shift.
Posted by thinkall-agent-auto-reviewer
…d_shift A model fitted on dense data with a nonzero feature minimum now applies its fitted offset to sparse prediction input (by densifying it) instead of skipping it, so dense and sparse input always score identically. eps must be a finite positive number and threshold_shift a finite number; anything else raises ValueError. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Thanks Li Jiang (@thinkall). Both points are addressed in 90bc8be:
Is there anything else blocking this from being merged? |
Li Jiang (thinkall)
left a comment
There was a problem hiding this comment.
Overall review of the complete current PR: approved.
The final two blockers are resolved. A dense-fitted SEFR model now applies its learned offset to sparse prediction input before scaling, and the new regression tests compare identical dense/CSR probes with zero and nonzero fitted minima. eps now requires a finite positive real value and threshold_shift a finite real value, with tests for invalid and valid NumPy-scalar inputs. The previously reviewed calibration, weighting, sklearn-version, tag, and multiclass-resource fixes remain unchanged. No blocking code issues found. Build-matrix CI is still running on this head.
Posted by thinkall-agent-auto-reviewer
|
Li Jiang (@thinkall) The only failing check is |
Why are these changes needed?
FLAML's cheap linear slot for classification is
lrl1, and a singleLogisticRegression(saga)fit is not interruptible, so a smalltime_budgetcan be overrun by orders of magnitude (see #Issue 1612). This PR adds SEFR, a
linear-time, closed-form classifier, as an additional low-cost option.
SEFR derives one weight per feature plus a bias from class-conditional feature
means. Fitting is a closed form with no iterative optimization, and takes three
passes over the data whatever the number of classes: the feature range, the
per-class feature sums (one matrix product), and the spread of the training
scores. Beyond the scaled input it holds O(n_classes * n_features) memory.
Each binary SEFR head is
n_features + 1floats; the fitted classifier alsostores per-feature scaling parameters, one head per class for multiclass targets,
and one probability scale.
Two estimators are registered, both opt-in and deliberately left out of
default_estimator_list, matching the existing treatment ofhistgb,kneighborandsvc:"sefr"— SEFR itself. Feature scaling is part of the estimator, sinceSEFR's weight formula needs non-negative features and FLAML does not scale
anywhere in its pipeline. Dense input is min-max scaled and clipped to [0, 1]
at predict time; sparse input is scaled by its column max and rejected if it
has negative entries. Margins are calibrated because FLAML's default
multiclass metric is
log_lossand raw SEFR emits margins, not probabilities.Margins map to calibrated logits
a * margin + b(one shared scale, one biasper head) and to their sigmoid or softmax, so
predict,decision_functionand
predict_probaalways agree. The default"sigmoid"is closed form;"platt"fitsaandbby maximum likelihood and is in the search space.class_weightacts as misclassification costs on this calibrated decision.On iris (5-fold) log loss is 0.908 uncalibrated, 0.555 with
sigmoidand0.209 with
platt."sefr_boost"— AdaBoost over SEFR base learners, which gives the searcha real but cheap-to-traverse space. It also tunes the base learner's
calibration, starting at
"platt".No new dependencies. The algorithm is a handful of NumPy operations, and
FLAML's only required dependency is NumPy.
Budget adherence
Max wall-clock / requested
time_budgetover budgets of 1, 2, 5 and 10 seconds,on the classification datasets in
test/default/all/metafeatures.csv(pandas 2.3.3, scikit-learn 1.5.2):
sefrstays within 1.12x on 7 of 9 datasets and within 1.7x on all 9. With 355classes, a Dionis fit takes 1.2s (
sigmoid) or 4.5s (platt, whose likelihoodfit uses a row subsample once the margin matrix passes 1e7 entries).
sefr_boostfits a SEFR model per boosting round, so on Dionis it stilloverruns small budgets (18.1x at 2s, 3.8x at 10s).
On accuracy, plainly
SEFR does not beat the tree learners and this PR does not claim it does.
Binary, test AUC at a 10s budget:
Multiclass, negative log loss at a 10s budget (higher is better):
SEFR clearly beats the existing cheap linear slot on 4 of 6 binary tasks, is
indistinguishable from it on poker (both at chance) and on car (0.94098 vs
0.94113), and loses to it on 2 of 3 multiclass tasks. On Dionis it comes within
0.09 of
lgbm; elsewherelgbmis clearly better.sefr_boostbeat the default pipeline on dilbert (2-10s) and connect-4 (1s). Italso scores higher on Dionis, but only while overrunning the budget (3.8x-18x),
so that is not a like-for-like win.
"sefr"itself has an effectively inert search space underroc_aucand is best understood as a constant-time baseline;"sefr_boost"isthe variant that improves with budget.
This is why both are opt-in rather than defaults.
Tests
test/automl/test_model.pyadds 12 unit tests, including checks of the fittedmodel against the closed form of Eqs. 3-9 recomputed independently (binary, and
each one-vs-rest head for multiclass), plus
sample_weight(its effect oncalibration, zero and invalid weights),
class_weightas costs, the binary Plattintercept against
LogisticRegression, the multiclass calibration and its rowsubsample, sparse/dense equivalence and rejection of negative sparse input, and
sklearn's
check_estimatorsuite, which passes in full on scikit-learn 1.0.2,1.5.2 and 1.9.1.
sefr_boostpasses the SEFR learner asbase_estimatoronscikit-learn < 1.2.
test/automl/test_extra_models.pyadds two integration tests alongside theother opt-in estimators.
Disclosure
I am an author of the SEFR paper. Happy to drop this if maintainers would rather
not carry it.
Related issue number
Related to #ISSUE 1612: #1612
Checks