Skip to content

feat: add SEFR and SEFR-boost as opt-in classification estimators - #1613

Open
Hamidreza Keshavarz (hamidkm9) wants to merge 5 commits into
microsoft:mainfrom
hamidkm9:feat/sefr-estimator
Open

Hamidreza Keshavarz (hamidkm9) wants to merge 5 commits into
microsoft:mainfrom
hamidkm9:feat/sefr-estimator

Conversation

@hamidkm9

@hamidkm9 Hamidreza Keshavarz (hamidkm9) commented Sep 26, 2026 •

Copy link
Copy Markdown

Why are these changes needed?

FLAML's cheap linear slot for classification is lrl1, and a single
LogisticRegression(saga) fit is not interruptible, so a small time_budget
can be overrun by orders of magnitude (see #Issue 1612). This PR adds SEFR, a
linear-time, closed-form classifier, as an additional low-cost option.

SEFR derives one weight per feature plus a bias from class-conditional feature
means. Fitting is a closed form with no iterative optimization, and takes three
passes over the data whatever the number of classes: the feature range, the
per-class feature sums (one matrix product), and the spread of the training
scores. Beyond the scaled input it holds O(n_classes * n_features) memory.
Each binary SEFR head is n_features + 1 floats; the fitted classifier also
stores per-feature scaling parameters, one head per class for multiclass targets,
and one probability scale.

Keshavarz, Saniee Abadeh, Rawassizadeh (2020),
"SEFR: A Fast Linear-Time Classifier for Ultra-Low Power Devices", arXiv:2006.04620

Two estimators are registered, both opt-in and deliberately left out of
default_estimator_list, matching the existing treatment of histgb,
kneighbor and svc:

  • "sefr" — SEFR itself. Feature scaling is part of the estimator, since
    SEFR's weight formula needs non-negative features and FLAML does not scale
    anywhere in its pipeline. Dense input is min-max scaled and clipped to [0, 1]
    at predict time; sparse input is scaled by its column max and rejected if it
    has negative entries. Margins are calibrated because FLAML's default
    multiclass metric is log_loss and raw SEFR emits margins, not probabilities.
    Margins map to calibrated logits a * margin + b (one shared scale, one bias
    per head) and to their sigmoid or softmax, so predict, decision_function
    and predict_proba always agree. The default "sigmoid" is closed form;
    "platt" fits a and b by maximum likelihood and is in the search space.
    class_weight acts as misclassification costs on this calibrated decision.
    On iris (5-fold) log loss is 0.908 uncalibrated, 0.555 with sigmoid and
    0.209 with platt.
  • "sefr_boost" — AdaBoost over SEFR base learners, which gives the search
    a real but cheap-to-traverse space. It also tunes the base learner's
    calibration, starting at "platt".

No new dependencies. The algorithm is a handful of NumPy operations, and
FLAML's only required dependency is NumPy.

Budget adherence

Max wall-clock / requested time_budget over budgets of 1, 2, 5 and 10 seconds,
on the classification datasets in test/default/all/metafeatures.csv
(pandas 2.3.3, scikit-learn 1.5.2):

dataset n_train n_feat n_cls lrl1 lgbm sefr sefr_boost
car 1,296 6 2 1.00x 1.02x 1.00x 1.03x
dilbert 7,500 2,000 5 24.8x 1.74x 1.54x 1.56x
Amazon_employee_access 24,576 9 2 1.14x 1.67x 1.01x 1.41x
adult 36,631 14 2 1.25x 1.71x 1.01x 1.13x
connect-4 50,667 126 3 8.47x 1.54x 1.05x 1.15x
Dionis 312,141 60 355 770x 3.33x 1.68x 18.1x
Albert 318,930 78 2 10.3x 1.21x 1.12x 2.70x
Airlines 404,537 7 2 5.16x 1.10x 1.04x 1.43x
poker 768,757 10 2 5.51x 1.24x 1.04x 1.45x

sefr stays within 1.12x on 7 of 9 datasets and within 1.7x on all 9. With 355
classes, a Dionis fit takes 1.2s (sigmoid) or 4.5s (platt, whose likelihood
fit uses a row subsample once the margin matrix passes 1e7 entries).
sefr_boost fits a SEFR model per boosting round, so on Dionis it still
overruns small budgets (18.1x at 2s, 3.8x at 10s).

On accuracy, plainly

SEFR does not beat the tree learners and this PR does not claim it does.

Binary, test AUC at a 10s budget:

dataset sefr sefr_boost lrl1 lgbm
car 0.941 0.947 0.941 1.000
adult 0.846 0.884 0.506 0.913
Amazon_employee_access 0.580 0.591 0.518 0.842
Albert 0.630 0.676 0.596 0.758
Airlines 0.596 0.619 0.578 0.721
poker 0.502 0.501 0.502 0.931

Multiclass, negative log loss at a 10s budget (higher is better):

dataset sefr sefr_boost lrl1 lgbm
connect-4 -0.657 -0.606 -0.593 -0.384
dilbert -1.181 -0.811 -0.238 -0.124
Dionis -2.170 -2.204 -5.872 -2.083

SEFR clearly beats the existing cheap linear slot on 4 of 6 binary tasks, is
indistinguishable from it on poker (both at chance) and on car (0.94098 vs
0.94113), and loses to it on 2 of 3 multiclass tasks. On Dionis it comes within
0.09 of lgbm; elsewhere lgbm is clearly better.

sefr_boost beat the default pipeline on dilbert (2-10s) and connect-4 (1s). It
also scores higher on Dionis, but only while overrunning the budget (3.8x-18x),
so that is not a like-for-like win. "sefr" itself has an effectively inert search space under
roc_auc and is best understood as a constant-time baseline; "sefr_boost" is
the variant that improves with budget.

This is why both are opt-in rather than defaults.

Tests

test/automl/test_model.py adds 12 unit tests, including checks of the fitted
model against the closed form of Eqs. 3-9 recomputed independently (binary, and
each one-vs-rest head for multiclass), plus sample_weight (its effect on
calibration, zero and invalid weights), class_weight as costs, the binary Platt
intercept against LogisticRegression, the multiclass calibration and its row
subsample, sparse/dense equivalence and rejection of negative sparse input, and
sklearn's check_estimator suite, which passes in full on scikit-learn 1.0.2,
1.5.2 and 1.9.1. sefr_boost passes the SEFR learner as base_estimator on
scikit-learn < 1.2.
test/automl/test_extra_models.py adds two integration tests alongside the
other opt-in estimators.

Disclosure

I am an author of the SEFR paper. Happy to drop this if maintainers would rather
not carry it.

Related issue number

Related to #ISSUE 1612: #1612

Checks

SEFR derives one weight per feature plus a bias from class-conditional
feature means, so fitting is a single O(n_samples * n_features) pass with
no iterative optimization and the model is n_features + 1 floats.

  Keshavarz, Saniee Abadeh, Rawassizadeh (2020), arXiv:2006.04620

This makes it the cheapest learner in the portfolio and the one most
likely to return a model under a very small time_budget on large data,
where a single un-interruptible lrl1 fit can overrun the budget.

Implemented in NumPy rather than added as a dependency: the algorithm is
a handful of array operations and FLAML's only required dependency is
NumPy. Two estimators are registered:

  - "sefr": SEFR itself. Feature scaling is part of the estimator, since
    SEFR's weight formula assumes non-negative features and FLAML does
    not scale anywhere in its pipeline. Margins are calibrated (Platt by
    default) because FLAML's default multiclass metric is log_loss.
  - "sefr_boost": AdaBoost over SEFR base learners, which gives the
    search a real but cheap-to-traverse space.

Both are opt-in and deliberately left out of default_estimator_list:
they help under short budgets on small and medium data, but dilute the
budget away from lgbm on large data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@hamidkm9

Copy link
Copy Markdown
Author
@microsoft-github-policy-service agree

@microsoft-github-policy-service agree

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete PR: changes are required.

  1. flaml/automl/contrib/sefr.py:107,234: signed sparse inputs silently use max-abs scaling, violating SEFR's required [0,1] domain and reversing predictions. Remove unsafe modes or implement/reject with guaranteed non-negative transforms.
  2. flaml/automl/contrib/sefr.py:137-158: sample weights affect the SEFR head but are ignored by probability calibration; zero-weight classes can also divide by zero. Validate weights and pass effective weights into calibration.
  3. flaml/automl/contrib/sefr.py:314: AdaBoostClassifier(estimator=...) is incompatible with FLAML's supported scikit-learn 1.0/1.1 floor. Support base_estimator for older APIs or raise the minimum consistently.
  4. flaml/automl/contrib/sefr.py:204: fitted-state, input, feature-count, sample-weight, NaN/inf, and sparse capability validation do not satisfy basic sklearn estimator contracts. Use sklearn validation helpers and accurate tags.
  5. flaml/automl/contrib/sefr.py:157 and related docs: default Platt calibration is iterative and multiclass/scaler/calibrator state exceeds the advertised single-pass and n_features + 1 model-size guarantees. Correct the implementation default or the claims.

Posted by thinkall-agent-auto-reviewer

- Keep SEFR inputs in its non-negative domain. The unsafe "maxabs" mode is
  removed; dense input is min-max scaled and clipped to [0, 1] at predict
  time, sparse input is scaled by its column max and rejected if it has
  negative entries, and scaling="none" rejects negative features.
- Validate sample_weight (shape, finite, non-negative), reject classes whose
  total weight is zero, and pass the effective weights into calibration and
  the scaler's feature range.
- Pass the SEFR base learner to AdaBoostClassifier as base_estimator on
  scikit-learn < 1.2, matching FLAML's >= 1.0 floor.
- Use sklearn validation (validate_data / _validate_data, check_is_fitted,
  check_classification_targets), accept dict class_weight, and declare
  sparse and positive-only tags. Passes check_estimator on scikit-learn
  1.0.2 and 1.9.1 except check_class_weight_classifiers, which SEFR cannot
  meet by construction and is listed as an expected failure.
- Calibration now uses one sigmoid scale shared by all heads, so predict,
  decision_function and predict_proba agree. The closed-form "sigmoid"
  mode is the default; "platt" is documented as iterative.
- Correct the single-pass and n_features + 1 claims in the docstrings and
  docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The shared-scale calibration from the previous commit kept predict,
decision_function and predict_proba consistent but dropped the per-class
intercepts of the old per-head Platt, which cost a lot of multiclass
log_loss (connect-4 0.656 -> 0.781, Dionis 3.07 -> 3.65).

Multiclass margins now map to calibrated logits a * m + b with one shared
scale and one bias per class, and to softmax(a * m + b); decision_function
returns these logits, so all three methods still agree. "sigmoid" stays
closed form (a = 1/std, b = log weighted class prior); "platt" fits a and
b by weighted maximum likelihood (L-BFGS), on a deterministic row
subsample that keeps every class once the margins exceed 2e7 entries.
Binary behaviour is unchanged.

sefr_boost now tunes the base learner's calibration, starting at "platt",
since AdaBoost's SAMME.R uses the base learner's probabilities.

On the budget sweep (10s budget) sefr's Dionis log loss goes from 3.069
to 2.163 and its max overrun from 18.5x to 8.6x; sefr_boost's goes from
3.049 to 2.185 and 81.9x to 47.2x.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@hamidkm9

Copy link
Copy Markdown
Author

Overall review of the complete PR: changes are required.

  1. flaml/automl/contrib/sefr.py:107,234: signed sparse inputs silently use max-abs scaling, violating SEFR's required [0,1] domain and reversing predictions. Remove unsafe modes or implement/reject with guaranteed non-negative transforms.
  2. flaml/automl/contrib/sefr.py:137-158: sample weights affect the SEFR head but are ignored by probability calibration; zero-weight classes can also divide by zero. Validate weights and pass effective weights into calibration.
  3. flaml/automl/contrib/sefr.py:314: AdaBoostClassifier(estimator=...) is incompatible with FLAML's supported scikit-learn 1.0/1.1 floor. Support base_estimator for older APIs or raise the minimum consistently.
  4. flaml/automl/contrib/sefr.py:204: fitted-state, input, feature-count, sample-weight, NaN/inf, and sparse capability validation do not satisfy basic sklearn estimator contracts. Use sklearn validation helpers and accurate tags.
  5. flaml/automl/contrib/sefr.py:157 and related docs: default Platt calibration is iterative and multiclass/scaler/calibrator state exceeds the advertised single-pass and n_features + 1 model-size guarantees. Correct the implementation default or the claims.

Posted by thinkall-agent-auto-reviewer

Thanks Li Jiang (@thinkall) for the review. All five points are addressed in the new commits:

  1. Signed sparse input. maxabs is removed. Dense input is min-max scaled and clipped to [0, 1] at predict time. Sparse input is scaled by column max and rejected with a clear error if it has negative entries, since it can't be shifted without densifying. scaling="none" rejects negative features. scaling is no longer in the search space, since only minmax is safe for arbitrary input.
  2. Sample weights. sample_weight is validated for shape, finite values and non-negativity, and classes with zero total weight raise instead of dividing by zero. The effective weights, including class_weight, now reach the calibration and the scaler's feature range. A zero weight gives the same fit as dropping the sample, and there's a test for that.
  3. scikit-learn 1.0/1.1. AdaBoost gets base_estimator on scikit-learn < 1.2 and estimator otherwise. Verified on 1.0.2.
  4. sklearn contracts. It now uses validate_data/_validate_data, check_is_fitted and check_classification_targets, accepts dict class_weight, and declares sparse and positive-only tags. A new test runs check_estimator, which passes on 1.0.2, 1.5.2 and 1.9.1 except check_class_weight_classifiers. That one is a documented expected failure: SEFR's class weights cancel in the per-class means and only move the Eq. 9 threshold, so it can't reach the 87% majority the check expects.
  5. Calibration and claims. Binary margins map to sigmoid(a·m). Multiclass margins map to softmax(a·m + b_k), with one shared scale and a bias per class, and decision_function returns those logits. So predict, decision_function and predict_proba always agree, which wasn't guaranteed before. The default sigmoid is closed form (a from the margin spread, b_k the log class prior). platt fits a and b_k by maximum likelihood and is documented as iterative. The docstrings, docs and PR description no longer claim single-pass or n_features + 1 for the full model.
    I re-ran the budget sweep on the current code and updated the tables in the description. Binary results are unchanged. On Dionis (355 classes), sefr log loss improved from 3.07 to 2.16, within 0.08 of lgbm, and its worst overrun dropped from 18.5x to 8.6x.

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: changes are still required.

  1. flaml/automl/contrib/sefr.py:291: binary Platt calibration disables the intercept, so it cannot learn imbalanced or weighted class priors. Fit and retain a binary intercept and derive prediction consistently from the calibrated score.
  2. flaml/automl/contrib/sefr.py:305: calibration subsampling can preserve a zero-weight class row while omitting every positive-weight row for that class. Preserve positive-weight coverage and revalidate effective class totals after subsampling.
  3. flaml/automl/contrib/sefr.py:322: advertised class_weight does not materially affect the decision rule and the standard sklearn contract check is suppressed. Implement it correctly or remove it from the public API/search space.
  4. flaml/automl/contrib/sefr.py:179: sparse support is not declared through legacy _more_tags() on supported sklearn 1.0-1.5, so meta-estimators treat SEFR as dense-only.
  5. flaml/automl/contrib/sefr.py:286 and documentation: multiclass fitting materializes multiple n_samples × n_classes arrays and performs a full pass per class, contradicting the constant-pass claim and creating multi-gigabyte allocations on documented datasets. Subsample/stream before materialization and document actual complexity.

Posted by thinkall-agent-auto-reviewer

…att intercept

Addresses the second review round:

- Multiclass fitting no longer copies X per class or materializes
  n_samples x n_classes arrays. Per-class feature sums come from one
  sparse one-hot matrix product, every one-vs-rest head and its Eq. 9
  bias follow from them in closed form, and score spreads are
  accumulated in row chunks. Fitting takes three passes over the data
  whatever the number of classes and reproduces the previous model to
  ~1e-14. On 355-class Dionis a fit drops from 13.3s / 4.27 GB peak to
  1.2s / 0.42 GB (sigmoid) and from 17.7s to 4.5s / 0.64 GB (platt).
- class_weight now acts as misclassification costs on the calibrated
  decision (log(c) added to each logit in closed form; cost-weighted
  likelihood for platt) instead of cancelling in the class means.
  check_class_weight_classifiers now passes, so check_estimator runs
  without exemptions.
- Binary platt fits an intercept as well as a slope, and predict follows
  the calibrated score. Platt uses a small fixed penalty on the slope
  only, which keeps separable data finite and is invariant to rescaling
  the weights (AdaBoost normalizes them).
- The calibration subsample keeps a positive-weight row of every class
  and revalidates class totals on the subsample.
- _more_tags declares sparse input for scikit-learn < 1.6.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@hamidkm9

Copy link
Copy Markdown
Author

Overall review of the complete current PR: changes are still required.

  1. flaml/automl/contrib/sefr.py:291: binary Platt calibration disables the intercept, so it cannot learn imbalanced or weighted class priors. Fit and retain a binary intercept and derive prediction consistently from the calibrated score.
  2. flaml/automl/contrib/sefr.py:305: calibration subsampling can preserve a zero-weight class row while omitting every positive-weight row for that class. Preserve positive-weight coverage and revalidate effective class totals after subsampling.
  3. flaml/automl/contrib/sefr.py:322: advertised class_weight does not materially affect the decision rule and the standard sklearn contract check is suppressed. Implement it correctly or remove it from the public API/search space.
  4. flaml/automl/contrib/sefr.py:179: sparse support is not declared through legacy _more_tags() on supported sklearn 1.0-1.5, so meta-estimators treat SEFR as dense-only.
  5. flaml/automl/contrib/sefr.py:286 and documentation: multiclass fitting materializes multiple n_samples × n_classes arrays and performs a full pass per class, contradicting the constant-pass claim and creating multi-gigabyte allocations on documented datasets. Subsample/stream before materialization and document actual complexity.

Posted by thinkall-agent-auto-reviewer

Thanks Li Jiang (@thinkall). All five points are addressed in 7fdd46f:

  1. Binary Platt intercept. Binary platt now fits a slope and an intercept by weighted maximum likelihood, and predict follows the calibrated score. A test checks it against LogisticRegression. A small fixed penalty on the slope only keeps separable data finite and doesn't depend on the scale of the weights (AdaBoost normalizes them).
  2. Subsample coverage. The calibration subsample keeps a positive-weight row of every class and re-checks class totals on the subsample. A test covers the case where every row on the subsample grid, and the first row of each class, have zero weight.
  3. class_weight. It is now implemented as misclassification costs on the calibrated decision. The closed-form calibration adds log(c) to each class's logit, and platt weights the likelihood by class cost. It no longer cancels in the class means. check_class_weight_classifiers now passes, so check_estimator runs with no exemptions on scikit-learn 1.0.2, 1.5.2 and 1.9.1.
  4. Legacy tags. _more_tags() declares X_types: ["2darray", "sparse"] for scikit-learn < 1.6.
  5. Memory and passes. Per-class sums come from one sparse one-hot matrix product. Every one-vs-rest head and its Eq. 9 bias follow in closed form, and score spreads are accumulated in row chunks. Fitting takes three passes over the data whatever the number of classes, and no n_samples × n_classes array exists except the Platt block, capped at 1e7 entries. It reproduces the previous model to about 1e-14. On Dionis (312k × 60, 355 classes), a fit went from 13.3s and 4.27 GB peak to 1.2s and 0.42 GB (sigmoid), and from 17.7s to 4.5s and 0.64 GB (platt). The docstrings and docs now describe this cost.
    I re-ran the budget sweep and updated the tables in the description. sefr now stays within 1.7x of the budget on all nine datasets (Dionis went from 8.6x to 1.7x).

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: changes are still required.

  1. flaml/automl/contrib/sefr.py:304: a model fitted on dense data learns a scaler offset, but sparse prediction input skips that offset. Identical values can produce different scores and labels solely from ndarray versus CSR representation. Apply the fitted transformation consistently or reject unsupported cross-format input.
  2. flaml/automl/contrib/sefr.py:239: eps and threshold_shift are not validated. Zero/NaN values can fit successfully with non-finite coefficients or probabilities. Require finite positive eps and finite threshold_shift.

Posted by thinkall-agent-auto-reviewer

…d_shift

A model fitted on dense data with a nonzero feature minimum now applies
its fitted offset to sparse prediction input (by densifying it) instead
of skipping it, so dense and sparse input always score identically.

eps must be a finite positive number and threshold_shift a finite
number; anything else raises ValueError.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@hamidkm9

Copy link
Copy Markdown
Author

Thanks Li Jiang (@thinkall). Both points are addressed in 90bc8be:

  1. Dense fit, sparse predict. A model fitted on dense data with a nonzero feature minimum now applies its fitted offset to sparse input too, by densifying it, since subtracting the offset would densify it anyway. Dense and sparse input therefore always get identical scores. A test checks this for zero and nonzero minima, including values outside the training range.
  2. Parameter validation. eps must be a finite positive number and threshold_shift a finite number. Anything else raises a ValueError. A test covers zero, negative, NaN, infinite, string, None and boolean values.

Is there anything else blocking this from being merged?

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: approved.

The final two blockers are resolved. A dense-fitted SEFR model now applies its learned offset to sparse prediction input before scaling, and the new regression tests compare identical dense/CSR probes with zero and nonzero fitted minima. eps now requires a finite positive real value and threshold_shift a finite real value, with tests for invalid and valid NumPy-scalar inputs. The previously reviewed calibration, weighting, sklearn-version, tag, and multiclass-resource fixes remain unchanged. No blocking code issues found. Build-matrix CI is still running on this head.

Posted by thinkall-agent-auto-reviewer

@hamidkm9

Copy link
Copy Markdown
Author

Li Jiang (@thinkall) The only failing check is build (ubuntu-latest, 3.12), which timed out in test/tune/ (unrelated to this PR). The same tune tests pass on the other 11 jobs, and all SEFR tests passed in that job before the timeout. Could you re-run the failed job?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants