Skip to content

Add initial anomaly_detection task with IsolationForest - #1567

Open
Muhammad Rashid (PhD) (rashidrao-pk) wants to merge 18 commits into
microsoft:mainfrom
rashidrao-pk:add-anomaly-detection-task
Open

Muhammad Rashid (PhD) (rashidrao-pk) wants to merge 18 commits into
microsoft:mainfrom
rashidrao-pk:add-anomaly-detection-task

Conversation

@rashidrao-pk

Copy link
Copy Markdown

Closes #413

This draft PR adds initial support for anomaly detection in FLAML using IsolationForest.

Scope:

  • Add anomaly_detection task type
  • Add is_anomaly_detection()
  • Register isolation_forest
  • Add IsolationForestEstimator subclassing SKLearnEstimator
  • Expose predict, score_samples, and decision_function
  • Keep sklearn convention: 1 = normal, -1 = anomaly
  • Add unit test for synthetic point anomalies

Notes:

  • This first PR intentionally keeps the scope minimal to lock the API.
  • Full label-free AutoML.fit(X_train=..., task="anomaly_detection") integration can be completed after maintainer feedback.

cc Li Jiang (@thinkall)

Copilot AI review requested due to automatic review settings July 1, 2026 14:14
@rashidrao-pk

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

@rashidrao-pk

Copy link
Copy Markdown
Author

cc Kevin Chen (@int-chaos) for review, as discussed in #413.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR introduces an initial anomaly detection task type to FLAML, centered around a first estimator implementation based on scikit-learn’s IsolationForest, plus a basic unit test to validate anomaly scoring behavior.

Changes:

  • Adds a new task identifier anomaly_detection and a corresponding is_anomaly_detection() helper on Task.
  • Registers a new estimator name isolation_forest and adds an IsolationForestEstimator (subclassing SKLearnEstimator).
  • Adds a unit test using synthetic point anomalies to validate prediction output shape and anomaly ranking quality.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 4 comments.

File Description
test/automl/test_anomaly_detection.py Adds a unit test validating IsolationForest-based anomaly scoring on synthetic data.
flaml/automl/task/task.py Introduces the ANOMALY_DETECTION task constant and Task.is_anomaly_detection().
flaml/automl/task/generic_task.py Registers the isolation_forest estimator and wires anomaly detection defaults (estimator + metric).
flaml/automl/model.py Adds IsolationForestEstimator implementation and exposes score_samples / decision_function.

Comment thread flaml/automl/model.py Outdated
Comment thread flaml/automl/model.py
Comment thread flaml/automl/task/generic_task.py
Comment thread flaml/automl/task/generic_task.py
@rashidrao-pk

Copy link
Copy Markdown
Author

Thanks for the review. I will fix the search-space init value, formatting, and Spark-dataframe error path.

For the ap metric comment: I agree this needs proper handling through anomaly scores (score_samples) in the AutoML evaluation path. Since this draft PR intentionally does not yet implement full label-free AutoML.fit(...) integration, I propose to address the metric evaluation path in the next PR together with the AutoML integration.

@thinkall

Copy link
Copy Markdown
Collaborator

Hi Muhammad Rashid (PhD) (@rashidrao-pk) , could you add an e2e test for anomaly_detection? Thanks a lot!

@rashidrao-pk

Copy link
Copy Markdown
Author

Hi Li Jiang (@thinkall),

Thanks for the suggestion.

I've added an end-to-end test for anomaly_detection and pushed the updates to this PR.

The test now exercises the AutoML.fit(...) workflow using the new anomaly_detection task and verifies that the estimator is correctly integrated into the AutoML pipeline, rather than only testing IsolationForestEstimator in isolation.

In addition, I reran the test suite and the project's pre-commit checks locally:

  • ✅ pytest test/automl/test_anomaly_detection.py
  • ✅ pre-commit run --all-files

Please let me know if you'd like the e2e test to cover any additional scenarios or edge cases.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 3 comments.

Comment thread test/automl/test_anomaly_detection.py Outdated
Comment on lines +5 to +6
from flaml import AutoML
from flaml.automl.model import IsolationForestEstimator
Comment on lines +1094 to +1096
elif self.is_anomaly_detection():
assert split_type in ["auto", "uniform", "time", "group"]
return split_type if split_type != "auto" else "uniform"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Muhammad Rashid (PhD) (@rashidrao-pk) , could you address this comment? Thanks.

Comment on lines +1377 to +1378
elif self.is_anomaly_detection():
return "ap"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Muhammad Rashid (PhD) (@rashidrao-pk) , could you address this comment? Thanks.

@thinkall
Li Jiang (thinkall) requested a balanced review from Copilot August 19, 2026 07:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated no new comments.

Suppressed comments (5)

flaml/automl/task/generic_task.py:1096

  • The new split type is accepted, but holdout preparation only creates a validation split for classification or regression (prepare_data() lines 969–1016). For anomaly detection without an explicit X_val, state.X_val remains None, and evaluation subsequently calls predict(None). Extend the uniform holdout branch to split anomaly-detection data as well.
        elif self.is_anomaly_detection():
            assert split_type in ["auto", "uniform", "time", "group"]
            return split_type if split_type != "auto" else "uniform"

flaml/automl/task/generic_task.py:1378

  • ap does not currently evaluate anomaly scores for this task. get_y_pred() only uses probability scores for binary tasks, so anomaly detection falls through to predict() and supplies hard labels where 1 means normal and -1 means anomalous. With the 0/1 anomaly labels used by the added test, average precision therefore ranks normals as positives. Add anomaly-specific prediction/label polarity handling (for example, using negated score_samples) before making ap the default.
        elif self.is_anomaly_detection():
            return "ap"

test/automl/test_anomaly_detection.py:6

  • This import is unused and will be flagged by the repository's Ruff F401 check. Remove it unless the test directly instantiates the estimator.
from flaml.automl.model import IsolationForestEstimator

flaml/automl/model.py:1531

  • IsolationForest supports n_jobs, so removing it forces every forest fit to run serially and ignores FLAML's configured worker count. Preserve this parameter as the other sklearn forest estimators do.
        params.pop("n_jobs", None)

test/automl/test_anomaly_detection.py:64

  • The PR adds decision_function() as part of the public anomaly-detection API, but this end-to-end test only exercises predict() and score_samples(). Add a call and shape assertion so regressions in the third exposed method are covered.
    preds = automl.predict(X_val)
    scores = automl.model.score_samples(X_val)

    assert preds.shape == y_val.shape
    assert scores.shape == y_val.shape

@thinkall

Copy link
Copy Markdown
Collaborator

Hi Muhammad Rashid (PhD) (@rashidrao-pk) , could you help addressing the latest comments? Thanks.

@rashidrao-pk

Copy link
Copy Markdown
Author

Hi Li Jiang (@thinkall),

Thanks for the follow-up. I've addressed the latest review comments and pushed the updates.

The changes now include:

  • anomaly-specific score handling for ap / roc_auc using continuous anomaly scores,
  • holdout splitting support for anomaly_detection,
  • preserving n_jobs for IsolationForest,
  • removal of the unused estimator import,
  • additional e2e coverage for decision_function().

I also verified locally that:

  • pytest test/automl/test_anomaly_detection.py -vv passes
  • pre-commit run --all-files passes
  • git diff --check is clean

Please let me know if any further adjustments are needed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Suppressed comments (1)

flaml/automl/model.py:1552

  • AutoML.score() delegates to this wrapper's inherited BaseEstimator.score(). With no metric, that method calls IsolationForest.score(), which does not exist; with metric="ap", it evaluates hard predict() outputs instead of the continuous -score_samples() used during tuning. Add anomaly-aware score handling so public scoring does not fail or disagree with the optimization metric.
    def score_samples(self, X):
        X = self._preprocess(X)
        return self._model.score_samples(X)

Comment on lines +1307 to +1310
if self.is_anomaly_detection():
if is_spark_dataframe:
raise ValueError("anomaly_detection does not support Spark dataframes yet. Use numpy/pandas data.")
estimator_list = ["isolation_forest"]
Comment thread flaml/automl/ml.py
Comment on lines +296 to +297
if task.is_anomaly_detection() and eval_metric in ["ap", "roc_auc"]:
y_pred = -estimator.score_samples(X)
Comment thread test/automl/test_anomaly_detection.py Outdated
Comment on lines +42 to +43
scores = automl.model.score_samples(X_val)
decision_scores = automl.model.decision_function(X_val)
Comment thread flaml/automl/model.py Outdated
Comment on lines +1520 to +1521
def size(cls, config):
return config.get("n_estimators", 100)

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The anomaly-detection integration is not ready to merge yet:

  1. The advertised label-free flow is unreachable: AutoML.fit(X_train=X, task="anomaly_detection") is rejected by generic validation before IsolationForestEstimator.fit() can ignore y_train. Define and implement X-only fitting semantics, including how model selection works without labeled validation data.
  2. Explicit estimator lists are not task-validated, so classification accepts isolation_forest and anomaly detection accepts learners such as rf; both fail later through missing predict_proba()/score_samples() APIs.
  3. Public AutoML.score() either calls nonexistent IsolationForest.score() or evaluates AP on hard {-1, 1} predictions, disagreeing with the continuous anomaly scores used during tuning.
  4. score_samples() and decision_function() bypass the fitted DataTransformer; add AutoML-level forwarding APIs so they accept the same raw inputs as predict().
  5. Restrict or correctly implement explicit anomaly metrics. Several currently accepted metrics use inverted hard labels or unsupported probability APIs.
  6. Strengthen the E2E test to verify anomaly ranking quality and cover X-only fitting/public scoring, not only output shape and label domain.

@rashidrao-pk

Copy link
Copy Markdown
Author

Thanks for the detailed review. I’ve pushed an update addressing the requested anomaly-detection issues.

The changes now:

  • support label-free AutoML.fit(X_train=X, task="anomaly_detection") using a one-shot initial configuration without fabricating labels;
  • reject unlabeled HPO requests that would require a supervised validation objective;
  • validate estimator/task compatibility in both directions;
  • add public AutoML.score_samples() and decision_function() APIs with the same preprocessing path as predict();
  • make AutoML.score() use continuous anomaly scores and preserve the fitted metric;
  • restrict built-in anomaly metrics to ap and roc_auc;
  • normalize both {0,1} and IsolationForest-style {-1,1} anomaly labels consistently;
  • strengthen E2E tests for ranking quality, X-only fitting, public scoring, metric validation, and estimator compatibility.

The targeted regression suite currently passes:
25 passed, 1 deselected.

Commit: 5b892aa

Thanks again for the review — happy to make any further adjustments.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Label-free training-metric logging can fail, and single-class labels currently permit meaningless hyperparameter optimization.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review details
  • Files reviewed: 8/8 changed files
  • Comments generated: 2
  • Review effort level: Balanced

Comment thread flaml/automl/automl.py Outdated
Comment on lines +2429 to +2433
unlabeled_anomaly = (
task.is_anomaly_detection()
and self._y_train_all is None
and self._state.y_val is None
)
Comment thread flaml/automl/ml.py
Comment on lines +653 to +656
y_train_for_metric = (
normalize_anomaly_labels(y_train)
if task.is_anomaly_detection() and eval_metric in ["ap", "roc_auc"]
else y_train

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed the commits since my last review. The earlier integration issues are substantially improved, but three blockers remain:

  1. flaml/automl/automl.py:2431: single-class anomaly labels are still accepted for HPO, producing meaningless optimization. Validate the effective holdout/CV labels and require both normalized classes before search.
  2. flaml/automl/ml.py:653: supported X-only training with labeled validation crashes when log_training_metric=True because training-label normalization receives None. Skip training-loss calculation when y_train is None.
  3. flaml/automl/ml.py:307: string anomaly labels are generically label-encoded before anomaly normalization, which can invert their meaning and the resulting metric. Normalize/validate anomaly labels before generic encoding, or explicitly reject nonnumeric labels.

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The merge from main is conflict-free but does not resolve the outstanding blockers:

  1. flaml/automl/automl.py:2516: single-class anomaly labels still permit meaningless HPO instead of rejecting evaluation data without both classes.
  2. flaml/automl/ml.py:658: X-only training with labeled validation and log_training_metric=True still attempts to normalize y_train=None and raises.
  3. flaml/automl/data.py:397: nonnumeric anomaly labels are still generically label-encoded before anomaly normalization, which can reverse normal/anomaly semantics.

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The merge from main is conflict-free but the outstanding blockers remain:

  1. flaml/automl/automl.py:2516-2520 and flaml/automl/ml.py:303-308: labeled anomaly HPO still accepts single-class labels, producing meaningless AP/ROC-AUC optimization.
  2. flaml/automl/ml.py:658-669: X-only training with labeled validation and log_training_metric=True still normalizes y_train=None and raises.
  3. flaml/automl/data.py:397-406: nonnumeric anomaly labels are still generically label-encoded before anomaly normalization, which can reverse normal/anomaly semantics.

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: changes are still required.

  1. flaml/automl/ml.py:303-308: labeled anomaly HPO accepts single-class {0} or {1} evaluation labels, producing meaningless AP/ROC-AUC optimization. Require the normalized evaluation-label set to be exactly {0, 1}.
  2. flaml/automl/ml.py:658-669: supported X-only training with labeled validation and log_training_metric=True still normalizes y_train=None and raises. Skip training-loss computation when training labels are absent.
  3. flaml/automl/data.py:397-406: generic LabelEncoder processing can reverse nonnumeric anomaly-label semantics before normalization. Use an explicit anomaly mapping first or reject nonnumeric labels.

Posted by thinkall-agent-auto-reviewer

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: changes are still required.

  1. flaml/automl/automl.py:2516-2520 and flaml/automl/ml.py:303-308: single-class normalized anomaly labels still permit meaningless AP/ROC-AUC HPO. Require the effective evaluation-label set to be exactly {0, 1}.
  2. flaml/automl/ml.py:658-669: X-only HPO with labeled validation and log_training_metric=True still calls anomaly-label normalization with y_train=None. Skip training-loss computation when labels are absent.
  3. flaml/automl/data.py:397-406 and flaml/automl/automl.py:892-893: generic label encoding reverses nonnumeric anomaly semantics and can make prediction fail on -1. Explicitly map anomaly labels before generic encoding, or reject nonnumeric labels.
  4. flaml/automl/automl.py:843-846 and flaml/automl/model.py:1574-1577: callable anomaly metrics can drive fitting but are ignored or rejected by public AutoML.score(). Either reject them during validation or support them consistently.

The current formatting workflow also fails because Black rewrites five changed files.

Posted by thinkall-agent-auto-reviewer

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: changes are still required.

  1. flaml/automl/ml.py:295-309: single-class anomaly evaluation labels still permit meaningless AP/ROC-AUC HPO. Require the normalized evaluation-label set to be exactly {0, 1}.
  2. flaml/automl/ml.py:658-669: X-only HPO with log_training_metric=True still normalizes y_train=None and raises. Skip training-loss calculation when training labels are absent.
  3. flaml/automl/data.py:397-406 and flaml/automl/automl.py:892-893: generic encoding reverses nonnumeric anomaly semantics and prediction can fail while inverse-transforming -1.
  4. flaml/automl/automl.py:843-846,2608-2613: callable anomaly metrics can drive fitting but are not honored by public AutoML.score().
  5. Current formatting CI fails because Black rewrites five changed files.

Posted by thinkall-agent-auto-reviewer

@rashidrao-pk

Copy link
Copy Markdown
Author

Thanks for the detailed review. I’ve pushed commit e5db672 addressing the remaining anomaly-detection issues:

  • Require both normal and anomaly classes for AP/ROC-AUC hyperparameter-search evaluation.
  • Skip training-metric computation when anomaly training is label-free (y_train=None).
  • Reject nonnumeric anomaly labels instead of passing them through generic label encoding.
  • Reject callable anomaly metrics for now, so AutoML.fit() and public AutoML.score() remain consistent.
  • Fixed Black/pre-commit formatting issues.
    I also added regression tests for single-class evaluation labels, X-only training with labeled validation and log_training_metric=True, nonnumeric training/validation labels, and callable metrics.
    Local checks:
  • anomaly tests: 14 passed
  • broader AutoML regression tests: 45 passed, 1 deselected
  • pre-commit: all hooks passed
    The GitHub Actions workflows are currently awaiting approval to run.
    Thanks again for the review.

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: changes are still required.

  1. flaml/automl/task/generic_task.py:1120,1147-1149 and flaml/automl/ml.py:638-644: labeled anomaly holdout/CV uses uniform/KFold splitting, so imbalanced but feasible datasets can produce single-class evaluation folds and fail. Stratify on normalized anomaly labels and validate minority count against n_splits.
  2. flaml/automl/automl.py:2683: label-free mode unconditionally changes max_iter=0 to 1, violating the documented explicit no-fit/search-space-only contract. Preserve explicit zero and default only unspecified one-shot requests.
  3. flaml/automl/model.py:1500-1585: current formatting CI still fails because Black rewrites mixed line endings. Commit the Black/pre-commit output.

Posted by thinkall-agent-auto-reviewer

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

support anomaly detection

3 participants