Code for classification with HyperDFS. A subset encoder represents the observed features, a hypernetwork generates the weights of a small predictor, and a learned policy chooses features one at a time.
This folder contains the tabular and image patch HyperDFS implementations and small demos on the paper's Proxy Substitution synthetic benchmark and MNIST. (The image demo downloads MNIST or FashionMNIST automatically.)
This repository focuses on the HyperDFS implementation, but if you need any other piece of code regarding baselines, for example, contact me at j.fumanal-idocin at essex dot ac dot uk.
From this folder:
python -m pip install -r requirements.txt
python demo.pyThe demo generates 1,200 Proxy Substitution samples locally, trains on CPU,
and prints accuracy with the most precise proxy alone and with acquisition
budgets 0 through 3. The paper uses 10,000 samples and cross-validation;
this smaller run illustrates the API and does not reproduce its reported scores.
paper_datasets.py contains the generator, so no download is needed. Features
0–4 are noisy views of one latent signal; features 5–9 are independent noise.
From this folder, after installing requirements.txt:
python -m image_experiments.run_imageThis downloads MNIST to ~/.cache/hyperdfs/image_datasets, extracts a 7 × 7
grid of 4 × 4 patches per image, and trains the original two-phase image
HyperDFS model on a small stratified sample. It prints accuracy and macro-F1
for patch acquisition budgets 0 through 3. The defaults are a CPU smoke run,
not paper settings or paper scores.
To run the HyperDFS arm of the image benchmark with five stratified folds over the combined MNIST splits and the image experiment's training settings:
python -m image_experiments.run_image --train-size 0 --test-size 0 \
--folds 5 --encoder isab --patch-embed-dim 16 --primary-hidden 64 \
--n-masks-per-step 3 --batch-size 128 --epochs 200 \
--policy-epochs 100 --min-budget 2 --budget 10 --device cuda \
--output results/image/mnist.jsonUse --dataset fashionmnist for FashionMNIST. --data-root changes the
download cache. The classifier in
image_experiments/models/hyperdfs_image.py accepts arrays shaped
(samples, patches, pixels_per_patch) and exposes fit, predict_with_mask,
predict, and budget_curve for other image datasets or evaluation protocols.
With --folds 0 (the default), the runner uses the official train/test split.
The cross-validation mode materialises image patches in memory. Comparison
methods from the larger research repository are not included here.
In general, this is a copy and paste-friendly code:
import numpy as np
from sklearn.model_selection import train_test_split
from hyperdfs import HyperDFSClassifier
from paper_datasets import make_proxy_substitution
X, y = make_proxy_substitution(n_samples=1200, seed=42)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=42,
)
model = HyperDFSClassifier(
encoder_type='deepsets',
policy='selection',
device='cpu',
verbose=False,
)
model.fit(X_train, y_train)
# Feature order must match the training matrix. 1 means observed.
mask = np.zeros(X.shape[1], dtype=np.float32)
mask[0] = 1.0
probabilities = model.predict_with_mask(X_test, mask)
# Select up to two features for each sample with the learned policy.
predictions = model.predict(X_test, budget=2)
budgets, scores = model.budget_curve(X_test, y_test, max_budget=2,
metric='accuracy')We got accepted at NeurIPS 2026! Until the final version is ready, cite the preprint if you find this work useful:
@article{fumanal2026hypernetworks,
title={Hypernetworks for Dynamic Feature Selection},
author={Fumanal-Idocin, Javier and Fernandez-Peralta, Raquel and Andreu-Perez, Javier},
journal={arXiv preprint arXiv:2605.12278},
year={2026}
}MIT License. Copyright (c) 2026 Javier Fumanal Idocin. See LICENSE.