Founding Engineer at Ethical AI. Previously ML Research Intern at Axiom and at Scale AI.
I work on applied machine learning, with a recent focus on security and power-systems problems: anomaly and fraud detection, text and image classification, and representation learning on small tabular datasets.
| Project | What it does | Result |
|---|---|---|
| dga-domain-detection | Detects malware-generated domain names from the string alone (lexical + char n-gram features, XGBoost / LR / RF) on 300K domains from a 2.9M-row dataset. Includes recall-at-fixed-FPR operating points and a hold-out-a-malware-family generalization test. | 0.9957 ROC-AUC; 90.5% recall at 0.5% FPR |
| malicious-url-detection | Flags malicious HTTP request URLs from the string alone: a character n-gram co-occurrence matrix factorized with TruncatedSVD into frozen embeddings, then a three-block 1D CNN in PyTorch on 60K web-firewall URLs. | 92.6% test accuracy; 0.95 recall on malicious URLs |
| malware-detection-transformers | Treats Linux system-call traces (ADFA-LD, benign plus six attack types) as text, embeds them with a frozen BERT encoder, compares verbatim against run-length-collapsed encodings, and classifies with logistic regression and an MLP. | 95.6% benign vs. malicious accuracy; 81.3% on the 7-way task |
| network-intrusion-representation-learning | PCA, kernel PCA and t-SNE projections of KDD Cup 1999 connections, then PyTorch autoencoders whose bottleneck features are clustered with K-means and Gaussian mixtures and scored against the true labels. | K-means purity 80.5% on raw features, 99.0% on the autoencoder bottleneck |
| binary-isa-identification | Classifies 49K short program binaries into 12 instruction set architectures from raw bytes alone, comparing byte-histogram, byte n-gram and hex-nibble n-gram TF-IDF features across four classifiers. | 99% test accuracy with n-gram TF-IDF; byte histograms cap at 92% |
| credit-card-fraud-detection | Feature analysis and decision-tree baselines on a 0.17%-positive dataset, measures what SMOTE and Borderline-SMOTE actually buy, then compares voting, bagging, random forest, boosting and stacking ensembles on the same split. | 0.84 F1 on the fraud class with a random forest, up from 0.75 for a single tree |
| Project | What it does | Result |
|---|---|---|
| twitter-spam-detection | Character-trigram count and TF-IDF features with Naive Bayes, logistic regression, and linear SVM on 50K tweets from CRESCI-2017. | 95.7% test accuracy |
| email-spam-detection | Bag-of-words spam filter for email: punctuation and stop-word removal, a 37K-word sparse count matrix, and logistic regression on 5.7K labeled messages, with the learned weights inspected to show which words drive the decision. | 99.1% test accuracy; 0.98 F1 on the spam class |
| fake-news-classification | Bag-of-words (unigram and bigram) features on 7.8K news articles across seven classifiers, plus a distance-metric study showing that swapping KNN from Euclidean to cosine lifts test F1 from 0.71 to 0.90 on length-varying count vectors. | 99.5% test accuracy with a random forest |
| Project | What it does | Result |
|---|---|---|
| captcha-recognition-cnn | End-to-end PyTorch CNN that reads all four CAPTCHA characters without segmentation, using zone-wise convolutional classifiers and on-the-fly affine augmentation. | 82.3% whole-CAPTCHA accuracy on augmented test images |
| Project | What it does | Result |
|---|---|---|
| power-quality-fault-detection | Kernel PCA and a multi-head autoencoder as 3-D embeddings for an RBF SVM; K-Means vs. GMM clustering evaluated against the true fault labels. | 99% test accuracy; K-Means recovers 4 of 5 classes at 98% purity or better |
| power-system-fault-classification | Compares regression, multi-label, and multi-class framings of the same fault-classification task with a small MLP; binary fault detector as an extension. | 98.8% fault / no-fault accuracy |
| power-plant-regression | Predicts a combined-cycle plant's hourly output from ambient conditions on data with injected outliers: Cook's distance removal, linear vs. Ridge vs. Lasso before and after cleaning, coefficient reliability, and full regularization paths. | Test R² rises from 0.64 to 0.93 after removing 120 influential points |
| walmart-sales-regression | Predicts weekly store sales from store identity, holiday flag and economic indicators; compares Random Forest, KNN, Elastic Net and Decision Tree pipelines under 5-fold CV, then stacks two of them with a Ridge meta-learner. | 0.941 test R² with a random forest |
| fuel-and-electricity-regression | Two regression case studies, fuel economy to horsepower and daily weather to building electricity use, comparing linear against degree 2 to 4 polynomial fits with a train/test gap analysis of where added capacity helps and where it collapses. | 0.91 test R² on horsepower; unregularized degree-4 collapses to -33 R² on the weather data |
| wine-dimensionality-reduction | Holds logistic regression fixed and varies only the scaling and the 2-D PCA projection on the 13-feature UCI Wine dataset, with decision-region plots, then compares kernel PCA, LDA and UMAP embeddings of the same data. | Min-max scaling before PCA reaches 98.1% test accuracy against 68.5% unscaled |
Python · PyTorch · scikit-learn · XGBoost · transformers · pandas · NumPy · Hugging Face datasets · Jupyter
