Skip to content
View Rithikesh-M's full-sized avatar

Block or report Rithikesh-M

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Rithikesh-M/README.md

Rithikesh Muddana

Founding Engineer at Ethical AI. Previously ML Research Intern at Axiom and at Scale AI.

I work on applied machine learning, with a recent focus on security and power-systems problems: anomaly and fraud detection, text and image classification, and representation learning on small tabular datasets.

Security and anomaly detection

Project What it does Result
dga-domain-detection Detects malware-generated domain names from the string alone (lexical + char n-gram features, XGBoost / LR / RF) on 300K domains from a 2.9M-row dataset. Includes recall-at-fixed-FPR operating points and a hold-out-a-malware-family generalization test. 0.9957 ROC-AUC; 90.5% recall at 0.5% FPR
malicious-url-detection Flags malicious HTTP request URLs from the string alone: a character n-gram co-occurrence matrix factorized with TruncatedSVD into frozen embeddings, then a three-block 1D CNN in PyTorch on 60K web-firewall URLs. 92.6% test accuracy; 0.95 recall on malicious URLs
malware-detection-transformers Treats Linux system-call traces (ADFA-LD, benign plus six attack types) as text, embeds them with a frozen BERT encoder, compares verbatim against run-length-collapsed encodings, and classifies with logistic regression and an MLP. 95.6% benign vs. malicious accuracy; 81.3% on the 7-way task
network-intrusion-representation-learning PCA, kernel PCA and t-SNE projections of KDD Cup 1999 connections, then PyTorch autoencoders whose bottleneck features are clustered with K-means and Gaussian mixtures and scored against the true labels. K-means purity 80.5% on raw features, 99.0% on the autoencoder bottleneck
binary-isa-identification Classifies 49K short program binaries into 12 instruction set architectures from raw bytes alone, comparing byte-histogram, byte n-gram and hex-nibble n-gram TF-IDF features across four classifiers. 99% test accuracy with n-gram TF-IDF; byte histograms cap at 92%
credit-card-fraud-detection Feature analysis and decision-tree baselines on a 0.17%-positive dataset, measures what SMOTE and Borderline-SMOTE actually buy, then compares voting, bagging, random forest, boosting and stacking ensembles on the same split. 0.84 F1 on the fraud class with a random forest, up from 0.75 for a single tree

Text classification

Project What it does Result
twitter-spam-detection Character-trigram count and TF-IDF features with Naive Bayes, logistic regression, and linear SVM on 50K tweets from CRESCI-2017. 95.7% test accuracy
email-spam-detection Bag-of-words spam filter for email: punctuation and stop-word removal, a 37K-word sparse count matrix, and logistic regression on 5.7K labeled messages, with the learned weights inspected to show which words drive the decision. 99.1% test accuracy; 0.98 F1 on the spam class
fake-news-classification Bag-of-words (unigram and bigram) features on 7.8K news articles across seven classifiers, plus a distance-metric study showing that swapping KNN from Euclidean to cosine lifts test F1 from 0.71 to 0.90 on length-varying count vectors. 99.5% test accuracy with a random forest

Vision

Project What it does Result
captcha-recognition-cnn End-to-end PyTorch CNN that reads all four CAPTCHA characters without segmentation, using zone-wise convolutional classifiers and on-the-fly affine augmentation. 82.3% whole-CAPTCHA accuracy on augmented test images

Power systems and regression

Project What it does Result
power-quality-fault-detection Kernel PCA and a multi-head autoencoder as 3-D embeddings for an RBF SVM; K-Means vs. GMM clustering evaluated against the true fault labels. 99% test accuracy; K-Means recovers 4 of 5 classes at 98% purity or better
power-system-fault-classification Compares regression, multi-label, and multi-class framings of the same fault-classification task with a small MLP; binary fault detector as an extension. 98.8% fault / no-fault accuracy
power-plant-regression Predicts a combined-cycle plant's hourly output from ambient conditions on data with injected outliers: Cook's distance removal, linear vs. Ridge vs. Lasso before and after cleaning, coefficient reliability, and full regularization paths. Test R² rises from 0.64 to 0.93 after removing 120 influential points
walmart-sales-regression Predicts weekly store sales from store identity, holiday flag and economic indicators; compares Random Forest, KNN, Elastic Net and Decision Tree pipelines under 5-fold CV, then stacks two of them with a Ridge meta-learner. 0.941 test R² with a random forest
fuel-and-electricity-regression Two regression case studies, fuel economy to horsepower and daily weather to building electricity use, comparing linear against degree 2 to 4 polynomial fits with a train/test gap analysis of where added capacity helps and where it collapses. 0.91 test R² on horsepower; unregularized degree-4 collapses to -33 R² on the weather data
wine-dimensionality-reduction Holds logistic regression fixed and varies only the scaling and the 2-D PCA projection on the 13-feature UCI Wine dataset, with decision-region plots, then compares kernel PCA, LDA and UMAP embeddings of the same data. Min-max scaling before PCA reaches 98.1% test accuracy against 68.5% unscaled

Tools

Python · PyTorch · scikit-learn · XGBoost · transformers · pandas · NumPy · Hugging Face datasets · Jupyter

Popular repositories Loading

  1. Masala-CHAI Masala-CHAI Public

    Forked from jitendra-bhandari/Masala-CHAI

    LLM circuit analysis

    Jupyter Notebook

  2. twitter-spam-detection twitter-spam-detection Public

    Spam-bot tweet classification with character n-gram count and TF-IDF features (Naive Bayes, logistic regression, linear SVM)

    Jupyter Notebook

  3. power-system-fault-classification power-system-fault-classification Public

    Comparing regression, multi-label, and multi-class neural-network formulations for three-phase power-system fault classification

    Jupyter Notebook

  4. captcha-recognition-cnn captcha-recognition-cnn Public

    End-to-end CAPTCHA recognition with a PyTorch CNN and no character segmentation

    Jupyter Notebook

  5. credit-card-fraud-detection credit-card-fraud-detection Public

    Credit card fraud detection on a 0.17%-positive dataset: feature analysis, decision-tree baselines with SMOTE comparison, and voting/bagging/random-forest/boosting/stacking ensembles

    Jupyter Notebook

  6. dga-domain-detection dga-domain-detection Public

    Detecting malware-generated (DGA) domain names from lexical and character n-gram features with XGBoost, Random Forest, and Logistic Regression

    Jupyter Notebook