Data Scientist | Applied Machine Learning | Statistical Modeling | Model Evaluation
I am a Data Scientist with a Ph.D. in Statistics and experience in applied machine learning, predictive modeling, experimentation, and business analytics. My work focuses on whether models and analytical systems are reliable enough to support real decisions—not only whether they achieve a strong headline metric.
I am particularly interested in model validation, probability calibration, threshold selection, stability analysis, subgroup performance, behavioral data, and translating statistical evidence into clear recommendations.
- Applied Machine Learning: XGBoost, Random Forest, Logistic Regression, classification, feature engineering, model validation, SHAP, PyTorch, and transformer-based NLP
- Statistical Evaluation: ROC-AUC, PR-AUC, probability calibration, threshold analysis, simulation-based testing, Kendall-tau stability, bootstrap uncertainty, and subgroup diagnostics
- Product and Business Analytics: A/B testing, KPI development, customer segmentation, conversion, retention, revenue analysis, SQL workflows, and interactive dashboards
- Programming and Tools: Python, SQL, R, SAS, C++, pandas, NumPy, SciPy, scikit-learn, PyTorch, Hugging Face Transformers, Git/GitHub, Tableau, Power BI, and AWS
| Project | Focus | Methods and Tools |
|---|---|---|
| Financial Market Sentiment Analysis | Reproducible NLP pipeline for extracting and comparing sentiment signals from financial text | FinBERT, VADER, Hugging Face Transformers, PyTorch, Python |
| Credit Default Risk Modeling | Leakage-controlled credit-risk modeling with calibrated probabilities and cost-sensitive decision thresholds | XGBoost, Random Forest, Logistic Regression, ROC-AUC, PR-AUC, calibration |
| Portfolio Optimization and Investment Risk Analysis | Multi-strategy portfolio construction with out-of-sample evaluation and explicit risk diagnostics | Modern Portfolio Theory, Black-Litterman, covariance shrinkage, walk-forward backtesting, VaR/CVaR |
| Customer Revenue and Retention KPI Dashboard | End-to-end business-intelligence workflow for revenue, customer, product, retention, and conversion analysis | SQL, Python, Plotly Dash, Power BI modeling, DAX |
| Diabetes Risk Prediction with Machine Learning and Deep Learning | Healthcare risk-prediction benchmark comparing traditional machine learning with a PyTorch CNN | XGBoost, Random Forest, Logistic Regression, 1D CNN, calibration, SHAP, Integrated Gradients |
| Missing Data Mechanisms and Imputation Reliability in Healthcare | Repeated-simulation study of how missingness mechanisms and imputation choices affect value recovery, prediction, calibration, coefficient bias, and sample retention | MCAR, MAR, MNAR, complete-case analysis, KNN, Bayesian iterative imputation, multiple imputation, statistical simulation |
I am continuing to develop production-oriented machine-learning and analytics projects that combine:
- rigorous statistical evaluation;
- reproducible data and modeling pipelines;
- business or domain-specific decision logic;
- clear documentation of assumptions, limitations, and responsible use.
Applied machine learning · Predictive modeling · Product analytics · Experimentation · Model reliability · Financial and healthcare analytics · Decision-support systems