Skip to content
View ejunior029's full-sized avatar
  • São Paulo - Brazil

Block or report ejunior029

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ejunior029/README.md

Hi, I'm Edson Júnior

Senior Data Scientist building AI Agents & LLM systems that run in production — at Braskem, Bradesco, B3 (Brazil's Stock Exchange) and beyond

6+ years turning messy business problems — fraud, AML/CFT, churn, credit risk — into ML systems that ship. These days that means AI Agents and LLM applications on Azure AI Foundry; before that it meant fraud models running on Databricks catching real money in real time.


Currently

  • Building multi-agent LLM systems and intelligent document processing at Braskem
  • Exploring: agent evaluation/observability, RAG architectures for regulated industries
  • Writing about ML fundamentals (metrics, cross-validation, bias-variance) on Medium

Pinned work — start here

Repo What it shows
brazil-municipal-gdp-regression Regression on real IBGE data (5,570 municipalities): EDA → baseline → 5-fold CV model comparison → Optuna Bayesian tuning → SHAP explainability. XGBoost hits R² 0.90 in CV, but SHAP traces the CV-vs-test RMSE gap straight back to a single mining-town outlier the model never learned to extrapolate to. Shipped as a Dockerized FastAPI service with Supabase-logged prediction history. Try it live →
fatal-accident-prediction-br Imbalanced binary classification (13:1) on 73k real PRF traffic-accident records: EDA → baseline → 5-model CV comparison (LogReg, RF, XGBoost, CatBoost, LightGBM) → Optuna tuning + threshold sweep → SHAP explainability. Exposes the "93% accuracy, 0 fatalities detected" trap, then fixes it — F1 climbs from 0.00 (dummy) to 0.39, ROC-AUC 0.835. Shipped as a Dockerized FastAPI prediction service.
municipios-br-clustering Unsupervised clustering on real IBGE data (5,570 municipalities): EDA → baseline KMeans → 7-algorithm comparison (KMeans, Agglomerative, DBSCAN, HDBSCAN, OPTICS, KModes, KPrototypes) → final model shipped as a public API. Rediscovers Brazil's Southeast/Northeast economic divide with zero labels — silhouette 0.41. Try it live →
anp-fuel-price-anomaly-detection Unsupervised anomaly detection on 52k real ANP fuel-price records across all 27 Brazilian states: EDA → baseline → 4-algorithm comparison → synthetic-anomaly evaluation (no real labels exist, so anomalies are injected to measure it). Exposes LocalOutlierFactor flagging zero anomalies despite a 0.91 ROC-AUC; winner is OneClassSVM, F1 0.68. Shipped as a live map, auto-updated monthly by GitHub Actions. Try it live →

Experience

Senior Data Scientist — Braskem (current)
AI Agents, multi-agent systems, LLMs, prompt engineering, and intelligent document processing on Azure AI Foundry — moving enterprise workflows from manual to automated.

Data Scientist — Banco Bradesco
Built a LightGBM fraud-detection model on large-scale transaction data: benchmarked candidates with Databricks AutoML, handled severe class imbalance through sampling, and validated with time-based cross-validation to avoid look-ahead bias. Tuned with Optuna, tracked with MLflow, and orchestrated end-to-end through Databricks Jobs. Shipped to production at an 80% fraud detection rate.

CRM Data Scientist — Banco Sofisa
Lead propensity models, customer segmentation, and recommendation systems feeding credit and CRM decisions. Within 3 months of deployment, the propensity models drove R$6M+ in new credit risk originated for the bank

Data Scientist — B3 (Brazil's Stock Exchange)
Built econometric revenue-forecasting models (frequentist and Bayesian) for macroeconomic risk monitoring, and unsupervised fraud detection (KMeans, DBSCAN) to profile investors behind fraudulent trading activity — plus risk-control tooling that automated the monthly monitoring of funds trading assets they weren't authorized to hold. DBSCAN isolated fraud cases into their own distinct cluster — turning a manual investigation into a repeatable detection signal.


Stack

Languages Python · SQL · R ML scikit-learn · XGBoost · LightGBM · CatBoost · TensorFlow · PyTorch GenAI / Agents GPT · Claude · Azure AI Foundry · RAG · Multi-Agent Systems · Prompt Engineering Data & MLOps Databricks · PySpark · Delta Lake · MLflow · Feature Engineering Cloud Microsoft Azure Other Git · Docker · Power BI


Teaching

Instructor at FCCD and Universidade dos Dados — courses on Machine Learning, Statistics, and Python. Explaining a concept clearly to a room of students is a good forcing function for actually understanding it.


GitHub stats

GitHub Stats GitHub Stats


Background

Electrical Engineering (UNESP) · Postgraduate in AI & Big Data (USP)

Languages

🇧🇷 Portuguese — Native · 🇺🇸 English — Professional working proficiency · 🇫🇷 French — Intermediate


📫 ejunior029@gmail.com — open to conversations about AI Agents, LLM applications, and ML in production.

Pinned Loading

  1. brazil-municipal-gdp-regression brazil-municipal-gdp-regression Public

    Regression project predicting GDP per capita of Brazilian municipalities from IBGE open data — includes EDA, baseline models, cross-validation model comparison, and Bayesian hyperparameter tuning w…

    Jupyter Notebook 1

  2. fatal-accident-prediction-br fatal-accident-prediction-br Public

    Predicting fatal traffic accidents from Brazil's Federal Highway Police (PRF) 2024 data — an imbalanced binary classification project (EDA → baseline → model comparison → Optuna tuning)

    Jupyter Notebook 1

  3. municipios-br-clustering municipios-br-clustering Public

    Unsupervised clustering of Brazilian municipalities using IBGE open data — KMeans, KModes, KPrototypes, DBSCAN, HDBSCAN and OPTICS compared end-to-end, from EDA to final model evaluation.

    Jupyter Notebook 1

  4. anp-fuel-price-anomaly-detection anp-fuel-price-anomaly-detection Public

    Unsupervised anomaly detection on Brazilian fuel price data (ANP), comparing Isolation Forest, LOF, and One-Class SVM.

    Jupyter Notebook 1

  5. supplyChain_fraud_prediction supplyChain_fraud_prediction Public

    Jupyter Notebook 1

  6. Census_Income Census_Income Public

    Jupyter Notebook 1