Data scientist

I build models that have to earn their launch.

I work the whole path a model takes to a decision: framing the question, building the feature and the baseline it has to beat, designing the experiment that judges it, and then pricing the outcome in money rather than error.

A lower metric isn't a result. So my projects carry pre-registered decision rules, guardrails, confidence intervals, and a written verdict — including the two iterations where the more accurate model lost money and did not ship.

Open to Data Scientist roles 4 end-to-end projects, all source public Shipped on AWS · GCP · Azure
Selected work

Four projects, each carried to a decision.

Forecasting and experimentation, drift monitoring, streaming detection, and the pipelines underneath them. Every one runs on real data, states what was measured and what wasn't, and links to its full source. Every card flips — the front is the claim, the back is the evidence.

01 Forecasting · experimentation

Demand Forecasting, Cost-Aware A/B Tested

Does the better forecast actually save money? Tested — and twice refused.

01

A newsvendor cost simulation and a pre-registered A/B test over the full M5 panel. Two iterations returned HOLD; ordering the cost-optimal quantile instead of the mean is what shipped.

inventory cost, CI [−15.3, −12.9]
−14.1%
paired series in the test
30,474
02 MLOps · drift · feature store

Customer Segmentation, Monitored

Segmentation treated as a model with a lifecycle, not a one-off notebook.

02

Windowed RFM features, a six-algorithm bakeoff, and drift thresholds calibrated on the data's own variation, with MLflow champion/challenger retraining behind a Feast store.

drift alerts after calibration
21/22 → 4
customer-window rows
47,170
03 Data engineering · Spark · Snowflake

Airline On-Time Performance Pipeline

A decade of US flight data, cleaned, warehoused, and kept current on its own.

03

Spark cleans a decade of BTS flights into partitioned Parquet; dbt models it in Snowflake behind 36 tests and a Great Expectations gate. A monthly Airflow DAG ingested a new month unattended.

rows across 137 files
72.9M
smaller on disk after cleaning
94%
04 Streaming MLOps · Azure

Real-Time Crypto Anomaly Detection

Live anomaly scoring for three assets on a streaming feature store.

04

Coinbase WebSocket to Event Hubs, a stateful feature engine, a Redis feature store, and an Isolation-Forest plus autoencoder ensemble. One feature definition serves training and serving.

feature-store reads
<5 ms
retrain, register, redeploy
Weekly
Toolkit

What I reach for.

Tools I've used to ship something on this page — not a list of everything I've read about.

ML & statistics
LightGBM, PyTorch, scikit-learn, time-series forecasting, anomaly detection, clustering, quantile / pinball loss, rolling-origin backtesting, bootstrap intervals, calibration
Experimentation
Experiment design, pre-registered decision rules and guardrails, paired and A/B tests, power analysis, causal inference, cost-aware evaluation, honest negative results
Data engineering
PySpark, dbt, Snowflake, Airflow, Dagster, Great Expectations, Pandera contracts, DuckDB, Parquet medallion layouts, DVC
MLOps & cloud
MLflow model registry, Feast feature store, drift detection and retraining gates, Docker, Terraform, GitHub Actions, AWS · GCP · Azure, monitoring and alerting
Analytics & BI
Power BI (including REST push refresh), Tableau, Marimo, Plotly, matplotlib
Languages & APIs
Python, SQL, FastAPI, Pydantic, Streamlit, REST and WebSocket, server-sent events
About

Careful analysis, engineering that follows through.

A model that does well in validation is only the start. The questions I find more interesting come after it: does this hold on data it hasn't seen, does it change a decision anyone is actually making, and is it worth what it costs to run?

That shows up throughout this work. A forecasting model is judged by a newsvendor cost simulation and a pre-registered A/B test rather than by WMAPE — and the gate said HOLD twice before it said SHIP. A segmentation model is treated as something with a lifecycle, with drift thresholds calibrated against the data's own month-to-month variation instead of textbook constants, which cut false retrain alerts from 21 windows out of 22 down to the four real regime changes.

I also report what didn't work. Two full iterations went into a model that was more accurate and still lost money, and both are written up next to the version that finally earned its launch. A result you'd bury when it disagrees with you isn't a measurement.

I'm comfortable across the whole workflow — framing the problem, moving the data with Spark, dbt and Airflow, building and tuning the model, wiring the serving layer, and setting up the monitoring that catches drift — on AWS, GCP and Azure.

Contact

If this looks relevant to what your team is building, I'd like to hear from you.

I'm looking for Data Scientist roles. The fastest way to reach me is email.