AI engineer

I build LLM systems that survive contact with production.

I work the whole path an AI feature takes to get real: retrieval that finds the right evidence, a model tuned and served inside a latency and cost budget, guardrails that keep it from asserting things it can't support, and an eval harness that decides whether any of it shipped.

The demos are easy. What I care about is whether the thing holds — so my projects carry baselines they could lose to, confidence intervals, cost per request, and a written verdict when the answer was no.

Open to AI Engineer roles 6 end-to-end systems, all source public Shipped on AWS · GCP · Azure
Selected work

Six systems, each carried to a decision.

Retrieval, agents, fine-tuning, serving, and the measurement around them. Every one runs on real data, states what was measured and what wasn't, and links to its full source. Every card flips — the front is the claim, the back is the evidence.

01 Healthcare · RAG · LLM serving

ICD-10 Clinical Coding RAG

Clinical notes to ICD-10-CM codes, retrieved and reranked at production scale.

01

Hybrid BM25 + MedCPT retrieval over 72,750 codes, cross-encoder reranking, and a QLoRA-tuned Llama-3.2 served on vLLM behind a semantic cache. Scored on 1,000 held-out notes.

recall@5 over 1,000 notes
74%
cache-hit latency
<50 ms
02 Graph RAG · evaluation

GraphScout

Three retrieval arms over one corpus, measured with confidence intervals.

02

Baseline, graph-RAG and agentic arms scored on one corpus with bootstrap CIs. Agentic retrieval nearly doubles recall — and answer F1 barely moves, so retrieval was never the constraint.

recall@8, dense → agentic
0.454 → 0.838
all-support@8
14% → 65%
03 Multi-agent · guardrails

Atlas — Agentic Trip Planner

Six agents plan a real trip without ever being trusted with a fact.

03

Six agents with narrow tool sets each. Every price is checked against the payload it came from, all arithmetic is Python, and uncertainty is labelled — on free-tier APIs throughout.

tests + a 12-case eval set
608
cold run vs resumed
87.2s → 1.4s
04 Agentic AI · verified execution

AgentForge

A coding agent that has to prove its patch works before it can finish.

04

Writes a failing acceptance test first, patches, then verifies through deterministic gates: digest-checked tests, static checks before any model review, bounded repair, enforced budgets.

automated tests passing
206
05 Data product · QLoRA · GCP

Job Market Intelligence

Job postings turned into a queryable company and tech-stack signal.

05

Scraping, rule-based labeling, a QLoRA-tuned extractor served on vLLM, a DuckDB warehouse, and a Marimo dashboard on Cloud Run.

instruction examples
10,000
06 NLP · LangChain

AI Study Notes Generator

PDF chapters into summaries, flashcards, a glossary, and topic clusters.

06

gpt-4o-mini paired with classical NLP — KeyBERT and spaCy for terms, UMAP + HDBSCAN for clusters — and graded with ROUGE against ground truth.

Toolkit

What I reach for.

Tools I've used to ship something on this page — not a list of everything I've read about.

LLM & GenAI
Hybrid and graph RAG, QLoRA / PEFT fine-tuning, vLLM serving, LangGraph, MCP, multi-agent orchestration, cross-encoder reranking, semantic caching, structured output
Agents & orchestration
Tool design and least-privilege tool sets, bounded repair loops, provenance and claim verification, deterministic guardrails, budget and rate-limit control, checkpointed resume
Evaluation
Versioned eval harnesses, golden sets, LLM-as-judge calibration, bootstrap confidence intervals, regression gates in CI, cost and latency per request, honest negative results
Retrieval & storage
pgvector, Milvus / Zilliz, Qdrant, FAISS, Neo4j, BM25 + RRF fusion, PostgreSQL, Redis, Kafka / Event Hubs
Serving & MLOps
FastAPI, vLLM, Docker, AWS EC2, GCP Cloud Run, Azure ML and ACI, Terraform, MLflow, GitHub Actions, Langfuse, monitoring and alerting, SkyPilot
Languages & APIs
Python, SQL, FastAPI, Pydantic, Streamlit, REST and WebSocket, server-sent events
About

Systems that hold up, not demos that impress.

An LLM feature that works in a notebook is the easy 20%. The questions I find more interesting come after it: does the retrieval surface the right evidence on inputs nobody curated, can the system tell you when it doesn't know, what does a request cost, and how would you notice if it quietly got worse?

That shows up throughout this work. A semantic cache and a cross-encoder make a clinical retrieval system fast and accurate enough to trust. A multi-agent planner checks every price against the tool payload it came from and does its arithmetic in Python, because a model that mis-adds a budget is worse than no budget at all. A coding agent has to produce an observed failing test and a passing containerized run before it's allowed to call itself done.

I also report what didn't work. One project retracts its own earlier retrieval finding once the ablation was run properly — the corrected result is indistinguishable from zero, and it sits in the README next to the wins. Another measures a near-doubling of retrieval recall and says plainly, in the same table, that answer F1 barely moved. A result you'd bury when it disagrees with you isn't a measurement.

I'm comfortable across the whole path — retrieval and indexing, fine-tuning, the serving layer, the eval harness, and the tracing and cost accounting that keep it honest — on AWS, GCP and Azure.

Contact

If this looks relevant to what your team is building, I'd like to hear from you.

I'm looking for AI Engineer roles. The fastest way to reach me is email.