ICD-10 Clinical Coding RAG
A production-style system that maps free-text clinical notes to ICD-10-CM diagnosis codes. Hybrid sparse+dense retrieval over 72,750 codes, cross-encoder reranking, and a QLoRA-fine-tuned Llama-3.2-3B served on vLLM — fronted by a Redis semantic cache that answers repeat queries in under 50 ms.
01 Overview
Medical coders read a clinical note and assign the diagnosis codes a hospital bills against — slow, and inconsistent between coders. This project treats it as a retrieval-augmented generation problem: find the handful of plausible ICD-10 codes for a note, then let a fine-tuned LLM make the final, exact call.
The interesting engineering isn't the LLM — it's everything around it. Extracting the diagnosis block before encoding, fusing keyword and semantic search, reranking with a joint cross-encoder, and caching semantically so the system is fast and cheap enough to actually run. The whole demo runs for under $5 of cloud time.
02 Results
Measured on 1,000 held-out clinical notes (evaluation/results.json), scoring retrieval before and after reranking plus end-to-end code selection.
Retrieval carries this system: the true code reaches the top-50 candidate pool 77.8% of the time and survives reranking into the top-5 at 74.0%, so the reranker costs almost nothing in coverage. End-to-end exact-code selection is 43.2% against the 82% target this project set itself — the gap is the model choosing among five plausible codes, not the retrieval finding them, and that is the next piece of work.
03 Architecture
The request path at a glance — then each stage in detail below. Every step exists to fix a specific failure mode of the one before it.
Diagnosis extraction
Clinical notes are 1,000+ words of SOAP boilerplate; ICD descriptions are 5–10 word phrases. Pulling only the Assessment/Diagnosis block before encoding lifted recall@5 from 50% to 64% in the ablation that justified it (50-note sample).
Query encoding
The extracted text is embedded with MedCPT's Query Encoder into a 768-d vector — a bi-encoder trained on biomedical literature, so clinical language lands in the right neighbourhood.
Semantic cache check
A Redis Stack HNSW index checks for a near-duplicate past query (cosine ≥ 0.92). On a hit, the cached code returns in under 50 ms and the rest of the pipeline is skipped.
Hybrid retrieval
On a miss, Milvus (Zilliz Cloud) runs BM25 sparse + MedCPT dense search in one query (α=0.5) over all 72,750 codes → top-50 candidates. BM25 catches rare terms dense vectors miss.
Cross-encoder rerank
The MedCPT Cross-Encoder jointly scores each (note, candidate) pair — far sharper than a dot product — and narrows 50 candidates to the top 5.
LLM selection
Llama-3.2-3B, QLoRA-fine-tuned (r=16, NF4 4-bit) on synthetic clinical notes, picks the exact code from the 5 candidates — served via vLLM in Docker with the Alpaca prompt it was trained on.
Return & persist
The final code, description, and confidence are returned via FastAPI and written back to the Redis cache so the next similar note is instant.
04 Key design decisions
A few choices that made this reliable rather than just a demo, and the trade-offs behind each one.
Extract before you encode
The single biggest accuracy win came from parsing, not modelling: isolating the diagnosis block before the encoder moved recall@5 up 14 points with zero model changes (measured on a 50-note ablation sample, before the 1,000-note run above).
Hybrid over pure dense
BM25 catches rare clinical terms ("dysthymic", "nephrotic") that dense models blur; dense catches paraphrases. Weighted fusion (α=0.5) beats either alone.
Right encoder for each side
MedCPT is a bi-encoder system — Article Encoder for documents, Query Encoder for queries. Using one for both (a common RAG bug) measurably degrades retrieval.
vLLM in Docker, pinned
Pinning vLLM 0.7.3 via the official image sidesteps the CUDA / PyTorch / transformers version hell on the AWS Deep Learning AMI — reproducible serving, every time.
Self-hosted Redis Stack
Vector caching needs RediSearch, which plain ElastiCache lacks; the Enterprise tier is $90+/mo. A Redis Stack container on the same box is free and uses ~50 MB.
Match the training prompt at inference
The model was fine-tuned in Alpaca format, so inference uses the raw completions API with that exact template — not chat tokens the model never saw. Distribution stays aligned.
05 Stack
Want the technical detail?
The repo has the full provisioning walkthrough, evaluation harness, and cost breakdown.