Healthcare NLP Hybrid RAG LLM Fine-tuning AWS

ICD-10 Clinical Coding RAG

A production-style system that maps free-text clinical notes to ICD-10-CM diagnosis codes. Hybrid sparse+dense retrieval over 72,750 codes, cross-encoder reranking, and a QLoRA-fine-tuned Llama-3.2-3B served on vLLM — fronted by a Redis semantic cache that answers repeat queries in under 50 ms.

Role
Solo · end-to-end
Domain
Clinical coding
Serving
vLLM · FastAPI
Infra
AWS EC2 · Zilliz

01 Overview

Medical coders read a clinical note and assign the diagnosis codes a hospital bills against — slow, and inconsistent between coders. This project treats it as a retrieval-augmented generation problem: find the handful of plausible ICD-10 codes for a note, then let a fine-tuned LLM make the final, exact call.

The interesting engineering isn't the LLM — it's everything around it. Extracting the diagnosis block before encoding, fusing keyword and semantic search, reranking with a joint cross-encoder, and caching semantically so the system is fast and cheap enough to actually run. The whole demo runs for under $5 of cloud time.

02 Results

Measured on 1,000 held-out clinical notes (evaluation/results.json), scoring retrieval before and after reranking plus end-to-end code selection.

Retrieval carries this system: the true code reaches the top-50 candidate pool 77.8% of the time and survives reranking into the top-5 at 74.0%, so the reranker costs almost nothing in coverage. End-to-end exact-code selection is 43.2% against the 82% target this project set itself — the gap is the model choosing among five plausible codes, not the retrieval finding them, and that is the next piece of work.

72,750
ICD-10-CM codes indexed with both dense & sparse vectors
74.0%
recall@5 after reranking — the true code in the five codes shown
<50ms
latency on a semantic cache hit (vs ~2.4s cold)
<$5
total EC2 cost for a full build + tune + eval demo

03 Architecture

The request path at a glance — then each stage in detail below. Every step exists to fix a specific failure mode of the one before it.

Input
Clinical note
SOAP text
Preprocess
Diagnosis extractor
regex
Encode
MedCPT encoder
768-d vector
Cache
Redis semantic cache
hit → <50ms
Retrieve
Milvus hybrid
BM25 + dense
Rerank
Cross-encoder
top-50 → 5
Generate
Llama-3.2-3B
QLoRA · vLLM
Output
ICD-10 + confidence
Backing services
Zilliz Cloud — 72,750 codesRedis Stack (HNSW)AWS EC2 g4dn (T4)FastAPI
1

Diagnosis extraction

Clinical notes are 1,000+ words of SOAP boilerplate; ICD descriptions are 5–10 word phrases. Pulling only the Assessment/Diagnosis block before encoding lifted recall@5 from 50% to 64% in the ablation that justified it (50-note sample).

regexSOAP parsing
2

Query encoding

The extracted text is embedded with MedCPT's Query Encoder into a 768-d vector — a bi-encoder trained on biomedical literature, so clinical language lands in the right neighbourhood.

MedCPT Query Encoder768-d
3

Semantic cache check

A Redis Stack HNSW index checks for a near-duplicate past query (cosine ≥ 0.92). On a hit, the cached code returns in under 50 ms and the rest of the pipeline is skipped.

Redis Stack HNSW
4

Hybrid retrieval

On a miss, Milvus (Zilliz Cloud) runs BM25 sparse + MedCPT dense search in one query (α=0.5) over all 72,750 codes → top-50 candidates. BM25 catches rare terms dense vectors miss.

Milvus / ZillizBM25 + dense
5

Cross-encoder rerank

The MedCPT Cross-Encoder jointly scores each (note, candidate) pair — far sharper than a dot product — and narrows 50 candidates to the top 5.

MedCPT Cross-Encoder
6

LLM selection

Llama-3.2-3B, QLoRA-fine-tuned (r=16, NF4 4-bit) on synthetic clinical notes, picks the exact code from the 5 candidates — served via vLLM in Docker with the Alpaca prompt it was trained on.

Llama-3.2-3B · QLoRAvLLM 0.7.3
7

Return & persist

The final code, description, and confidence are returned via FastAPI and written back to the Redis cache so the next similar note is instant.

FastAPIcache write-back

04 Key design decisions

A few choices that made this reliable rather than just a demo, and the trade-offs behind each one.

Extract before you encode

The single biggest accuracy win came from parsing, not modelling: isolating the diagnosis block before the encoder moved recall@5 up 14 points with zero model changes (measured on a 50-note ablation sample, before the 1,000-note run above).

Hybrid over pure dense

BM25 catches rare clinical terms ("dysthymic", "nephrotic") that dense models blur; dense catches paraphrases. Weighted fusion (α=0.5) beats either alone.

Right encoder for each side

MedCPT is a bi-encoder system — Article Encoder for documents, Query Encoder for queries. Using one for both (a common RAG bug) measurably degrades retrieval.

vLLM in Docker, pinned

Pinning vLLM 0.7.3 via the official image sidesteps the CUDA / PyTorch / transformers version hell on the AWS Deep Learning AMI — reproducible serving, every time.

Self-hosted Redis Stack

Vector caching needs RediSearch, which plain ElastiCache lacks; the Enterprise tier is $90+/mo. A Redis Stack container on the same box is free and uses ~50 MB.

Match the training prompt at inference

The model was fine-tuned in Alpaca format, so inference uses the raw completions API with that exact template — not chat tokens the model never saw. Distribution stays aligned.

05 Stack

Retrieval
Milvus / Zilliz CloudBM25 (sparse)MedCPT (dense)Cross-encoder rerank
Models
Llama-3.2-3BQLoRA (NF4)UnslothMedCPT encoders
Serving
vLLM 0.7.3FastAPIPydanticDocker
Cache
Redis StackHNSW vector index
Infra
AWS EC2 g4dn.xlarge (T4)Zilliz serverless

Want the technical detail?

The repo has the full provisioning walkthrough, evaluation harness, and cost breakdown.