Job Market Intelligence Pipeline
An end-to-end ML pipeline that turns raw job postings into structured company & tech-stack signals — mirroring the go-to-market knowledge-graph products built by B2B data-intelligence companies. Scraping, a fine-tuned LLM extractor, a 4-stage entity-normalization pipeline, a DuckDB warehouse, and a live dashboard on GCP Cloud Run.
01 Overview
B2B go-to-market teams want to know which companies use which technologies, at what seniority they're hiring, and how their stack is changing. This pipeline builds exactly that signal from public job postings — the same shape of data product a company like Sumble sells, built end-to-end.
The hard part isn't the model, it's the messiness of real web data: "Google LLC" vs "Google", "React.js" vs "ReactJS", near-duplicate postings across boards. So the centre of gravity is a 4-stage entity-normalization pipeline and a quality-scoring layer that makes the output trustworthy enough to query — plus a fine-tuned LLM extractor that degrades gracefully to rules if it's ever down.
02 Results
03 Architecture
Raw postings in, a queryable knowledge graph out — with a fine-tuned model trained on labels the pipeline generates itself. The data flow at a glance, then each step below.
Scrape
A Selenium-based ATS scraper plus public-API and SERP sources pull job postings from company career pages — checkpointed so an interrupted run resumes instead of restarting.
Rule-based labeling
spaCy + regex (patterns compiled once at load) extract seniority, tech stack, and remote type as zero-cost "silver labels" — the training data for the LLM, bootstrapped without any manual annotation.
QLoRA fine-tune Gemma-2B
Silver labels + a HuggingFace dataset are formatted into 10,000 Alpaca instruction pairs and used to QLoRA-fine-tune Gemma-2B-it (4-bit NF4) — on a free Colab T4 or a GCP spot GPU via SkyPilot.
Enrich & normalize
A vLLM client (exponential backoff, batches of 10 for KV-cache sharing, rule-based fallback) extracts structured fields, then the 4-stage normalizer resolves company names (fuzzy), canonicalizes tech, de-dupes with FAISS, and scores quality.
Warehouse & knowledge graph
Enriched records land in a DuckDB analytical warehouse (enriched_jobs, tech_signals, company_signals) and a Postgres knowledge graph modelling org structure, tech stack, and hiring signals.
Serve
A reactive Marimo dashboard is containerized and deployed to GCP Cloud Run; the fine-tuned model is exposed behind an OpenAI-compatible vLLM inference API.
04 Engineering decisions
4-stage normalization
Strip legal suffixes → exact match → rapidfuzz token-sort for word-order variants → tech alias resolution → sentence-transformer + FAISS dedup. Messy web data becomes joinable entities.
Explainable quality scores
An additive 100-point score stores a per-dimension breakdown, not just a total — so a consumer asking "why is this record low quality?" gets a root-cause answer.
Knowledge graph as the product
The Postgres graph — org structure, tech stack, key projects per company — is the actual deliverable, modelled directly on how GTM-intelligence products are consumed.
Graceful degradation
The LLM isn't a hard dependency — if the vLLM endpoint is down, enrichment falls back to the rule-based labeler and the pipeline keeps producing data.
DuckDB over BigQuery
In-process, reads Parquet/JSONL directly, and can ATTACH to Postgres as a foreign-data wrapper — one query layer over operational and analytical data, zero cloud credentials to develop.
Spot-GPU training
SkyPilot launches QLoRA fine-tuning on GCP spot instances — the same job runs free on Colab or cheap on preemptible cloud GPUs.
05 Stack
Try the live dashboard
The deployed Marimo app is public — explore the tech and company signals, then read how it's built.