Data Product LLM Fine-tuning Data Engineering GCP

Job Market Intelligence Pipeline

An end-to-end ML pipeline that turns raw job postings into structured company & tech-stack signals — mirroring the go-to-market knowledge-graph products built by B2B data-intelligence companies. Scraping, a fine-tuned LLM extractor, a 4-stage entity-normalization pipeline, a DuckDB warehouse, and a live dashboard on GCP Cloud Run.

Role
Solo · end-to-end
Core product
GTM knowledge graph
Warehouse
DuckDB + Postgres
Deploy
GCP Cloud Run

01 Overview

B2B go-to-market teams want to know which companies use which technologies, at what seniority they're hiring, and how their stack is changing. This pipeline builds exactly that signal from public job postings — the same shape of data product a company like Sumble sells, built end-to-end.

The hard part isn't the model, it's the messiness of real web data: "Google LLC" vs "Google", "React.js" vs "ReactJS", near-duplicate postings across boards. So the centre of gravity is a 4-stage entity-normalization pipeline and a quality-scoring layer that makes the output trustworthy enough to query — plus a fine-tuned LLM extractor that degrades gracefully to rules if it's ever down.

02 Results

10,000
instruction-tuning examples generated for fine-tuning
4-stage
entity normalization: company → tech → dedup → quality
100-pt
quality score with a stored per-dimension breakdown
Live
interactive dashboard deployed on GCP Cloud Run

03 Architecture

Raw postings in, a queryable knowledge graph out — with a fine-tuned model trained on labels the pipeline generates itself. The data flow at a glance, then each step below.

Sources
ATS · API · SERP
Scrape
Selenium
checkpointed
Label
spaCy + regex
silver labels
Fine-tune
QLoRA Gemma-2B
10k examples
Enrich
vLLM extractor
rule fallback
Normalize
4-stage pipeline
FAISS dedup
Store
DuckDB + Postgres KG
Serve
Marimo dashboard
Cloud Run
Supporting
FAISS · sentence-transformersrapidfuzzSkyPilot — spot GPUsGCP Cloud Run · Docker
1

Scrape

A Selenium-based ATS scraper plus public-API and SERP sources pull job postings from company career pages — checkpointed so an interrupted run resumes instead of restarting.

Seleniumundetected-chromedriverSERP API
2

Rule-based labeling

spaCy + regex (patterns compiled once at load) extract seniority, tech stack, and remote type as zero-cost "silver labels" — the training data for the LLM, bootstrapped without any manual annotation.

spaCyregex
3

QLoRA fine-tune Gemma-2B

Silver labels + a HuggingFace dataset are formatted into 10,000 Alpaca instruction pairs and used to QLoRA-fine-tune Gemma-2B-it (4-bit NF4) — on a free Colab T4 or a GCP spot GPU via SkyPilot.

QLoRA · Gemma-2BPEFT · bitsandbytesSkyPilot
4

Enrich & normalize

A vLLM client (exponential backoff, batches of 10 for KV-cache sharing, rule-based fallback) extracts structured fields, then the 4-stage normalizer resolves company names (fuzzy), canonicalizes tech, de-dupes with FAISS, and scores quality.

vLLMrapidfuzzFAISS dedup
5

Warehouse & knowledge graph

Enriched records land in a DuckDB analytical warehouse (enriched_jobs, tech_signals, company_signals) and a Postgres knowledge graph modelling org structure, tech stack, and hiring signals.

DuckDBPostgreSQL
6

Serve

A reactive Marimo dashboard is containerized and deployed to GCP Cloud Run; the fine-tuned model is exposed behind an OpenAI-compatible vLLM inference API.

Marimo · PlotlyCloud Run · Docker

04 Engineering decisions

4-stage normalization

Strip legal suffixes → exact match → rapidfuzz token-sort for word-order variants → tech alias resolution → sentence-transformer + FAISS dedup. Messy web data becomes joinable entities.

Explainable quality scores

An additive 100-point score stores a per-dimension breakdown, not just a total — so a consumer asking "why is this record low quality?" gets a root-cause answer.

Knowledge graph as the product

The Postgres graph — org structure, tech stack, key projects per company — is the actual deliverable, modelled directly on how GTM-intelligence products are consumed.

Graceful degradation

The LLM isn't a hard dependency — if the vLLM endpoint is down, enrichment falls back to the rule-based labeler and the pipeline keeps producing data.

DuckDB over BigQuery

In-process, reads Parquet/JSONL directly, and can ATTACH to Postgres as a foreign-data wrapper — one query layer over operational and analytical data, zero cloud credentials to develop.

Spot-GPU training

SkyPilot launches QLoRA fine-tuning on GCP spot instances — the same job runs free on Colab or cheap on preemptible cloud GPUs.

05 Stack

Scraping
Seleniumundetected-chromedriverSERP API
Labeling
spaCyrapidfuzzregex
Fine-tuning
QLoRAGemma-2B-itPEFT · bitsandbytesvLLM
Data
DuckDBFAISSPostgreSQLsentence-transformers
Serving
Marimo · PlotlyGCP Cloud RunDocker · SkyPilot

Try the live dashboard

The deployed Marimo app is public — explore the tech and company signals, then read how it's built.