AgentForge
A repository-level coding agent built around a strict plan → code → review → container test → repair lifecycle. A LangGraph supervisor coordinates specialized workers, governed MCP tools, repository-scoped retrieval, durable checkpoints, and deterministic evidence gates so completion means the patch was actually observed and verified.
01 Overview
AgentForge starts from a repository selected through an allowlisted project picker. It maps the codebase deterministically, creates an exact-revision Git worktree, retrieves only repository- and commit-scoped context, then gives each worker the minimum tool capability required for its role.
The point isn't agent count; it's trustworthy autonomous execution. The system refuses to complete without an observed workspace change, independent review and security approval, and a successful test run in a disposable Docker container. Failed gates feed a bounded repair loop; contradictions fail closed.
02 Results
03 Architecture
A single supervisor owns lifecycle state, budgets, approvals, and termination. Workers operate on one isolated repository revision and can advance only when deterministic gates produce evidence.
Select and understand the repository
The user selects an allowlisted Git repository. AgentForge records its exact commit, creates a detached worktree, detects languages, dependencies, entry points and tests, then builds a repository-scoped index.
Plan with bounded context
Qwen produces a structured plan from the repository map and reranked AST-aware retrieval results. Token reservations and task budgets are applied before every provider call.
Implement through governed tools
GPT-OSS 20B edits only the task worktree through a capability-limited MCP scope. Protected paths, secret detection, write and command budgets, and human approval tokens constrain risky actions.
Verify the observed artifact
The supervisor hashes the real diff, sends that delivered change set through review and security gates, then runs the repository test command in a disposable container with network, CPU, memory, PID and timeout restrictions.
Repair or terminate truthfully
Failed gates return structured feedback to a bounded repair pass. Completion is allowed only when every required criterion points to recorded evidence; cancellation, failure and recovery remain explicit terminal states.
04 Reliability engineering
The engineering focus is controlled autonomy: useful agents, narrow permissions, reproducible context, and truthful outcomes.
Disposable execution
Generated commands run in short-lived Docker containers with no network, a read-only root filesystem, dropped capabilities, and explicit CPU, memory, PID, output and wall-time limits. Missing Docker fails closed.
Durable execution control
PostgreSQL queue records, worker leases, heartbeats and graph checkpoints support cancellation, resume, and startup recovery without treating a transport queue as the source of truth.
Four-layer memory
Working state, verified run episodes, reusable repository facts, and proven procedures have different scopes and retention policies. Only cross-validated completed runs create durable memories.
Governed MCP tools
Filesystem, terminal, GitHub and browser capabilities are task-scoped and actor-specific. Protected paths, secret redaction, call budgets, idempotency and signed one-use approvals constrain side effects.
Cost-aware inference
Role-based model routes, conservative token reservations, bounded retries, exact caching for safe read-only work, batched embeddings, and cache-savings telemetry control free-tier usage.
Repository-scoped retrieval
Tree-sitter symbol chunks, dense retrieval, BM25, reciprocal-rank fusion and deterministic reranking are keyed by repository, commit and content hash to prevent cross-project contamination.
05 Stack
See the graph in code
The repository includes the full orchestration graph, governed tool layer, Docker executor, retrieval and memory evaluations, dashboard, and 206-test regression suite.