The harness for governed multi-agent engineering

The discipline your AI tools are missing.

CortexBuild is the governed harness for the AI coding tools you already use — Copilot, Claude Code, Codex, and Cursor. It renders your team's rules and a disciplined review workflow into those tools, so the agents you already run can build on your existing repos and systems — new features, services, endpoints, and refactors, as well as legacy modernisation and backlog burn-down — under enterprise governance and your own LLM keys.

The model is a commodity; the harness is the moat. CortexBuild is the harness — the control loop, governance, and guardrails that make any model ship production-grade work.

BYO LLM keys User-controlled approval gates Open-core engine

Runs inside Copilot, Claude Code, Codex & Cursor · routes your own LLM keys · works where your code already is

Claude Gemini OpenAI-compatible Ollama
GitHub GitLab Bitbucket Azure DevOps

The two jobs you can't fund

Every engineering org carries two jobs it can't fund.

01

Modernising legacy code that works but rots

The code that keeps the business running is the code no one has time to touch. Every quarter it gets riskier to change and more expensive to maintain.

02

A backlog of small tickets that pile up

The fixes, chores, and tidy-ups that never reach the top of the sprint. Individually small, collectively a tax on every team and every roadmap.

How it works

From repo to reviewable PR — in your normal git flow.

No rip-and-replace. CortexBuild meets your code where it already lives and works the way a disciplined team would.

  1. 1

    Connect your repo

    Point CortexBuild at your repository on GitHub, GitLab, Bitbucket, or Azure DevOps.

  2. 2

    It analyses your conventions

    The brownfield analyser reads your code and infers your patterns — not a generic template.

  3. 3

    Approve the roadmap

    Review a plain-English plan and approve what the agents are allowed to work on.

  4. 4

    You ship reviewable PRs, your way

    Your existing agents in Copilot, Claude Code, Codex, or Cursor work the roadmap through a rendered plan → review → verify workflow; you open the PR in your normal git flow when it's ready.

Onboarding promise A working .cortexbuild/ — your repo analysed, conventions captured as rules, and the harness rendered into your AI tools — in 30 minutes or less.

The differentiator

A multi-agent wave model — not a single bot.

CortexBuild renders work as a sequence of specialised review roles into the AI coding tools you already run. The Critic reviews the plan before any code is written, and read-only analyzers run in parallel — so your tools follow the way a senior team actually ships.

Adversarial review, before code

The Critic challenges the plan's assumptions and missed requirements before a single line is written — catching the most expensive failure: shipping the wrong thing correctly.

Parallel read-only analysis

Tester, Security, and Optimizer all read the same diff at once. Three independent checks, no write conflicts, faster wall-clock per task.

One writer at a time

Only the Builder mutates code, and only after the plan is approved — so changes stay coherent and easy to review.

Why the harness is the moat

The model is a commodity. The harness is the moat.

Frontier models converge on the same capabilities and turn over every few months. What compounds is the harness around them — the control loop, governance, evals, and feedback that turn any model into a disciplined engineering team. CortexBuild is that harness, and it is model-agnostic by design.

A rendered control loop

CortexBuild renders a fixed Planner → Critic → Builder → Tester → Security → Optimizer → PR-Reviewer workflow into your AI tools, so your work follows the same plan → review → build → verify order every time, regardless of which model you use. The harness defines the workflow — not the model.

Governance rendered to every surface

Rules are authored once and rendered to every agent surface, so the same policy holds no matter which model or tool runs the task.

Deterministic checks, not vibes

Conformance scoring, the offline eval harness, render --check, and recurrence-ledger guards are deterministic: no LLM in the loop, same result whichever model produced the change. The Critic, Tester, Security, and Optimizer roles add LLM review on top. CortexBuild flags failures; it does not enforce merges by itself — your existing approvals decide what lands.

A feedback loop that compounds

The recurrence ledger records every fixed class of failure as a guard, so the harness catches the same mistake earlier on the next run.

Model-agnostic, BYO-keys

A provider router lets you bring your own keys and route LLM providers: native Claude, native Gemini, and any OpenAI-compatible endpoint (OpenAI/Codex models, gateways, and local servers such as Ollama when run in OpenAI-compatible mode) — without rewriting your harness.

Swap the model, keep the moat

When the next frontier model lands, you point the router at it. The governance, evals, and feedback that make it safe to ship are already yours.

Open-standard interop: MCP & A2A

CortexBuild speaks both halves of the agent-interoperability standard. A read-only MCP server lets any IDE or agent host (Claude, VS Code, Cursor, Codex) consult your repo's conformance, drift, context gaps, and corrective proposals — and run an engineering persona as a prompt. An A2A Agent Card lets other agents discover CortexBuild's roster and skills. Both are read-only, offline, and opt-in: they expose what CortexBuild already knows and never write code or send data anywhere.

The roster

Ten specialised agents. One disciplined team.

Planner

Breaks work into a sequenced, repo-specific roadmap.

Critic

Adversarially reviews the plan before any code is written.

Builder

Implements changes — the only agent that mutates code.

Tester-Validator

Runs the test matrix and synthetic probes for evidence.

Security-Compliance

Checks for security, privacy, and tenant-isolation regressions.

Optimizer

Evaluates scalability, reliability, performance, and cost.

PR-Reviewer

Final merge gate — verifies correctness and coverage.

Explore

Fast, read-only codebase research and Q&A.

Recurrence-Guard

Turns past failures into permanent guardrails.

Enterprise-Orchestrator

Enforces strict phase sequencing across the team.

Built for control

Governance and quality, baked in.

Rule single source of truth

Write a rule once; render it to Copilot, Claude, Cursor, and Codex automatically. No drift across tools.

Recurrence ledger

Agents learn from past failures. Every fixed class of mistake becomes a guard that blocks it from recurring.

BYO LLM keys

Your models, your spend. Code and inference stay under your control — never routed through us.

Approval gates

Sensitive actions are consent-gated behind an explicit approval, so you decide what agents may do without a human. On the roadmap: a configurable matrix of action classes across autonomy levels.

Brownfield analyser

Infers your conventions from your existing code — no greenfield assumptions, no template lock-in.

Self-hostable engine

The open-core engine installs and runs entirely on your own infrastructure — your perimeter, your keys. On the roadmap: a self-hostable VM sandbox for the agent runtime.

Audit log & governance

Every permission-affecting change is recorded — SOC2-friendly trails for compliance reviewers.

Conformance score

A 0–100 measure of how well a change adheres to your standards, surfaced before you merge.

Cost dashboard & budget gates

Track model spend per task and set budget gates that pause work before costs run away.

Supports local & self-hosted models

Point the harness at any OpenAI-compatible endpoint — cloud, or a model you serve on-box with Ollama, LM Studio, or vLLM. LLM egress is locality-gated: a loopback model on the same machine runs zero-egress with no consent prompt; LAN and cloud endpoints pass an explicit egress gate first. If you expose an OpenAI-compatible endpoint from workstation-class AI hardware — for example an NVIDIA DGX Spark — CortexBuild reaches it through the same generic endpoint path.

Always-current code graph

Every repo gets a module-dependency graph — god-nodes, communities, per-language stats — built on onboarding and kept current incrementally (it rebuilds only when your sources actually change). Across a multi-repo workspace, a cross-repo dependency graph shows how features connect (frontend → backend → shared libs), so a control plane can see which repos a change ripples into. Deterministic and offline — no code leaves the machine.

Portable memory — switch device, hand off teammate

Save a compact, tag-only snapshot of where a session left off and restore it on another device or hand it to another teammate — "continue where Alice left off." Scoped to a stable repo identity, attribution is opt-in, and it carries summaries and tags only — never raw code, prompts, or secrets. Hosted sync is consent-gated and erasable.

Feature utilization & subscription control

See every capability in one place — its plan tier, whether you're entitled, and whether it's actually used — so you can spot unused features and right-size your subscription. Usage is captured at the enforcement point, so features used only via the API still count. Admins enable or disable features within the plan ceiling; every change is subtract-only, signed, and audit-logged.

Use cases

Build new work, modernise the old, burn down the backlog — one governed team.

Modernise without the rip-and-replace

Point CortexBuild at the module everyone is afraid to touch. It infers your conventions, proposes an incremental roadmap, and drives your AI tools to prepare governed changes — each one tested, security-checked, and reviewable — that you open as PRs. Legacy debt comes down a safe step at a time.

  • Incremental, reversible changes — never a big-bang rewrite
  • Tests and security findings attached to every PR
  • Estimated $0.50–$5 per modernisation PR in model cost

ROI calculator

Estimate what your backlog is costing you.

Move the inputs to see an estimate of hours and dollars saved each month. The math is deliberately conservative and fully shown below.

Estimated monthly impact

60 engineer-hours saved / month
$5,392 net saved / month
Gross labour saved
$5,400
Estimated LLM spend
$8
Tickets automated / month
20

Estimate — your results will vary. For planning only; not a guarantee of performance.

The formula, in plain terms

automated      = tickets/mo × (autonomous share ÷ 100)
hours saved    = automated × hours per ticket
gross labour $ = hours saved × engineer $/hour
LLM spend $    = automated × LLM $ per ticket
net saved $    = gross labour $ − LLM spend $

How we compare

Built differently — on purpose.

Comparison of CortexBuild against single-agent tools and closed autonomous agents
Capability CortexBuild Single-agent tools Closed autonomous agents
Multi-agent wave model Yes No Partial
Adversarial Critic (before code) Yes No No
Rule single-source-of-truth fan-out Yes Partial No
Recurrence ledger / learning from failures Yes No Partial
Bring your own LLM key Yes Partial No
Self-hostable VM Yes No No
Approval governance Yes Partial No
Open-core Yes No No

Born in production

Proven before it was a product.

CortexBuild is the generic extraction of an agent governance system proven on CortexVigil, an enterprise surveillance platform. The wave execution model, Rule single-source-of-truth, recurrence ledger, adversarial Critic, and self-audit discipline were all forged shipping real software under real constraints.

  • Rule SSoT
  • Wave execution
  • Recurrence ledger
  • Adversarial Critic
  • Self-audit YAML

Pricing

Start free and self-hosted. Scale when you're ready.

Indicative pricing for early access. With BYO LLM keys, you pay model costs directly — roughly $0.05–$0.80 per backlog ticket and $0.50–$5 per modernisation PR.

Open Core

Free

OSS engine, self-hosted

  • cortexbuild-core engine
  • Self-host on your infrastructure
  • Bring your own LLM keys
  • Community rule packs
View on GitHub

Enterprise

Custom

For regulated & large teams

  • SSO & role-based access
  • 7-year audit retention
  • Private rule packs & support
  • Self-host option
Talk to us

Indicative / early access pricing — subject to change before general availability.

FAQ

Questions, answered plainly.

Put your backlog on autopilot — with guardrails.

Join the early access list. Connect your repo at launch and get it analysed, with your conventions captured as rules and the harness rendered into your AI tools — in under 30 minutes.

Early access — we'll email you when you can connect your repo. No spam.