Building a Hybrid Coding Agent Architecture with Local and Frontier AI Models

Quick Answer
Most teams treat AI model selection as a binary choice — cloud frontier models with strong output but real API cost, or self-hosted models that avoid third-party token billing but may struggle with complex reasoning. This architecture takes a third path: frontier models handle planning and failure diagnosis; a local quantized model handles the execution loop; and a deterministic verification gate — triggered by the verification suite, not the model — decides when to re-engage the frontier. The result moves a significant share of routine inference off the API bill while keeping frontier capability available at the stages that need it most.
Split work by cognitive load, and let a real verification signal — not the model's own claim — decide when to escalate.
The Problem With Current AI Dev Setups
Development teams running AI-assisted workflows face a persistent trade-off: powerful cloud frontier models produce strong plans and complex reasoning, but every token costs money — and in a team with many developers making many changes per day, that cost compounds quickly. Self-hosted models eliminate third-party per-token API charges — although they still incur infrastructure and operational costs — and keep code on-premises, but they struggle on tasks that require reasoning across a whole codebase, and they have no reliable mechanism to know whether their output actually works.
The less-discussed problem with local-only setups is not just raw output quality. It is that local models will confidently report completion on code that fails immediately. Without an external verification signal, you are trusting the model's self-assessment — which is not the same as trusting the build.
This post describes a reference architecture that addresses both problems: a pipeline where the frontier model handles planning and failure diagnosis, a local model handles the execution loop, and a deterministic verification gate — driven by the verification suite — decides when escalation is needed.
What Is a Hybrid AI Development Workflow?
A hybrid AI development workflow splits the coding pipeline between two categories of model based on the cognitive demands of each task.
Frontier models (Claude Sonnet, GPT-5.x) are large cloud-hosted models with strong reasoning, broad context windows, and the ability to plan coherently across a complex codebase. They are well-suited for upfront design, cross-file analysis, and targeted failure diagnosis. They are also billed per token by a third-party provider.
The local tier refers to self-hosted or organisation-managed inference running on hardware the organisation controls — a developer workstation, an on-premises GPU server, or a private cloud GPU instance. Inference on this tier has no third-party per-token API charge, though it does carry hardware, hosting, electricity, and operational costs. The model used here is Qwen3.6 35B-A3B, a Mixture-of-Experts (MoE) model, run at Q4_K_M quantization via Ollama. The local tier is not fixed to this choice — as the ecosystem matures, the model can be replaced through the inference configuration without changing the architecture.
On MoE architecture: A Mixture-of-Experts model routes each token through a subset of specialised sub-networks (“experts”) rather than the entire model. In Qwen3.6 35B-A3B, the “A3B” indicates approximately 3.6 billion parameters are activated per inference pass, out of a total of 35 billion. This reduces computation per token relative to a comparably-sized dense model, which means faster inference and lower energy usage during generation.
One important clarification: active parameter count does not equate to a smaller memory footprint. The complete set of expert weights — all 35 billion parameters — must still reside in memory so the router can dispatch tokens to any expert at runtime. MoE reduces compute per token; it does not reduce the stored model size. At Q4_K_M quantization (approximately 4 bits per weight), the full model occupies around 22–24 GB, which is what determines your VRAM requirement.
On quantization: Full-precision models store each weight as a 16-bit or 32-bit float. Q4_K_M compresses weights to approximately 4 bits using a mixed-precision scheme that applies finer quantization to layers where precision matters more. The practical effect: model size drops by roughly 75% versus fp16, and for code generation, 4-bit K-quants typically show little perceptible degradation versus the full-precision model.
A hybrid workflow assigns each task to the appropriate model tier and uses a deterministic mechanism — not a prompt asking the model if it is finished — to decide when to move between them.
The Core Architecture: Right Model for the Right Job
The central insight is task decomposition: not every part of development requires the same reasoning capacity.

Planning requires reasoning across the whole codebase, understanding trade-offs, and producing a coherent strategy. That is where frontier models create the most value relative to their cost.
Execution — writing actual code changes guided by a clear plan — is where local models perform well on well-scoped tasks. They avoid provider token billing, keep code within your network, and cycle through edit-apply-verify loops quickly.
The escalation trigger is the key design decision: re-engagement of the frontier is driven by a failed verification signal, not by asking the model whether it is done.
Where Spec-Driven Workflows Such as BMAD Fit
This article describes the model-execution and verification layer — which model runs at each stage, how verification gates escalation, and how context is scoped. A complementary layer is work decomposition and specification: defining what should be built, why, and how it breaks into implementable units.
In a spec-driven (BMAD-style) workflow, durable artefacts — product requirements, architecture decisions, implementation stories, and acceptance criteria — are produced before implementation begins. The frontier model can use them as direct input when producing a plan; well-scoped stories give the local model the bounded context it performs best in; acceptance criteria supplement tests in the verification suite; and when escalation fires, the frontier receives the story, acceptance criteria, failure output, and scoped code context together. BMAD is an example integration, not a dependency — the architecture is equally compatible with internal SDLC processes or lightweight specification files.
Four Components That Make It Work

1. Frontier Wrapper — Reuses Existing Provider Authentication
The frontier wrapper is a lightweight local server that speaks the standard OpenAI API format. OpenCode sends requests to it as if it were any cloud provider. The wrapper translates those requests into CLI calls to Claude Code or Codex — tools most developers already have authenticated and configured. This avoids adding a new model provider integration or a new set of credentials. On developer machines, existing CLI authentication may be reused; in CI, non-interactive authentication must be supplied through appropriately scoped CI-managed secrets.
2. Automatic Planning Delegation
When a developer opens the Plan tab, the request routes automatically to the frontier model. No manual switching per task. The plan lands in the shared session context where the local model reads it as part of the conversation history. This delegation is architectural: the routing is enforced by the configuration, not by developer discipline.
3. Deterministic Verification-Gated Escalation

The escalation decision is driven by deterministic verification signals rather than the model's own completion claim. A language model — local or frontier — cannot determine from generated code alone whether a change compiles, passes tests, or satisfies type constraints. The agent harness must execute those checks and return the results. A verification failure is a hard, external signal that model self-assessment cannot replicate.
When escalation fires, the frontier call begins with the failure output, the edited files, and a concise summary of the attempted change. The harness may then selectively retrieve related tests, definitions, interfaces, or call sites when needed. The full codebase is not automatically re-ingested, which keeps per-escalation token usage predictable. A retry limit (default: 3 attempts) prevents runaway loops, surfacing genuinely broken requirements as an explicit developer-review state.
On what “verification” means here: the gate is as strong as what you put in it. A production-grade configuration should include static type checking, linting, build success, and where applicable contract tests or security scans — not tests alone. The diagram says “verification passes” rather than “tests pass” precisely because the gate should be composed of multiple signals.
4. Knowledge Graph for Codebase Navigation

When the relevant files are not already known, the agent can query a structural map of the codebase before expanding its context. The map is built by static analysis — no AI involved — and stored in FalkorDB, a graph database that answers traversal questions like “what imports auth.js?” or “where is the checkout function defined?” in a single query.
A graph database is used here rather than a vector store because the query is structural, not semantic. “What imports this module?” has an exact answer derivable from the AST — it does not require embedding similarity. Graph queries return precise results with near-zero token cost and latency well below a model round-trip. For large codebases, this changes navigation from “load everything and let the model skim” to “query first, read only what matters.”
Where Frontier Tokens Actually Get Spent

A typical feature involves one planning call and zero to a small number of escalations. Dozens of edit-apply-verify cycles run on self-hosted infrastructure with no third-party API charge. In codebases with a meaningful test suite and clear task decomposition, this architecture concentrates frontier spend at the stages where it creates the most leverage — upfront reasoning and targeted failure diagnosis — while shifting the repetitive execution loop to organisation-managed inference.
The actual cost reduction will depend on codebase complexity, test coverage, task scope, and the local model's first-pass success rate on well-specified plans. Teams should measure their own baseline before projecting savings.
How Developers Actually Use It
The day-to-day experience collapses to four steps:
- Open the editor.
opencode /your/project— one command. - Plan the task. Switch to the Plan tab, describe what you want. The frontier model writes a structured implementation plan.
- Execute. Switch to the default tab, say “implement the plan.” The local model works through it — editing files, running commands, iterating.
- Wait for green. If verification fails, escalation is handled automatically. When everything passes, review the diff and commit.
Non-secret configuration — model routing rules, escalation retry limits, and verification commands — can be version-controlled in a single file checked into the repo. Provider credentials and CI authentication must come from environment variables or secret-management systems and should never be committed.
Is This Architecture Right for Your Team?
Strong fit:
- Multiple engineers where frontier API cost is a meaningful and growing line item.
- A codebase with real test and type coverage — the verification gate is only as strong as what it runs.
- Code privacy requirements where adding another model provider endpoint is not acceptable.
- Teams already using Claude Code or Codex CLI who want to extend that investment without new accounts.
Consider a different approach if:
- You have little or no test coverage — without a meaningful verification signal the escalation gate has nothing to trigger on.
- You are building short-lived prototypes where verification infrastructure does not yet exist.
- Your team is small (1–2 engineers) and API cost is not yet a problem worth solving.
Interested in how this architecture applies to your stack?
We'll look at your verification setup, codebase, and cost profile, and map the fastest path to a hybrid agent workflow.
Talk to our engineering team →Reference Configuration and Hardware
The reference configuration for this architecture:
- Frontier tier: Claude Sonnet via Claude Code CLI — planning and escalation diagnosis.
- Local tier: Qwen3.6 35B-A3B at Q4_K_M quantization, served via Ollama.
- Knowledge graph: FalkorDB, one instance per project.
- Graph updates: Incremental — changed files are re-parsed and affected nodes updated on each CI run; a full rebuild runs periodically for consistency.
- Escalation limit: 3 retries before human review; every frontier call logged as structured JSON.
- Deployment time: new project setup including the initial graph build is a same-day task.
Hardware for the local model tier: at Q4_K_M, the full Qwen3.6 35B-A3B model occupies roughly 22–24 GB of quantized weights. The requirement is that the entire model fits in GPU VRAM — offloading to system RAM degrades inference speed enough to break interactive developer flow. You need enough VRAM headroom above the model weight size to hold the active context and KV cache for your typical task size.
Summary
The architecture's core claim is not that it eliminates frontier model use. It constrains frontier use to the stages where reasoning quality matters most, keeps routine execution on self-hosted infrastructure, and makes escalation deterministic and auditable.
| What | How |
|---|---|
| Planning | Frontier model — automatic routing, no manual switching. |
| Execution | Local tier (Qwen3.6 35B-A3B Q4_K_M via Ollama) — no API token charge, within your network. |
| Escalation trigger | Verification-suite failure (tests, type checks, build) — not model self-assessment. |
| Codebase navigation | Knowledge graph (FalkorDB) — structural queries, not embeddings. |
| Provider credentials | Reuses existing CLI authentication on dev machines; CI uses secret-managed credentials. |
| Audit trail | Structured log per frontier call: trigger, metadata, result, retry count, usage. |
| VRAM requirement | Model weights ~22–24 GB; headroom needed above that for context and cache. |