LOCAL FIRST · CODEX · QUALITY GATES · ESCALATION

A local model performs the stage.
Verification decides if it is accepted.

In the Digi Anton AI lab, suitable bounded work goes to a local model; Codex defines the goal, architecture, routing and acceptance criteria; deterministic tests and evidence gates evaluate the result; and a cloud model is brought in selectively—only for genuinely difficult or high-value stages. This article explains how the cycle works, which failures it has already exposed and what it does not promise.

Problem

Both extremes produce equally poor results

Sending everything to the strongest cloud model lets routine tasks consume limited cloud capacity that hard ones need, and increases provider dependence. Sending everything to a local model and accepting the answer on faith turns savings into a hidden loss of quality. There is also a third, less visible trap: the controlling agent misconfigures the experiment itself and then declares the model weak. A working route is therefore not a model choice but a chain: bounded stage → verification → a decision to accept or escalate. The overall role map is described on the model routing page.

Roles

Who is responsible for what

LOCAL · PRIMARY WORKER

DeepSeek V4 Flash

Runs locally on two NVIDIA GB10 nodes, without sending the request to a cloud provider. Handles substantive analysis, long-context work, fact extraction and bounded code implementation in an isolated scope.

LOCAL · ROUTINE

A lighter local model

Classification, extraction, short drafts, tool calls and recurring tasks. It can also act as an independent local critic where that fits the task better.

CONTROLLER

Codex

Defines the goal, architecture, decomposition, route and acceptance criteria, and reviews the diff and failure behaviour. It fixes small residual defects itself but does not silently rewrite the local model's entire stage.

DETERMINISTIC

Tests and evidence gates

Tests, linters, schema and format checks, finish reason, and refusal to accept empty, truncated or malformed output. These checks do not depend on how confident the model sounds.

CLOUD · SELECTIVE

Claude and other cloud models

Claude is a selective route for implementation or independent review of complex, high-value, disputed and final results. It is not a mandatory stage of every task, and its findings are verified too.

HUMAN

Major forks

A failed local-model stage, moving work to the cloud or to Codex, changing a model's role, choosing a product and publishing are decided by a human. Routine reversible details are not.

Work cycle

From task to accepted result

01 · GOAL

Codex writes the contract

An observable outcome, the smallest sufficient slice of sources, explicit non-goals and an acceptance criterion that is machine-checkable where possible.

↓
02 · LOCAL STAGE

The model performs a bounded part

Extracting facts, proposing, implementing, testing and critiquing are separate assignments. A single call is not asked to digest an entire corpus, invent a solution and accept it itself.

↓
03 · VERIFY

Deterministic gates

Tests, lint, format, completeness and truncation signals. The raw response, finish reason and parse errors are preserved so truncation is not confused with weak reasoning.

↓
04 · REVIEW

Codex reviews the diff and behaviour

Technical correctness and applied usefulness are checked separately, against the original goal rather than only the local contract.

↓
05 · REPAIR

Exact defects go back to the model

The local model receives a specific defect list for a bounded fix. A fundamental miss calls for a new decomposition or route, not a silent rewrite.

↓
06 · DECISION

Acceptance or deliberate escalation

The result is accepted, or the problematic stage is handed to a stronger route with a stated reason. Cloud escalation is explicit, bounded and auditable: it consumes limited cloud capacity.

Escalation

When a cloud model is genuinely needed

The cloud is justified when a local result fails a task-specific check and the task genuinely requires stronger reasoning, when the cost of error is high, or when a result is disputed or final and needs an independent view. It is not justified for scheduled routine jobs, simple extraction and classification, or as a safety net at every step. After two material failures of the same route on the same class of task, the route is not repeated: the model, decomposition, context or acceptance method changes, and that decision goes to a human. A failed local stage is a decision point, not automatic permission to hand the whole job to Codex or the cloud.

Diagnostics

A controller error must not pass for model weakness

01 · CONTEXT

How much context the model actually received

An advertised window means nothing if a wrapper truncated the input. Context window, batch size and output limit are three different parameters.

02 · OUTPUT

Why generation stopped

A “length” finish reason with no final answer, because reasoning consumed the whole allowance, is a configuration failure, not a model-quality failure.

03 · TIME

Whether the timeout was long enough

The local cluster wakes on demand, and a cold start takes minutes. A short client timeout looks like a model failure.

04 · STATE

RUNNING, COMPLETE, FAILED or UNKNOWN

Missing output, a timeout or an aborted outer tool does not prove failure. Before an expensive retry, the process, session and persisted result are checked.

05 · ROUTE

Wrapper, permissions and routing

Before drawing conclusions about a model, the whole parameter and data path is verified—not one line of configuration.

06 · MODEL TRAITS

Real limitations exist too

If a defect reproduces on a correct path, it is a property of the model. It is recorded and reflected in routing, not hidden.

Case from practice

8 thousand tokens instead of an available million

A fact from the lab's working records: in one experiment Codex limited a local model to roughly 8 thousand tokens even though the configured context window was up to 1,048,576 tokens, conflated the context window with batch size and output limit—and could then have declared the model or tool unusable. That is a controller defect: the fix belongs in the experiment setup, not in routing. A defect that still reproduces after the complete parameter and data path—context, limits, wrapper, routing and timeouts—has been verified is different. Only such a result counts as a genuine model limitation and a reason to change the route. Interpretation: both situations look like “the model did poorly”, but they lead to opposite decisions—in the first the controller is fixed; in the second the routing changes.

Acceptance

PASS, files and reports are not yet a result

Fact from practice — reports, files, PASS statuses, tests and registries were taken for user-facing results, and one work cycle was wrongly presented as complete although its original goals had not been met.

What now exists — which capability or external result exists that did not exist before the stage? If the answer is only a document, a test or a status, there is no progress.

Check against the original goal — acceptance compares the result with the original task, not only with the stage's local contract.

Remaining gap — what is still unachieved is stated explicitly; “done” is not used while the goal remains open.

Human minutes — if supervision takes more time than the result saves, the contour is not autonomous.

Did the decision change — a useful stage changes the next practical action; if it does not, it is research overhead.

How to read this article

Fact, interpretation and practical lesson

FACT

What the records confirm

DeepSeek V4 Flash runs locally on two GB10 nodes with a configured window of up to 1,048,576 tokens. Controller errors in experiment setup, mishandling of uncertain runs and acceptance of technical statuses as results are documented.

INTERPRETATION

What follows from it

Conclusions about a local model's quality or cost are only as reliable as the experiment that produced them. Codex is a strong bounded technical executor, but not a self-validating controller and not an AGI.

PRACTICE

What to do

Separate stages, give the model the smallest verifiable context, keep raw responses, verify the parameter path before judging a model, escalate explicitly and accept results against the original goal.

Boundaries

What this approach does not promise

It does not guarantee savings, full autonomy or benchmark superiority of one model over another. It is a working practice in an experimental lab: routes and builds change after verification, and some of the original goals remain open. What holds is the order in which a cheap stage, deterministic checks, review and selective escalation produce a result that can be trusted. Related parts of the architecture: model routing, the local AI lab and shared AI-agent memory.

How shared AI-agent memory works →