DUAL GB10 · LOCAL AI · LONG CONTEXT

An AI lab
where models run locally.

Two NVIDIA GB10 compute nodes form an experimental environment for substantive analysis, engineering, RAG and testing new AI workflows. Local models are used when they fit the task—not simply for the sake of running locally.

Purpose

More control over compute, data and experiment cost

The lab can process large document corpora and long contexts without sending every request to an external provider. Important outputs still pass a separate quality gate, while a genuinely difficult stage can be deliberately escalated to a stronger cloud route.

Working environment

What the laboratory contains

COMPUTE

Two NVIDIA GB10 nodes

A connected two-node environment for models requiring more shared memory and compute.

CORE ANALYSIS

DeepSeek V4 Flash

The primary local route for substantive analysis, research and complex drafts.

LONG CONTEXT

Up to 1,048,576 tokens

The verified configuration can work with large source sets without an artificial few-thousand-token window.

MEMORY

QMD and RAG

Retrieval finds relevant documents and decisions first, then supplies the model with context and sources.

ORCHESTRATION

The task chooses the model

Extraction, analysis, coding, critique and visual work are separated and routed to suitable workers.

ACCEPTANCE

Tests and quality gates

Results are judged by facts, completeness, reproducibility and practical value—not by confident wording.

Workflow

From task to accepted result

01 · GOAL

Applied task

The observable outcome comes before the choice of model or technology.

↓
02 · CONTEXT

Sources and memory

RAG retrieves the smallest sufficient, verifiable context.

↓
03 · COMPUTE

Local route

DeepSeek or a specialist model performs a bounded stage.

↓
04 · ACCEPTANCE

Test or expert review

A weak result is repaired or deliberately escalated.

Practical lessons

A local model is part of a system, not the goal

Context must be real — a large advertised window is useless if the wrapper truncates input or output.

Stages should be separated — extracting facts, proposing, implementing and criticising require different assignments.

A routing defect is not model weakness — parameters, limits, finish reason and actual delivered context are checked first.

Accepted work matters — speed and savings count only when the result is good enough.

Related topic

Compute becomes useful through shared memory

The next architectural layer stores decisions, documents and retrieval indexes for primary and backup working environments.

How shared memory for AI agents works →