Two NVIDIA GB10 nodes
A connected two-node environment for models requiring more shared memory and compute.
DUAL GB10 · LOCAL AI · LONG CONTEXT
Two NVIDIA GB10 compute nodes form an experimental environment for substantive analysis, engineering, RAG and testing new AI workflows. Local models are used when they fit the task—not simply for the sake of running locally.
Purpose
The lab can process large document corpora and long contexts without sending every request to an external provider. Important outputs still pass a separate quality gate, while a genuinely difficult stage can be deliberately escalated to a stronger cloud route.
Working environment
A connected two-node environment for models requiring more shared memory and compute.
The primary local route for substantive analysis, research and complex drafts.
The verified configuration can work with large source sets without an artificial few-thousand-token window.
Retrieval finds relevant documents and decisions first, then supplies the model with context and sources.
Extraction, analysis, coding, critique and visual work are separated and routed to suitable workers.
Results are judged by facts, completeness, reproducibility and practical value—not by confident wording.
Workflow
The observable outcome comes before the choice of model or technology.
RAG retrieves the smallest sufficient, verifiable context.
DeepSeek or a specialist model performs a bounded stage.
A weak result is repaired or deliberately escalated.
Practical lessons
Context must be real — a large advertised window is useless if the wrapper truncates input or output.
Stages should be separated — extracting facts, proposing, implementing and criticising require different assignments.
A routing defect is not model weakness — parameters, limits, finish reason and actual delivered context are checked first.
Accepted work matters — speed and savings count only when the result is good enough.
Related topic
The next architectural layer stores decisions, documents and retrieval indexes for primary and backup working environments.
How shared memory for AI agents works →