TASKS · DATA · LOAD · HYBRID ROUTING
A local AI server is not the first decision.
The first is what work it will do.
Businesses are often offered a local AI server as a way to save on the cloud or to protect data. Sometimes it genuinely fits, but the first constraint is usually not compute power — it is tasks, data, quality checks and operations. Below: when local infrastructure fits, when the cloud is better, why a hybrid route often wins, what work begins after the purchase and in what order to make the decision. Experience from the Digi Anton local AI lab is cited as practical evidence with explicit limits, not as a finished product or a reference configuration to buy.
The starting question
A server is a means, not a goal
The question “should we buy an AI server?” is usually asked too early. A more useful starting point: which recurring tasks the AI system must perform, on what data, at what load and at what cost of error — and which route delivers an accepted result at lower total cost and risk. A local server is only one possible answer. “Local” does not by itself mean cheaper, safer or better quality: it means that compute, updates, availability and protection become your responsibility. Facts from the lab's working records are marked explicitly in this article; all other statements are general implications and recommendations for a business owner, worth testing against your own tasks and data.
Local route
When local infrastructure is appropriate
A regular, predictable flow of tasks
A lot of similar work every day: extraction, classification, drafts, document search. Ownership economics depend on utilisation: an idle server is an expensive way to handle occasional requests.
Data that cannot or should not leave the company
Internal policy, a contract or a regulator restricts sending documents to an external model provider. Local processing removes that particular transfer channel, but it does not replace access control, logs and backups.
Large internal corpora and long documents
Contracts, archives, technical documentation, correspondence — when they must be analysed regularly in full or in large parts. Long context helps, but it does not replace data preparation and search.
You need control over the model version
It matters that the model and its behaviour do not change without your decision: a provider may update, replace or retire a model, or change prices and terms. Locally, you decide when to update — and you are responsible for the updates.
An open model is good enough for your task
A model that can run locally passes your own test on real examples, not just general benchmarks. If only the strongest cloud model delivers acceptable quality, a local server will not solve that task.
Someone will run the environment
A specialist or contractor owns updates, monitoring, backups, availability and regular quality checks. Without an operations owner, a server quickly becomes an unused asset.
Cloud route
When the cloud is the more sensible choice
A pilot, seasonality or occasional requests
Until volume is measured, paying per use is more sensible than capital expenditure. First find out how many requests, tokens and peaks actually occur.
You need the strongest model
Complex analysis, rare hard tasks and final review often require the most capable models, and these are frequently available only in the cloud.
You need to test a hypothesis quickly
A cloud API lets you test a scenario's value within days, without buying, installing and configuring hardware. If the scenario does not prove its value, no idle hardware is left behind.
No one to run a local environment
The provider takes care of availability, updates and scaling. Without in-house skills and time, a local server is more likely to add risk than to remove it.
Load changes sharply
The cloud scales to peak demand within the provider's limits, while your own server has fixed capacity: at peaks requests queue, and the rest of the time the hardware sits idle.
The process cannot stop, and there is no standby
A single local server is a single point of failure: a breakdown, a power cut or a failed update stops the process until recovery. The cloud can be unavailable too, but your own standby means separate spending on second hardware and its upkeep.
Hybrid
Why a hybrid route often wins
A real process usually contains both routine work and rare hard stages. A hybrid route separates them: a local model performs bounded recurring stages — extraction, classification, first drafts, search over internal documents; deterministic checks evaluate the result; and a cloud model is brought in explicitly for complex, high-value, disputed or final stages. The question “local or cloud” then turns into a different one: which stage goes where, under what rule, and how acceptance is checked. Fact from the lab's working records: recurring tasks there are performed by local models when their quality is sufficient; paid cloud models are not an implicit route for scheduled jobs, and escalation to them is explicit and bounded. Recommendation: decide in advance what happens when the local environment is unavailable. For data that must not leave the company, the fallback can only be another local route or waiting — not the cloud. How stages and checks are divided is covered in the article on routing and quality gates.
Operations
The work that begins after the purchase
Corpus preparation
Document owners, current versions, duplicates, scans and spreadsheets, access rights. A model leans on an outdated document as confidently as on a current one, so data quality is the first line of work.
Search that finds the right passage
Document chunking, hybrid search, source citations, a sufficiency threshold and index updates when documents change. More in the article on RAG for a company knowledge base.
The layer between the model and the process
The model server, the client, response formats, parsing, retries and timeouts. Many “model errors” actually originate here: a wrapper truncates the input, cuts off the answer or loses the result.
What the model actually receives
The advertised context window, the input actually delivered and the output limit are different quantities. Long input increases processing time and memory load, so the model needs the smallest sufficient slice, not the whole archive.
Real tasks and an acceptance criterion
Examples from your own work with expected answers; checks for completeness, truncation and format; a re-run after every change of model, prompt or software. Without this, any comparison is an opinion.
Electricity, cooling, downtime
Power draw, heat removal, noise, an uninterruptible power supply, space. If hardware is switched off or put to sleep to save energy, the first request after waking waits for the model's cold start.
Updates, backups, rollback
Drivers, the runtime and models get updated and are not always compatible with one another. You need backups of configuration and indexes, a change log and a tested way to roll back.
Local does not mean protected
Access rights, network exposure, action logs, secret storage, physical access to the hardware and protection of backups remain your responsibility.
Minutes for supervision and fixes
Monitoring, failure analysis and acceptance of results take specialists' time. If supervision takes more than the automation saves, the environment has not become autonomous.
Evidence from the lab
Fact, general implication and recommendation
Context window versus scheduler batch size
In the lab, DeepSeek V4 Flash runs locally on two NVIDIA GB10 nodes in a verified configuration with a context window of up to 1,048,576 tokens. Separately, the inference server has its own scheduler batch-size control: this is neither the model's context window nor the output limit.
A large window is a ceiling, not a guarantee
The context window, the scheduler batch size and the output limit are different controls, and quality depends on what the model actually received. Confusing them is a controller or measurement error, not a model weakness, and it easily leads to a wrong conclusion about the model, the route or a hardware purchase.
Verify the data path before drawing conclusions
Before judging a model, check the context actually delivered, the output limit and why generation stopped, and do not mistake scheduler parameters for model limits. Give the model the smallest sufficient slice through search, not the whole archive.
Wake on demand and cold start
The lab's local cluster wakes on demand, and a cold start takes minutes. In one case model loading stalled while the processes stayed active; a separate check of start-up progress was added afterwards.
“Server on” and “model answering” are different states
Saving energy through sleep or shutdown is paid for with first-response latency and extra logic. A short client timeout looks like a model failure.
Plan for outages in advance
Monitoring should check a real model response, not just a running process; timeouts should allow for a cold start. For each process, decide: queue, wait or another route.
A known model limitation
The local version of the model used in the lab has a known limitation: inconsistent arithmetic. Disabling one of the generation-acceleration mechanisms did not remove it and made the work noticeably slower: the cause lies in the model, not the cluster.
Hardware does not fix model properties
Running locally does not make a model stronger, and tuning the hardware does not remove its limitations. Quality has to be tested on your own tasks, not inferred from the model's name.
Compute numbers in code and keep a test set
Verify calculations with deterministic code, not the model's text. Run a set of real tasks before buying and after every update of the model or runtime.
Economics
Count the total cost, not the server price
A comparison of “server price versus the monthly API bill” is almost always incomplete. On the local side: the hardware and its obsolescence, electricity and cooling, space, backup power, backups, monitoring, updates, specialists' time and downtime after failures. On the cloud side: the cost of requests at measured rather than assumed load, peaks, changes to a provider's prices and terms, restrictions on data transfer and dependence on a single vendor. The decisive variable is utilisation: the cost per task on your own hardware falls as load grows and rises sharply when the hardware sits idle. The second condition is quality: savings and speed matter only when the result passes acceptance; otherwise a cheap route is paid for with people's time spent on corrections. The figures for the calculation are best taken from a pilot — choosing a scenario and checking quality are covered in the article on implementing AI in a business process. For a first estimate, the free local AI economics calculator will do: it exposes its assumptions and does not promise a financial result.
Decision sequence
Seven steps before buying hardware
List the tasks and a measurable result
What should appear regularly: a verified report, an answer with a source reference, a processed document, code. For each task — the cost of error and who accepts the result.
Classify data by permitted route
What may go to a cloud provider, what only after anonymisation, and what must not leave the company's environment. This constraint shapes the architecture more than the price of hardware.
Build a task set for quality checks
Real examples with expected answers and an acceptance rule. Without it, choosing a model and hardware remains guesswork.
Test the route without buying
Through a cloud API or rented compute — where possible on the same open model you plan to run locally; for restricted data, on an anonymised sample. Measure quality, request and token volume, latency and people's minutes spent on supervision.
Calculate the total cost at measured load
Compare cloud, hybrid and local options on pilot figures rather than expectations — including maintenance, downtime and specialists' time.
Appoint an owner and an unavailability plan
Who updates, monitors, makes backups and repeats the quality check after changes. What happens to each process if the local environment is unavailable.
Choose the route, then the hardware
Cloud, hybrid or local — based on the results of steps 1–6. Hardware is sized to measured load, and after installation the quality check is repeated on the local version of the model. Fact from the lab's working records: its architecture also allows hardware expansion only on measured demand.
Stop signals
When the purchase should wait
No task list — there is a wish to “adopt AI”, but no defined result that should appear regularly.
Load not measured — nobody knows how many requests, tokens and peak periods there will actually be.
Quality not tested — the model was chosen by name or ranking, not by a run on your tasks.
No operations owner — it is unclear who updates, monitors and answers for downtime.
“It is safer” without analysis — access rights, logs, backups and network exposure have not been examined.
Savings expected, people's time not counted — supervision, corrections and downtime are missing from the calculation.
Boundaries
What this article does not promise
It does not guarantee savings, confidentiality or security: local processing removes one specific risk — sending requests to an external model provider — and does not replace access control, logs and backups. The lab is an experimental working environment, not a finished commercial product and not a reference configuration to buy; its facts show which questions arise in practice, but the figures for your business will come only from your own measurements. Related reading: the local AI lab, model routing, local models, Codex and quality gates, RAG for a company knowledge base and implementing AI in a business process.
How the local AI lab works →