Technologies

Edge and on-premise AI

The hardware question is the last one to answer, not the first. Establish what the system has to do and under what constraints, find out whether that is achievable on edge hardware at all, then pick the class of hardware that fits — and only then argue about part numbers.

Tiers
Machine · Site · Data centre
Decided By
Six questions, not a datasheet
Hardware Sales
None, no reseller ties
First Step
Two to three weeks
The Wrong Order

Most edge AI conversations start with a device and work backwards.

That is how organisations end up with a fleet of modules too weak for the model they eventually needed, or a unit cost that makes the business case collapse at the hundredth installation. We help clients through the opposite sequence — and about a third of the time the useful outcome is finding out that the workload does not belong on a device at all.

01

A fleet too weak for the model you needed

The device was chosen against the model you had at the time, and the model that eventually worked did not fit on it.

02

A unit cost that collapses at scale

The difference between a forty-euro module and a two-thousand-euro one is a rounding error in a pilot and the entire business case across a thousand installations.

03

A workload that never belonged on a device

High concurrency, long context, frequent retraining, or an accuracy requirement only a frontier model meets. These belong on a server, and finding that out early is a result.

About a third of the time, the useful outcome of an edge assessment is finding out that the workload does not belong on a device.

Capability Today

What edge AI can actually do today

The capability ceiling has moved a long way in three years, and much of the received wisdom is out of date. A rough picture of what is comfortable now — and of what still is not.

Classical perception is solved territory

Object detection, segmentation, tracking, OCR, anomaly detection and multi-camera pipelines run at real-time frame rates within a few watts to a few tens of watts. If your problem is a vision problem, it almost certainly runs on a device.

Audio and time-series run even lower

Keyword spotting, acoustic anomaly detection, vibration analysis and predictive maintenance fit on microcontroller-class hardware in the milliwatt to single-watt range — often on the sensor itself.

Small language and vision-language models now run on device

A 7 to 13 billion parameter model gives usable throughput on mid-range embedded hardware: enough for structured extraction, classification, short instructions and constrained dialogue. Top-end modules hold much larger models, but throughput rather than capacity is usually the limit.

What remains awkward

High concurrency, long context windows, anything needing frequent retraining on the device, and workloads where only a frontier model meets the accuracy requirement. These are signals that the work belongs on a server.

Deriving the Requirement

Six questions decide almost every edge architecture

Answer these before anyone opens a datasheet. Several of them settle the question on their own.

01

How fast does the answer have to arrive, at the percentile that matters rather than on average? A control loop closing in under ten milliseconds settles the question by itself. A quality report within two seconds does not.

02

Can the data leave the site? Export control, classified work, GDPR, works council agreements and simple unwillingness to put process data or design IP on third-party infrastructure are all hard constraints — and they are about the building rather than the machine.

03

What happens when the network is gone? Intermittent connectivity is normal on vessels, in mines, at substations and in vehicles. A system that degrades to nothing without a link is usually the wrong design.

04

What is the power, thermal and space envelope? More edge designs fail on heat and enclosure than on compute.

05

What does this cost at fleet scale? A module price that is a rounding error in a pilot is the whole business case across a thousand installations.

06

How long does the product have to live, and does it need certification? A five-year service commitment, functional safety requirements or a regulated environment constrain the hardware and the update mechanism far more than performance does.

Three Tiers

On the machine, on the site, or in your own data centre

The expensive confusion sits between the first two. Specifying a fielded device for work that belongs in a rack inflates cost across the whole fleet; specifying a desktop box as production infrastructure produces a system that cannot be maintained or certified. Both mistakes are common, and both are avoidable in a week of analysis.

01

On the machine

An embedded module or an accelerator inside the equipment. Chosen when latency, power, size, connectivity or fleet economics demand it.

02

On the site

A local server, industrial PC or workstation in the plant, on the vessel, in the substation. This is where most requirements described as “we need an edge LLM” actually belong, and it is the cheapest way to satisfy a data-residency constraint.

03

In your own data centre

When scale, multi-tenancy or model size make anything smaller uneconomic.

Five embedded compute modules of ascending size laid out on an anti-static mat beside a steel rule
Deployment vs Development

Deployment hardware and development hardware are not the same class, and the current NVIDIA line-up illustrates it well.

Jetson modules are deployment hardware: industrial temperature ranges, long production lifecycles, ruggedisation options, and a supply commitment you can build a product around.

DGX Spark is development hardware: a desktop machine with a large pool of unified memory, built so an engineer can run and fine-tune a sizeable model locally without sending anything to a cloud provider. For an organisation that cannot share data with a cloud, that is genuinely valuable. It is not a production inference server, it has no industrial environmental rating, and its memory bandwidth becomes the limit under real concurrency.

Deployment Hardware Brings
Industrial temperature range Long production lifecycle Ruggedisation options A supply commitment
A Development Box Does Not Bring
An industrial environmental rating
A long-term embedded lifecycle
Headroom under real concurrency
Anything you can certify a product against
But it does keep your data off a cloud provider
Build or Buy

Can you build your own edge AI system?

Usually yes, and rather more cheaply than people expect — provided you are honest about which layer you are building. The pragmatic answer for most industrial companies is to buy the module and the base platform, integrate with a partner if the enclosure is demanding, and own the model and the operations outright.

01

Silicon — almost nobody should build

Custom accelerators and ASICs pay back only at very high volume with a fixed, stable function, and the development cost and schedule risk are substantial.

02

Board and system — it depends on volume

Off-the-shelf carriers and industrial computers are the right starting point for pilots and volumes in the hundreds. A custom carrier becomes worthwhile when the mechanical envelope, connector set, sensor interfaces or unit cost at volume justify the engineering and the certification work that follows.

03

Model and application — you should build

This is where your data, your process knowledge and your sensor set produce something a supplier cannot sell you. It is also where the long-term operational work lives: updating fielded models, rolling back a bad release, detecting drift without shipping data home, and keeping the whole thing reproducible.

How We Help

The analysis that makes the hardware decision the easy part.

Compute is rarely what kills an edge design. Power, thermals, enclosure, update paths and supply lifecycle are — and most of them are decided long before anyone benchmarks a model.

Quantisation is a trade rather than a free optimisation, so the only useful version of that question is measured on your task, against the throughput it buys you.

Characterised First
Real input rates Real concurrency The latency percentile that matters
What the Engagement Covers
Workload characterisation against what each tier can achieve
The accuracy cost of running smaller or in lower precision
Power, thermal and enclosure budgets, sized early
Model update and rollback across a fleet
Drift detection without a data link home, and air-gapped operation
Supply and lifecycle risk across module generations

We sell no hardware and hold no reseller agreements, so the recommendation is the one we would make with our own capital.

Where This Applies

Programmes where the computation has to sit somewhere specific

Manufacturing and industrial inspection

Inline decisions at line rate, in an enclosure on the machine, with a product lifecycle measured in years rather than release cycles.

Aerospace and Defense

Export-controlled and classified programmes, where the boundary the data may not cross is the first architectural constraint.

Energy and critical infrastructure

Substations, vessels and remote sites where intermittent connectivity is normal and a system that stops without a link is the wrong design.

Remote Sensing

Onboard and ground-segment processing traded against downlink bandwidth, which is usually the scarcest resource in the system.

Deployment architecture review

How a first engagement looks

Two to three weeks. You bring the workload, the constraints and the target unit economics.

1

Characterise the workload

Real input rates, real concurrency, the real latency requirement — and the percentile that matters rather than the average.

2

Establish what each tier can achieve

Whether the requirement is reachable on the machine, on the site, or only in a data centre — including the case where the answer is that it should not be on a device.

3

Shortlist hardware, with the consequences stated

A hardware shortlist with the accuracy and throughput consequences stated plainly, including what running smaller or in lower precision costs on your task.

4

Split build from buy, and design the operations

A build-or-buy split across silicon, board and model layers, with the update, rollback and drift-detection design that goes with it.

Book a Strategy & Architecture Review

Bring us the workload before you bring us a part number.

You bring the workload, the constraints and the target unit economics. You get a tier recommendation, a hardware shortlist with the consequences stated plainly, a build-or-buy split, and the operations design that goes with it.

Book a Strategy & Architecture Review
nAIxt Technologies GmbH
Am Forst 2
82166 Gräfelfing, Germany
+49 89 54196515
info@naixt-technologies.de