A fleet too weak for the model you needed
The device was chosen against the model you had at the time, and the model that eventually worked did not fit on it.
The hardware question is the last one to answer, not the first. Establish what the system has to do and under what constraints, find out whether that is achievable on edge hardware at all, then pick the class of hardware that fits — and only then argue about part numbers.
That is how organisations end up with a fleet of modules too weak for the model they eventually needed, or a unit cost that makes the business case collapse at the hundredth installation. We help clients through the opposite sequence — and about a third of the time the useful outcome is finding out that the workload does not belong on a device at all.
The device was chosen against the model you had at the time, and the model that eventually worked did not fit on it.
The difference between a forty-euro module and a two-thousand-euro one is a rounding error in a pilot and the entire business case across a thousand installations.
High concurrency, long context, frequent retraining, or an accuracy requirement only a frontier model meets. These belong on a server, and finding that out early is a result.
About a third of the time, the useful outcome of an edge assessment is finding out that the workload does not belong on a device.
Capability Today
The capability ceiling has moved a long way in three years, and much of the received wisdom is out of date. A rough picture of what is comfortable now — and of what still is not.
Object detection, segmentation, tracking, OCR, anomaly detection and multi-camera pipelines run at real-time frame rates within a few watts to a few tens of watts. If your problem is a vision problem, it almost certainly runs on a device.
Keyword spotting, acoustic anomaly detection, vibration analysis and predictive maintenance fit on microcontroller-class hardware in the milliwatt to single-watt range — often on the sensor itself.
A 7 to 13 billion parameter model gives usable throughput on mid-range embedded hardware: enough for structured extraction, classification, short instructions and constrained dialogue. Top-end modules hold much larger models, but throughput rather than capacity is usually the limit.
High concurrency, long context windows, anything needing frequent retraining on the device, and workloads where only a frontier model meets the accuracy requirement. These are signals that the work belongs on a server.
Answer these before anyone opens a datasheet. Several of them settle the question on their own.
How fast does the answer have to arrive, at the percentile that matters rather than on average? A control loop closing in under ten milliseconds settles the question by itself. A quality report within two seconds does not.
Can the data leave the site? Export control, classified work, GDPR, works council agreements and simple unwillingness to put process data or design IP on third-party infrastructure are all hard constraints — and they are about the building rather than the machine.
What happens when the network is gone? Intermittent connectivity is normal on vessels, in mines, at substations and in vehicles. A system that degrades to nothing without a link is usually the wrong design.
What is the power, thermal and space envelope? More edge designs fail on heat and enclosure than on compute.
What does this cost at fleet scale? A module price that is a rounding error in a pilot is the whole business case across a thousand installations.
How long does the product have to live, and does it need certification? A five-year service commitment, functional safety requirements or a regulated environment constrain the hardware and the update mechanism far more than performance does.
The expensive confusion sits between the first two. Specifying a fielded device for work that belongs in a rack inflates cost across the whole fleet; specifying a desktop box as production infrastructure produces a system that cannot be maintained or certified. Both mistakes are common, and both are avoidable in a week of analysis.
An embedded module or an accelerator inside the equipment. Chosen when latency, power, size, connectivity or fleet economics demand it.
A local server, industrial PC or workstation in the plant, on the vessel, in the substation. This is where most requirements described as “we need an edge LLM” actually belong, and it is the cheapest way to satisfy a data-residency constraint.
When scale, multi-tenancy or model size make anything smaller uneconomic.
Jetson modules are deployment hardware: industrial temperature ranges, long production lifecycles, ruggedisation options, and a supply commitment you can build a product around.
DGX Spark is development hardware: a desktop machine with a large pool of unified memory, built so an engineer can run and fine-tune a sizeable model locally without sending anything to a cloud provider. For an organisation that cannot share data with a cloud, that is genuinely valuable. It is not a production inference server, it has no industrial environmental rating, and its memory bandwidth becomes the limit under real concurrency.
Usually yes, and rather more cheaply than people expect — provided you are honest about which layer you are building. The pragmatic answer for most industrial companies is to buy the module and the base platform, integrate with a partner if the enclosure is demanding, and own the model and the operations outright.
Custom accelerators and ASICs pay back only at very high volume with a fixed, stable function, and the development cost and schedule risk are substantial.
Off-the-shelf carriers and industrial computers are the right starting point for pilots and volumes in the hundreds. A custom carrier becomes worthwhile when the mechanical envelope, connector set, sensor interfaces or unit cost at volume justify the engineering and the certification work that follows.
This is where your data, your process knowledge and your sensor set produce something a supplier cannot sell you. It is also where the long-term operational work lives: updating fielded models, rolling back a bad release, detecting drift without shipping data home, and keeping the whole thing reproducible.
Compute is rarely what kills an edge design. Power, thermals, enclosure, update paths and supply lifecycle are — and most of them are decided long before anyone benchmarks a model.
Quantisation is a trade rather than a free optimisation, so the only useful version of that question is measured on your task, against the throughput it buys you.
We sell no hardware and hold no reseller agreements, so the recommendation is the one we would make with our own capital.
Where This Applies
Inline decisions at line rate, in an enclosure on the machine, with a product lifecycle measured in years rather than release cycles.
Export-controlled and classified programmes, where the boundary the data may not cross is the first architectural constraint.
Substations, vessels and remote sites where intermittent connectivity is normal and a system that stops without a link is the wrong design.
Onboard and ground-segment processing traded against downlink bandwidth, which is usually the scarcest resource in the system.
Deployment architecture review
Two to three weeks. You bring the workload, the constraints and the target unit economics.
Real input rates, real concurrency, the real latency requirement — and the percentile that matters rather than the average.
Whether the requirement is reachable on the machine, on the site, or only in a data centre — including the case where the answer is that it should not be on a device.
A hardware shortlist with the accuracy and throughput consequences stated plainly, including what running smaller or in lower precision costs on your task.
A build-or-buy split across silicon, board and model layers, with the update, rollback and drift-detection design that goes with it.
The perception workloads that most often end up in an enclosure on the machine rather than in a rack.
Several perception chains at real-time rates plus a fused estimator — the workload where device-class selection stops being a detail.
What changes when the model that takes actions has to run inside your perimeter rather than behind an API.
Where the training and evaluation data comes from when the real data is the thing that cannot leave.
The wider architecture around the device decision — where inference runs, how it is operated, and what it costs to keep running.
The build-or-buy split across silicon, board and model layers, decided against requirements rather than vendor slides.
You bring the workload, the constraints and the target unit economics. You get a tier recommendation, a hardware shortlist with the consequences stated plainly, a build-or-buy split, and the operations design that goes with it.
Book a Strategy & Architecture Review