Technologies

Multimodal sensor fusion

The reason to fuse sensors is not more data. It is that they fail differently — and the combination is available in conditions where any single channel is not.

Fusion Levels
Early · Feature · Late
Default Start
Late fusion
Tested On
Degraded conditions
Usual Shape
Learned front end, classical state

Why Fuse At All

Every sensor is blind to something different

Fusion is worth its considerable complexity when the combination is available in conditions where any single channel is not. That is a specific claim, and it is the one to test before committing to a multi-sensor architecture — because a second sensor that fails in the same conditions as the first has added cost, weight, power and calibration burden while buying almost nothing.

Camera

Easy to interpret and useless in fog, at night, and pointed into the sun.

Radar

Works through weather and darkness, and produces something no human finds intuitive.

Thermal

Sees heat rather than detail, which is exactly right for some questions and useless for others.

Lidar

Gives geometry directly, and struggles with rain and reflective surfaces.

The same jetty and tank farm shown three ways: daylight optical, thermal infrared, and synthetic aperture radar
What We Combine

The same scene, through sensors that disagree about what is interesting.

Vision with thermal and process data on a production line. Radar with camera, where weather and darkness decide availability. Electro-optical with thermal infrared. Lidar and geometry. SAR with optical, where radar provides all-weather availability and optical provides interpretability. Radio frequency and other passive sources.

And the non-sensor context that is frequently the cheapest accuracy improvement available — usually cheaper than adding another sensor, and almost always overlooked in favour of one.

Sensor Combinations
Vision + thermal Radar + camera EO + thermal infrared Lidar and geometry SAR + optical RF and passive sources
Non-Sensor Context
Process data from the machine itself
Plant and maintenance records
Weather
Terrain models
Known traffic and vessel data
Where to Fuse

Early, feature-level, or late

For safety-relevant and mission-relevant systems we usually start from late fusion and move earlier only where the accuracy gain is measured and the alignment can be guaranteed.

01

Early fusion

Combines raw or near-raw data before the model. Best accuracy when the sensors are well aligned in time and space, and unforgiving when they are not.

02

Feature-level fusion

Combines learned representations. The common middle ground, and it tolerates moderate misalignment.

03

Late fusion

Combines decisions from independent per-sensor chains. Less accurate at the top end and much better at everything else: it degrades gracefully when a sensor drops out, each chain can be tested and certified on its own, and you can explain to an operator or an auditor why the system concluded what it did.

Method

Classical estimation and learned fusion, usually in the same system.

Kalman filters, factor graphs and Bayesian estimators are still the right answer for tracking, state estimation, and anything where the dynamics are known and the behaviour has to be predictable and explainable. Learned fusion earns its place where the relationship between modalities is complicated and hard to write down, such as matching a radar signature against an optical detection.

Most production systems we work on are hybrids, with a learned perception front end per sensor and a classical estimator maintaining the fused state. The choice is an engineering one, and it should follow from what has to be certified, explained and debugged at three in the morning.

Classical Estimation
Kalman filters Factor graphs Bayesian estimators
Where Learned Fusion Earns Its Place
Relationships between modalities that are hard to write down
Matching a radar signature against an optical detection
Perception front ends, per sensor
With a classical estimator still holding the fused state
The Hard Parts

The hard parts are not the model

None of these are solved by a better architecture, and every one of them looks like model degradation when it goes wrong.

01

Time synchronisation

Sensors sample at different rates with different latencies. An unmodelled offset of a few tens of milliseconds becomes a position error that scales with relative velocity, and nothing downstream repairs it.

02

Calibration drift

Intrinsic and extrinsic parameters drift with temperature, vibration and time. A system that assumes a factory calibration holds for five years on a vehicle, a mast or a machine frame will slowly get worse in a way that looks like model degradation and is not.

03

Coordinate frames and geometry

Different sensors, different projections, different mounting points — and one world model that all of them have to agree on.

04

Calibrated uncertainty

Fusion only works if each channel reports a calibrated confidence, rather than a score that merely happens to lie between zero and one.

05

Models are badly calibrated by default

Most deep models are, and fixing that is often the single largest accuracy improvement available anywhere in a fusion stack.

06

Missing inputs

A sensor will fail, be occluded, be blinded or be jammed. The system needs a behaviour defined in advance for that case, rather than whatever the model happens to do when it receives zeros.

Evaluation

Average-case benchmarks hide exactly the behaviour you are buying fusion for.

The tests that matter are the ones where the sensors disagree, and the ones where a channel is degraded or gone: fog, darkness, glare, heavy rain, occlusion, interference, a failed camera, a blocked radome.

01

Test where the sensors disagree

Disagreement is the whole reason the second sensor is there. A benchmark that never produces it is not testing the fusion.

02

Test with a channel degraded or gone

Fog, darkness, glare, heavy rain, occlusion, interference, a failed camera, a blocked radome — each as its own case, not averaged away.

03

Report the degraded case separately

An aggregate number lets a strong clean-weather result hide a collapse. The degraded-mode figure is the one the system was bought for.

A fusion system that outperforms on clean data and collapses when one input drops has not solved the problem it was bought for.

Where This Applies

Anywhere one sensor runs out before the requirement does

Fusion is not a remote-sensing technique that occasionally escapes into industry. It applies wherever a single channel stops being sufficient — which is most places, once the conditions get difficult enough.

Manufacturing and industrial inspection

Vision, thermal and process data together catching what none of them catches alone — surface defects, thermal anomalies and a machine that was already drifting.

Robotics and autonomous systems

Where a machine has to keep a usable world model while a sensor is occluded, blinded or simply pointed the wrong way.

Energy and critical infrastructure

Assets that have to stay observed through weather, darkness and intermittent access, often with nobody on site.

Aerospace, defense and remote sensing

Surveillance and earth observation, where the conditions you care about are precisely the ones a single channel cannot hold, and where space-based and airborne sources have to be reconciled.

Architecture assessment

How a first engagement looks

If the honest answer is that one well-chosen sensor does the job, that is a cheaper system and we will tell you.

1

Is the sensor set genuinely complementary?

Whether the channels actually fail differently in the conditions you care about — or whether the second one fails alongside the first.

2

Where should fusion happen?

Early, feature-level or late, decided against what has to be certified, explained and debugged rather than against a benchmark.

3

What keeps it accurate in the field?

The calibration and synchronisation infrastructure the system needs to stay correct once temperature, vibration and time get to it.

4

What is the compute budget, really?

Several perception chains at real-time rates plus the fused estimator, costed in watts and euros before a device class is chosen.

Book a Strategy & Architecture Review

Tell us which conditions your system has to survive.

We will tell you whether your sensor set is genuinely complementary in those conditions, where the fusion should happen, and what the compute budget will actually be.

Book a Strategy & Architecture Review
nAIxt Technologies GmbH
Am Forst 2
82166 Gräfelfing, Germany
+49 89 54196515
info@naixt-technologies.de