Running models at the edge when the network cannot be trusted
Edge inference is usually framed as a latency problem. In the field it is a power, memory and update problem — and the model is the easiest part to change.
The budget comes first
Before anyone picks an architecture we write down three numbers: peak RAM available to inference, the energy per inference the power budget allows, and the worst-case latency the product can tolerate. Those numbers eliminate most model families immediately, which is a good thing — it turns an open search into a short list.
Quantisation and pruning then get evaluated against measured accuracy on device data, not the benchmark set. A model that loses two points of accuracy and half its memory footprint is almost always the right trade.
Three numbers — peak RAM, energy per inference, worst-case latency — eliminate most model families before the search begins.
Offline means deciding without a second opinion
When the network is unavailable, the device cannot escalate an uncertain result to a bigger model. So the interesting design work is in the fallback: what the device does when confidence is low. Usually it records the input, acts conservatively, and flags the event for review once a link returns.
That decision belongs to the product, not the model. We surface confidence as a first-class output and let the application choose the threshold.
Updates are the long-term requirement
A model that cannot be replaced in the field is a model frozen at launch quality. Signed, resumable, A/B-partitioned updates go into the first firmware image — retrofitting them across a deployed fleet is the most expensive work in this category.




