Edge AI
Edge AI runs model computation close to the source of data, such as a phone, camera or embedded device. The competence balances model quality with latency, memory, power and runtime compatibility. It includes the local pipeline and update strategy, rather than merely selecting a small model or claiming that local execution solves every privacy concern.
What it is
An edge system performs some or all inference on a device instead of depending entirely on a remote service. Models can be compressed or selected for the hardware, and runtimes map supported operations to CPUs, GPUs or other accelerators. The deployment must also handle input processing, buffering, outputs and intermittent connectivity. Edge AI is broader than small language models: vision, audio and conventional predictive models can all run locally. Local execution changes the placement of computation and data, but does not by itself define the model's architecture, guarantee real-time performance or establish the security of storage, telemetry and updates.
What the work involves
Specify the device and end-to-end response requirement before choosing a model. Measure memory, startup, sustained latency and energy behavior on actual hardware with representative inputs. Check operator support and quantify any quality changes introduced by conversion or compression. Design local error handling, update rollback and the boundary between local and remote work. The deliverable is a tested device pipeline with explicit resource and quality tradeoffs, including behavior during limited connectivity and prolonged use rather than only a successful desktop demonstration.
Illustrative example
Suppose, illustratively, a mobile app labels objects from the camera. An engineer converts a model to an on-device runtime, verifies preprocessing and compares predictions with the original implementation. They measure camera-to-result latency while the app runs continuously, observing whether heat changes performance. Uncertain results can prompt review or optional remote processing under the application's data policy. The deployment decision accounts for the full camera and UI pipeline, not just the isolated model invocation.
Limits and common mistakes
A small weight file can still require substantial activation memory or preprocessing time. Conversion and quantization can alter accuracy, and unsupported operations may fall back to a slower path. Thermal and power constraints affect sustained behavior. Local processing reduces some data transfers but does not secure every stored or logged artifact. Edge deployment differs from model compression: compression is one possible means. Test the final device and workload before extrapolating from accelerator specifications or desktop benchmarks.
Prerequisites
- mediumTransformer Architecture
SLMs are compressed/distilled Transformers — understanding what is being compressed requires understanding the original architecture
- mediumModel Quantization
Edge deployment almost always requires quantization — these skills go hand in hand
Related skills
- → is subcategory of: AI
- → is subcategory of: Machine Learning
- ← is an instance of: NVIDIA Jetson
- ← is an instance of: LiteRT (TensorFlow Lite)
Sources and further reading
- Google: LiteRT
On-device model conversion, runtime and hardware acceleration.
- NVIDIA: Jetson Orin Nano Quick Start
Device-based execution and hardware setup context.
Last updated: 2026-10-10