Emerging
Edge AI
Deploying AI models on local devices (phones, IoT, embedded systems) rather than cloud servers, enabling low-latency inference and data privacy.
Explained at five levels
Level 1
AI that runs right on your phone or device instead of needing the internet — so it works even when you're offline.
Level 2
Running AI models directly on devices like phones, cars, or cameras instead of sending data to the cloud. It's faster and more private.
Level 3
Deploying AI models on local devices (phones, IoT, embedded systems) rather than cloud servers, enabling low-latency inference and data privacy.
Level 4
On-device inference using quantized, distilled, or purpose-built models optimized for constrained compute environments — trading model size for latency, privacy, and offline capability.
Level 5
Inference at the network edge using hardware-aware model optimization (quantization, pruning, knowledge distillation, neural architecture search) — targeting NPU/DSP accelerators with constraints on power, memory, and thermal envelope.
Definitions are educational summaries. Terminology can vary by source and context.
Sources
- NIST Trustworthy and Responsible AI Resource Center: Glossary — National Institute of Standards and Technology. Accessed 2026-07-20.