Emerging

Edge AI

Deploying AI models on local devices (phones, IoT, embedded systems) rather than cloud servers, enabling low-latency inference and data privacy.

Explained at five levels

Level 1

AI that runs right on your phone or device instead of needing the internet — so it works even when you're offline.

Level 2

Running AI models directly on devices like phones, cars, or cameras instead of sending data to the cloud. It's faster and more private.

Level 3

Deploying AI models on local devices (phones, IoT, embedded systems) rather than cloud servers, enabling low-latency inference and data privacy.

Level 4

On-device inference using quantized, distilled, or purpose-built models optimized for constrained compute environments — trading model size for latency, privacy, and offline capability.

Level 5

Inference at the network edge using hardware-aware model optimization (quantization, pruning, knowledge distillation, neural architecture search) — targeting NPU/DSP accelerators with constraints on power, memory, and thermal envelope.

Definitions are educational summaries. Terminology can vary by source and context.

Sources