THE EDGE AI COMPUTE SPECTRUM
Edge AI spans a 1,000x compute range: microcontrollers with 256 KB SRAM, mobile NPUs with 10-30 TOPS INT8 (7B models in 4-bit), and edge server GPUs with 50-200 TOPS (Jetson Orin, L4). The decision between edge and cloud is driven by latency (sub-50ms needs edge), data privacy, connectivity reliability, and TCO.
A mobile app running a 7B LLM on-device costs $0.0001-0.0005 per query (amortized) versus $0.001-0.005 for cloud API, making on-device 3-10x cheaper at volume. However, the device hardware cost ($20-50 for edge NPU) and deployment complexity shift the balance toward cloud for low-volume scenarios.
MODEL COMPRESSION: QUANTIZATION, PRUNING, AND DISTILLATION
Weight-only INT4 quantization on Llama-3.1-8B reduces memory from 16 GB to 2 GB with 1.5-3 percent MMLU drop. Activation quantization requires SmoothQuant or QuIP for handling outliers, enabling INT8 activation with <1 percent quality loss.
Structured 2:4 sparsity plus INT8 quantization achieves 4-6x compression with 2-4 percent degradation, enabling 7B models on mobile GPUs. Knowledge distillation (Phi-3-mini, TinyLlama) produces purpose-built edge models 40-60 percent smaller with minimal quality loss.
For production edge AI: distilled model + INT4 quantization achieves 5-8x total compression versus the original teacher.
MOBILE GPU VERSUS NPU: HARDWARE ACCELERATION TRADEOFFS
Mobile NPUs (Apple Neural Engine, Qualcomm Hexagon) offer 2-3x better power efficiency than GPUs for attention-based models but support fewer operations. Apple's Apple Intelligence uses hybrid deployment: smaller models run entirely on ANE, larger models use GPU fallback for unsupported operations.
Hybrid execution achieves 80-120 tok/s on Apple A17 Pro for a 3B INT4 model versus 40-60 tok/s on GPU-only. The edge server tier (Jetson Orin NX 16GB, 100 TOPS INT8) runs full 7B INT4 models at 30-50 tok/s on 15W TDP, replacing $10,000-50,000 in cloud GPU inference per year.
HYBRID EDGE-CLOUD ARCHITECTURE FOR AI INFERENCE
The practical pattern: device runs a small model for common cases, falls back to cloud GPU for complex queries. A lightweight router (MobileBERT, TinyBERT) predicts query difficulty. For mobile LLM assistants, 60-80 percent of queries are on-device (3B INT4), 20-40 percent route to cloud (70B). This reduces cloud GPU costs by 60-80 percent.
The cloud gateway must handle smaller request sizes, higher latency tolerance, and device-specific authentication. The cloud GPU pool for edge fallback uses batch sizes of 1-4 because fallback requests arrive asynchronously. During network outages, the on-device model serves all queries with 10-15 percent accuracy degradation.
DEPLOYMENT FRAMEWORKS AND TOOLCHAINS
Apple's CoreML + ANE deployment requires Xcode model conversion (`coremltools`), operator fallback registration, and on-device profiling. Operator coverage is 70-80 percent for modern architectures. Qualcomm's AI Hub provides QNN SDK with ~90 percent LLM operator coverage and HTP backend for INT4 via AWQ calibration.
For cross-platform deployment, Google's MediaPipe and NVIDIA TensorRT are primary options. TensorRT provides maximum performance on NVIDIA edge hardware with explicit operator fusion and kernel auto-tuning, but requires per-device compilation (30-60 minutes per model variant), precluding runtime model adaptation.
B200 AND THE EDGE-CLOUD CONTINUUM
B200 lowers cloud inference costs by 3-5x versus H100. At $0.50-0.80/GPU-hour reserved, cloud cost for a 70B model is $0.001-0.002 per query. For 10M queries/month, this is $10,000-20,000 (B200) versus $30,000-60,000 (H100). This makes cloud-only architectures more viable for mid-scale deployments.
However, network round-trip (20-50 ms wired, 100-500 ms cellular) versus zero network delay for on-device means the user-perceptible difference persists. The long-term architecture is three-tier: on-device NPU for real-time (30 ms), edge server GPU for medium (50-100 ms), and B200 cloud for complex queries (100-200 ms).