Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Local and edge inference: design for the device and the trust boundary

By Anup Rai5 min readReviewed September 2026

On-device inference runs the model on the user's device. Edge inference runs near the data source, such as a site gateway or local appliance. Self-hosted inference means the organization operates the serving system; it can still run in a cloud data center. These locations offer different control, connectivity, capacity, and maintenance tradeoffs.

A Learnastra design answer starts with a concrete constraint—offline operation, data handling, response latency, or workload economics—then selects a model and runtime that meet it.

The Runtime Stack

Layer or role Examples What to evaluate
Desktop workflow and local service Ollama, LM Studio Artifact management, API behavior, concurrency, authentication and operational fit
Portable inference engine llama.cpp Supported architectures, GGUF variants, CPU/GPU backends and server capabilities
Apple-silicon array framework and ecosystem MLX and MLX-based model tooling Model support, conversion, unified-memory use and measured performance
Shared serving runtime vLLM, SGLang, TensorRT-LLM Scheduling, supported hardware, quantization, parallelism and production controls
Embedded/application deployment ExecuTorch, Core ML, ONNX Runtime, MLC LLM Conversion, operator support, accelerator backend, app integration and updates

These roles overlap. A GUI can expose a service, and an inference engine may include a server. Avoid calling one tool “the production answer” without defining the workload. Primary references: llama.cpp, MLX, ExecuTorch, Core ML, ONNX Runtime generation, MLC LLM.

A local prototype needs a production evaluation

Ollama documents parallel request processing and queue/concurrency controls. It is incorrect to claim it always serializes all requests. More parallel work also needs more memory, and overload can still create queueing. See the current Ollama FAQ.

LM Studio documents configurable API-token authentication in supported versions. Its default local API behavior is different from a server with authentication explicitly enabled. See LM Studio authentication. Conversely, the Ollama local API does not require authentication by default; network exposure needs deliberate controls.

A team moving to a shared endpoint should test concurrent arrivals, tail latency, cancellations, access control, quotas, memory pressure, monitoring, updates, and recovery. A dedicated serving engine may offer better capacity or control, but no universal 16–20× speedup follows from the product names. Equal model, precision, hardware, input/output lengths, and latency targets are prerequisites for a meaningful comparison.

As checked in September 2026, Hugging Face TGI is in maintenance mode. Existing deployments need their own migration assessment; new designs should consider actively developed alternatives and required compatibility.

When Local Beats Cloud (and When It Does Not)

Requirement Why local/edge may help Remaining limitation
Offline operation No inference round trip is required Updates, first installation, and external tools may still need connectivity
Restricted data movement Processing can stay inside the required boundary Logs, telemetry, backups, retrieval and fallback can still transmit data
Interactive latency Removes one network path Local compute, loading and thermal limits can dominate
Predictable high volume Can use owned or reserved capacity efficiently Staff, idle redundancy, power, upgrades and quality still cost money
Highly variable or low volume Cloud API can avoid idle dedicated capacity Provider limits, network dependence and data terms must fit

Do not infer legal compliance from “runs locally,” or infer adequate privacy from a provider's retention label. Establish the actual data flow and contractual requirements. A hybrid architecture is useful only when data can legitimately cross the fallback boundary.

Architecture / visual model
flowchart TD A[Request and data classification] --> B{Local capability meets task?} B -->|Yes| C[Run approved local model] B -->|No| D{Cloud transfer permitted and connected?} D -->|Yes| E[Send permitted minimum context] D -->|No| F[Explain limitation or defer] C --> G[Validate result] E --> G
Read diagram source
flowchart TD
  A[Request and data classification] --> B{Local capability meets task?}
  B -->|Yes| C[Run approved local model]
  B -->|No| D{Cloud transfer permitted and connected?}
  D -->|Yes| E[Send permitted minimum context]
  D -->|No| F[Explain limitation or defer]
  C --> G[Validate result]
  E --> G

Quantization for Local Serving

Start with weight and cache arithmetic, not a device/model-size slogan. For ideal four-bit weights, the payload alone is:

Stored parameters Ideal payload
3 billion 1.5 GB
8 billion 4 GB
32 billion 16 GB
70 billion 35 GB
200 billion 100 GB

These are decimal GB before scales, unquantized tensors, cache, temporary buffers, and the operating system. A claimed 200B four-bit model fitting wholly into 48 GB would require further assumptions such as offloading or a different representation; active MoE parameters do not replace total stored weights in this calculation.

GGUF variants such as Q4_K_M, Q5_K_M, and Q8_0 have different layouts and mixed tensor choices. They do not correspond to fixed universal accuracy losses. Choose a supported conversion and compare the actual workload; the largest model or highest precision that fits is not necessarily the best product choice if latency or battery use is unacceptable.

Hardware

  1. Discrete GPU: account for VRAM, memory bandwidth, supported kernels and host/device transfer. Do not assume a fixed market-wide VRAM ceiling.
  2. Unified memory: CPU and GPU can share a pool, but the application, operating system, and other processes share that capacity too. Total installed RAM is not all available to model state.
  3. NPU: peak TOPS does not establish LLM speed. Check supported operators, numerical formats, model conversion, memory traffic and fallback execution.
  4. Phone or battery-powered device: measure sustained thermal behavior, energy, app memory limits, background lifecycle and model download size on representative devices.
  5. Edge appliance: include physical access, signed updates, storage failure, monitoring without sensitive payloads and replacement procedures.

For an illustrative app with a measured 6 GiB usable memory budget, 3.8 GiB weights and metadata, 1.1 GiB cache, and 0.7 GiB runtime buffers leave only 0.4 GiB headroom. A longer context or a concurrent request can exhaust it. Measure the peak instead of adding an arbitrary fixed overhead percentage.

Prototype to Production

  1. Select a small baseline model that can pass the task and important failure cases.
  2. Run the exact converted artifact on the lowest supported device class.
  3. Test cold start, sustained load, memory pressure, cancellations and offline behavior.
  4. Pin model, tokenizer, template, runtime and conversion versions; verify artifact integrity.
  5. Add access controls, bounded queues, quotas and monitoring for a shared service.
  6. Ship staged updates with rollback and a truthful unsupported-device or offline fallback.
  7. Recheck data handling when adding cloud fallback, remote logging, retrieval or tools.

An API-compatible endpoint can reduce integration effort, but tool parsing, streaming events, defaults and error behavior can still differ. Test the application contract during runtime migration.

Interview practice

  1. Does self-hosting mean on-device? No. An organization can self-host in a remote cloud data center.
  2. Why not pick hardware from TOPS? Operator support, bandwidth, memory, precision, software and sustained conditions determine useful performance.
  3. Must an Ollama prototype be discarded? Evaluate production requirements and load. Replace components when measured capacity or controls justify it.
  4. Does local inference guarantee privacy? No. Audit every path that handles data, including logs, updates and fallback.
  5. When is cloud fallback inappropriate? When data movement is forbidden, connectivity is absent, or the provider contract cannot meet the task's requirements.
  6. How do you compare cost? Match quality and workload, then include hardware/hosting, idle capacity, staff, power, maintenance, failure handling and refresh work.

Recall card and closing

Location → data boundary → device budget → runtime → operations. Close with the supported device/workload envelope, the behavior when it is exceeded, and the update and rollback plan.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Diffusion language models: parallel refinement with a measurable contract
NEXT LESSONPrompt Engineering Fundamentals →

Explore the diagram