# edge0-android — Release Build Guide
English | [中文](README_zh.md) | [日本語](README_ja.md) | [Español](README_es.md) | [Français](README_fr.md)
This document is for **developers building or evaluating the Android app from source**. It covers: quick start (engine → app → models → test), measured performance on the reference device, and the technical choices behind the stack.
edge0 is published as a **monorepo** — [`Edge0-AI/edge0`](https://github.com/Edge0-AI/edge0) — whose top level holds the shared engine supply (`vendor.llama.pin` + the scripts-materialized `vendor/llama.cpp`, `patches/llama.cpp/` band-sets) and the platform subprojects (`windows/` = the desktop companion, `android/` = this app). Inference runs on pinned upstream llama.cpp, patched as a replayable patch set, fully on-CPU with ARM-NEON kernels and a demand-paged expert pool.
Third-party attribution: see `NOTICE` in this directory. Model cards: [Edge0/Edge0-8B-A1B-preview](https://huggingface.co/Edge0/Edge0-8B-A1B-preview) · [Edge0/Edge0-35B-A3B-preview](https://huggingface.co/Edge0/Edge0-35B-A3B-preview).
---
## 1. Quick Start
### 1.1 Prerequisites
| component | requirement |
|---|---|
| Device | arm64-v8a, Android 13+ (API 33); Snapdragon 8 Elite class recommended |
| RAM | 8 GB+ runs 8B; 12–16 GB runs 35B (expert paging, see §3.2) |
| Toolchain | JDK 17+, Android SDK 35, **NDK r28** (`28.2.13676358`), CMake ≥ 3.21 + Ninja |
| Python | 3.10+ with `numpy` (model converter, runs from the sibling `windows/tools`) |
| Disk | ≥ 30 GB free for model sources and converted GGUF builds |
The NDK is **not** part of this repository — install it once (Android Studio:
*Android SDK → SDK Tools → NDK (Side by side)*, or via CLI):
```bash
sdkmanager --install "ndk;28.2.13676358"
```
Then point the build at it with `NDK_DIR` (or an exported `ANDROID_NDK_HOME`):
`build_vendor_libs.sh` reads the compiler toolchain **and** the `libomp.so`
runtime it stages (§1.2) from inside that NDK installation; a missing NDK
fails fast with that hint rather than producing a broken library set.
### 1.2 Build the engine
The native libraries come from the pinned upstream tree with this platform's
patch bands replayed into an isolated worktree — the vendor tree is **never
patched in place**:
```bash
git clone https://github.com/Edge0-AI/edge0
cd edge0/android
bash tools/llama/build_vendor_libs.sh
```
The script replays `../patches/llama.cpp/{common,android}` (6 + 14 bands) onto
the pinned llama.cpp tree — the pin lives in `../vendor.llama.pin` (currently
`7ab4ee7`, tag b11100); the tree is **not** a submodule, so on first run it is
cloned from upstream into `../vendor/llama.cpp` (gitignored; set
`EDGE0_LLAMA_URL` to use a mirror) and detached at the pin. Replays happen in a
gitignored consumer worktree, the script
asserts the golden result-tree hash, and produces the four engine shared
libraries (plus the NDK `libomp.so` runtime that `libggml-cpu` needs) with
headers into `build-dl/llama-libs/` (the app's jniLibs staging point).
`--replay` re-applies the bands after patch changes; mismatched tree ⇒ RED, the
build refuses to start. This step is required once before the app build —
the Gradle plugin reads these libraries from `build-dl/llama-libs/`.
### 1.3 Build & install the app
```bash
./gradlew :app:assembleDebug
./gradlew :app:installDebug # or adb install -r app/build/outputs/apk/debug/app-debug.apk
```
### 1.4 Models
GGUF builds are produced locally from the published checkpoints — everything
stays under this directory's `models/` (gitignored):
```bash
huggingface-cli download Edge0/Edge0-8B-A1B-preview --local-dir models/edge0-8b
python ../windows/tools/convert_mlx_to_gguf.py --dir models/edge0-8b
# → models/edge0-8b-gguf/{edge0-8b.gguf, lora_edge0_8b-gguf.gguf, manifest.json}
huggingface-cli download Edge0/Edge0-35B-A3B-preview --local-dir models/edge0-35b
python ../windows/tools/convert_mlx_to_gguf.py --dir models/edge0-35b
bash tools/model/push_models.sh --all # md5-gated staging onto the device
```
The converter needs only python3 + numpy (no MLX/torch runtime — "MLX" denotes
the on-disk checkpoint layout). It runs the r3 repack with numeric parity gates
and emits a sha256 manifest; a correct conversion reproduces the baseline
checksums listed in `push_models.sh` byte-for-byte. Models can also be copied
into `files/models/` with the in-app picker.
### 1.5 Run
Launch **Edge0 Chat**. The title bar shows the active model (8B / 35B), the
top-right button switches. The composer is send-only; temperature, thinking and
the system prompt live in the drawer settings. Each reply carries an inline
metrics line: `tokens · TTFT · prefill t/s · decode t/s · RSS`.
### 1.6 Test it
Instrumented regression (needs both models staged — reinstalling the test APK
wipes app data, so restage right before the run):
```bash
./gradlew :app:installDebugAndroidTest
bash tools/model/push_models.sh --all
adb shell am instrument -w -e class dev.edge0.runtime.app.LlamaRuntimeTest \
dev.edge0.runtime.app.test/androidx.test.runner.AndroidJUnitRunner
# expected: OK (8 tests), ~7 min on the reference device
```
Coverage: 8B/35B smoke, 35B↔8B in-process switching, prefix-reuse fidelity,
thinking on/off × system-prompt quadrants (8B gate + 35B off leak probe), and
identity retention across multi-turn rendering. Host-side logic tests:
`./gradlew :app:testDebugUnitTest` (26 tests, no device needed).
---
## 2. Performance
Reference device: **Lenovo TB322FC (Snapdragon 8 Elite, 16 GB RAM)**, shipping
config, sustained windows (first segments = boost clocks, tail = thermal steady
state — both reported rather than cherry-picked peaks).
| Model | TTFT (warm turn) | Decode | Prefill | Session RSS |
|---|---:|---:|---:|---:|
| **8B** (Q8-class GGUF + LoRA) | ≈ 1.4 s | 29–32 → ~10 t/s over a 480 s window | ~100 t/s | ≈ 250 MB |
| **35B** (mixed int8 MoE, demand-paged) | ≈ 1.1 s warm (≈ 10–15 s on the very first turn of a new topic — cold expert pool, flash-bound) | 6–9 t/s in-app (9.46 t/s sustained CLI) | ~1 s per incremental turn | pool budget 2–6 GB; resident ≪ file size |
Notes: "first turn of a new topic" pays the cold-pool tax — e.g. a 55-token
question measured prefill 12.8 s with 25956 expert loads, 73 % of wall time in
flash I/O wait; that is demand paging working as designed, not a regression.
Turn two onwards is warm: KV prefix reuse + keepwarm refill bring TTFT to ~1 s.
Decode decay along a long window is DVFS/thermal behavior on this SoC.
---
## 3. Technical Details
### 3.1 Architecture
```mermaid
graph TD
subgraph App ["Kotlin / Jetpack Compose"]
UI[ChatScreen · dark · send-only composer] --> VM[ChatViewModel]
VM --> RT[LlamaRuntime
coroutines + Flow events]
VM --> DB[(Room · threads & messages)]
ST[SettingsStore] --> VM
end
subgraph Native ["C JNI shell (llama_chat.c)"]
SHELL[generate loop · template-aware thinking control
segment-wise history render · UTF-8-safe streaming]
end
subgraph Engine ["patched llama.cpp @ b11100 · arm64 CPU-only"]
LLIB[libllama.so]
GCPU[libggml-cpu.so
NEON kernels + moe_pool]
end
RT -->|JNI| SHELL --> LLIB --> GCPU
GCPU -->|demand-paged expert IO| MODELS[GGUF on flash]
```
### 3.2 Why a 21.7 GB MoE fits on a phone — `moe_pool`
The 35B model activates only a moving subset of its 256-per-layer experts per
token, so the shipped design pages experts **on demand** from flash instead of
resident memory (`ggml/src/ggml-cpu/moe_pool.c`, developed as the 14-patch
android band): copy-in private frames, an equal-slot state machine, background
IO staging queues tuned to measured UFS throughput, pin/blob/trim controls, a
turn-end keepwarm refill within a byte budget, and full cross-model reset
enabling 8B↔35B switching inside one process. With the pool disabled the engine
resolves expert rows exactly like upstream (NULL resolver ⇒ zero perturbation,
verified by symbol-set diffing). This is the same mechanism family the desktop
project implements over NVMe + Vulkan; here it is CPU/NEON by design — GPU
backends stay out of the shipping config for deterministic numerics and a
single memory model.
### 3.3 Repo map
```
app/ Android app: Compose UI (src/main/java), JNI shell (src/main/cpp),
instrumented + unit tests (src/androidTest, src/test)
tools/llama/ build_vendor_libs.sh — engine rebuild from the pinned tree + bands
tools/model/ push_models.sh — md5-gated model staging to devices
../vendor/llama.cpp/ materialized by the build scripts from vendor.llama.pin (gitignored; never patched in place)
../patches/llama.cpp/ common(6) + android(14) hook-point bands + ledger README
../windows/ desktop companion — hosts the MLX→GGUF converter used in §1.4
```