EdgeLM is a shared, on-device AI runtime for Android. It runs chat, vision, speech and search directly on your phone — privately, offline, and shared by every app that wants to use AI. Download a model once, and any EdgeLM-powered app can use it, instead of each app bundling and running its own.
Think of it as a system service for on-device intelligence: one place that holds the models, so your apps get AI without sending your data to the cloud.
| Chat | streaming, multi-turn | EdgeLM.chat(...) |
| Vision | describe a photo, read a screenshot | EdgeLM.caption(...) |
| Speech → text | live captions and transcription | EdgeLM.liveTranscribe() |
| Voice Q&A | speak a question, get an answer | EdgeLM.askByVoice(...) |
| Search | on-device embeddings + local vector index | /v1/embeddings, /v1/edge/rag |
| Agents & tools | app-registered webhooks, brokered through the firewall | /v1/edge/agent |
Each is a separate capability an app must declare and the user can revoke, and each loads its model on demand and frees it when idle.
- Private by design. Your prompts and the AI's responses never leave your device. No cloud, no account, no tracking, no ads.
- Works offline. Once a model is downloaded, EdgeLM runs with no internet at all — the same on a plane, in a tunnel, or on airplane mode.
- Shared and efficient. EdgeLM loads a model once and serves every app from that single copy, so a second app adds almost no extra memory.
- No cloud cost, no round-trips. Nothing to pay per request, and no waiting on a network.
- Install EdgeLM Runtime and follow the quick welcome.
- Pick an AI. EdgeLM recommends the right model for your phone — one tap to download. Prefer to choose yourself? Switch to Advanced for the full catalog.
- Use it. Try it right away in the built-in playground, and any EdgeLM-powered app on your phone can now use on-device AI.
A small, always-available notification shows when the runtime is active, lets you free up memory on demand, and the runtime automatically releases memory when idle.
A curated catalog spans tiny-and-fast to more capable, each shown with a plain-language description, size, and what it's best for — with a warning if your phone may not have enough memory:
| In the app | Model | Good for |
|---|---|---|
| Quick Assistant | Qwen2.5 0.5B | Fast, light — quick questions on any phone |
| Everyday Assistant | Llama 3.2 1B | Chatting, writing, summarizing |
| Smart Assistant | Qwen2.5 1.5B | Sharper answers, better multilingual |
| Pro Assistant | Llama 3.2 3B | Higher-quality writing and thinking |
| Expert Assistant | Phi-3.5 mini | Tricky questions, math, coding help |
Keep several installed and switch between them instantly, with no re-download.
Beyond chat — install only what you use:
| In the app | Model | Size | Good for |
|---|---|---|---|
| Quick Listener | Whisper base.en | 58 MB | Live captions, transcription |
| Careful Listener | Whisper small.en | 190 MB | Accents, noise, proper nouns |
| Vision Assistant | SmolVLM 500M | 520 MB | Captioning photos |
| Sharp Vision | Qwen2.5-VL 3B | 2.2 GB | Reading screenshots and documents |
| Audio Companion | LFM2-Audio 1.5B | 1.5 GB | Questions about a recording |
| Search Brain | BGE small | 34 MB | Semantic search over your notes |
On choosing a speech model. For transcription, use a Whisper model — it's purpose-built and small. The speech-LLMs (
kind="audio") exist for questions about audio, where tone matters. Their cost is dominated by the audio encoder, and encoders that pad every clip to a fixed 30-second window are unusable on a phone: Ultravox measured 962 s for one short utterance on an 8 GB device and never completed, while LFM2-Audio's variable-length conformer did the same job in 25 s. Seedocs/PHASE2-SPEECH.md.
EdgeLM collects no personal data. Prompts and responses are processed on your device
and are never stored or sent anywhere. The only time EdgeLM uses the network is to
download the model you choose. Full policy: docs/play/PRIVACY.md.
Add on-device AI to your app in a few lines — no model weights to ship or manage.
1. Add the SDK (via JitPack):
// settings.gradle.kts
repositories { maven { url = uri("https://jitpack.io") } }
// app/build.gradle.kts
implementation("com.github.Chandra-Mauli-Sharma.EdgeLM:sdk:0.2.0")2. Call it. Every capability is a cold Flow<String> — collect to start, cancel the
scope to stop. The runtime app must be installed.
EdgeLM.initialize(context)
EdgeLM.chat("default", "Hello", sessionId = "chat-1").collect { print(it) }
EdgeLM.caption(context, photoUri).collect { print(it) } // what's in this image?
EdgeLM.liveTranscribe().collect { print(it) } // captions as you speak
EdgeLM.askByVoice(wav) { heard -> show(heard) } // speak a question, get an answerUSE_RUNTIME and package visibility come in via manifest merge. The other capabilities are
opt-in — declare only what you use, because an app that never touches images shouldn't
advertise that it might:
<uses-permission android:name="ai.edgelm.VISION" />
<uses-permission android:name="ai.edgelm.AUDIO" />
<uses-permission android:name="android.permission.RECORD_AUDIO" />Runnable sample: samples/hello-edgelm (clone →
./gradlew installDebug) · full guide: docs/INTEGRATION.md ·
architecture & release kit in docs/.
A desktop app that talks to the runtime on your phone over adb forward: pull models,
watch throughput and benchmarks, test chat/vision/speech, manage the firewall, and read the
device log without leaving the window.
Download from Releases (Windows, macOS, Linux), or run it from source:
cd hub-desktop && npm ci && npm startBuilds are unsigned — Windows SmartScreen warns on first run, and macOS needs right-click → Open.
git clone --recurse-submodules https://github.com/Chandra-Mauli-Sharma/EdgeLM
./gradlew :runtime-service:installDebug--recurse-submodules is not optional. llama.cpp and whisper.cpp are submodules.
If whisper.cpp is missing, CMake disables speech-to-text with a warning and the build still
succeeds — leaving you an APK that silently can't transcribe. Already cloned? Run
git submodule update --init --recursive.
Iterating against a single arm64 phone? -Pedgelm.abi=arm64-v8a halves both the native
build time and the APK's native payload.
On-device. Private. Offline. Shared by every app.





