AI summary
Runs llama.cpp server as a foreground service to turn your phone into an OpenAI-compatible API endpoint on your local network. Supports GPU/NPU acceleration for local LLM inference, with a web UI at localhost:8080. Future plans include Snapdragon Hexagon NPU optimization and distributed mesh hosting via RPC. Requires broad storage and network permissions to manage model files and serve requests.
Generated by AI. May contain inaccuracies.
About this app
Run llama.cpp server with GPU/NPU acceleration on Android
An Android app that runs llama.cpp's llama-server in a foreground service, turning a phone into a headless, OpenAI-compatible API endpoint on your LAN. Open the llama.cpp server web UI at http://localhost:8080 or use your preferred frontend.
The goal is inference on the Snapdragon Hexagon NPU with phones eventually meshing over llama.cpp's RPC server to host larger models (not built yet).
About this version
- Version
- 0.1.0 (1)
- Size
- 15.64 MB
- Requires Android
- 12
- Target SDK
- 31
- Architecture
- arm64-v8a
- Downloads
- 18
- Updated
- Sep 10, 2026
- Package
- com.lumarans30.hexamesh
Similar apps
Ratings & reviews
- 50
- 40
- 30
- 20
- 10