AI summary
Open-source AI chat app that runs LLMs locally on-device via GPU-accelerated llama.cpp and Google's LiteRT-LM, with optional cloud AI providers. Supports vision models, Stable Diffusion image generation, multi-session chat with branching, and a built-in OpenAI-compatible API server. Fixed GGUF model crash in release builds.
Generated by AI. May contain inaccuracies.
About this app
A cross-platform AI chat application with local on-device inference and multi-provider cloud AI support.
Runs LLMs directly on your Android device via GPU-accelerated llama.cpp and Google's LiteRT-LM runtime, with an optional built-in OpenAI-compatible API server.
Features
Local AI inference
- LLM inference via llama.cpp (GGUF models) with GPU acceleration (Vulkan / OpenCL) - LiteRT-LM inference via Google's LiteRT-LM runtime (.litertlm models) - Stable Diffusion 1.5 on-device image generation (safetensors) - Vision models — Qwen2-VL-2B, Gemma 4 E2B/E4B for image understanding - Streaming token generation with real-time tokens-per-second display - GPU crash recovery with automatic CPU fallback - Device-tier auto-configuration based on detected RAM - Real hardware identification (SoC and GPU renderer)
Inference parameters
- Auto Tune mode with optimal context and output limits per device - Manual mode with context window up to 1M tokens and output up to 128K tokens - Inference temperature control - Local sampling controls (Top-P, Top-K, repeat penalty) - RAM guard with safety clamping - Compute backend toggle (CPU / Vulkan / OpenCL)
Cloud AI providers
- OpenRouter, Hugging Face, xKiro, TokenRouter, AgentRouter, OrcaRouter, APInex - OpenAI, Anthropic, Google Gemini, DeepSeek, Z.AI, Groq, Mistral AI - Together AI, xAI Grok, Perplexity, Cerebras, Fireworks AI, Cohere, NVIDIA NIM - Stability AI for cloud image generation - Custom OpenAI-compatible endpoints with multiple profiles
Chat features
- Multi-session chat with history - Full-text searchable sidebar - Message actions — copy, share, regenerate, branch, edit with revision history - Prompt templates — 6 built-in plus custom - Multi-select bulk operations - Per-chat model pin - Side-by-side model comparison - Battle Arena — up to 4 cloud models race one prompt - Chat labels and hidden chats - Whole-chat PDF export
Web access
- Fetch live web content from links in messages - Clean text extraction for model context - Visible source chips with favicon and title
Skills
- Offline instruction extensions (markdown) - Intelligent per-prompt activation - 5 bundled starter skills - Import from file, URL, or Anthropic skills browser
Custom MCP Server
- Single remote connection (Streamable HTTP / SSE) - Optional bearer token authentication - Tool preview and ask-before-run dialog
Built-in OpenAI-compatible API server
- Expose local models on port 8080 - Optional API key authentication - Rate limiting
Disclaimer
CubicLM is an independent, open-source project. Not affiliated with, endorsed by, or connected to any AI model provider or third-party services. All trademarks belong to their respective owners.
License
MIT
What's new
Fixed a crash when loading GGUF models in release builds. - CI improvements.
About this version
- Version
- 1.15.1 (102024)
- Size
- 66.23 MB
- Requires Android
- 9
- Target SDK
- 28
- Architecture
- arm64-v8a
- Downloads
- 18
- Updated
- Sep 10, 2026
- Package
- com.cubiclm.app
Similar apps
Ratings & reviews
- 50
- 40
- 30
- 20
- 10