Skip to content
SoftSri
Log in
Mobile LM Server - Local AI
LATEST RELEASE

Mobile LM Server - Local AI

chaterminal

Run local LLMs with llama.cpp GGUF & LiteRT-LM,API server & LM model test tools.

APKTools

Kiểm tra tương thích trên thiết bị của bạn

7.3/ 10
0
Downloads
25
Reviews
v1.1.6
Latest version

Download

v1.1.6 50.75 MB arm64-v8akatex ✓ Verified
APP PREVIEW

Mobile LM Server - Local AI screenshots

Swipe to explore

ABOUT THIS APP

About Mobile LM Server - Local AI

Mobile LLM Server - Local AI, Offline LLM & OpenAI-Compatible API for Android

Mobile LLM Server turns your Android device into a powerful local AI computing engine. Run advanced large language models completely offline, without cloud dependency, login requirements, or data sharing. Experience private, fast, and secure AI directly on your phone.

This app is designed for developers, AI enthusiasts, and advanced users who want full control over on-device AI inference. It supports multiple runtime backends including LiteRT-LM and llama.cpp, enabling flexible execution of modern open-source models such as Llama, Gemma, Mistral, Phi, Qwen, and DeepSeek distilled models.

Key Features:

LOCAL LLM INFERENCE ON ANDROID Run state-of-the-art language models directly on your device. No server required. No internet dependency. Your data stays on your phone.

OPENAI-COMPATIBLE API SERVER Transform your phone into an AI server with OpenAI-style API endpoints. Easily connect your mobile AI engine to desktop applications, automation tools, bots, or custom workflows.

OLLAMA-COMPATIBLE API SUPPORT Seamlessly integrate with Ollama-style tooling and workflows. Your Android device becomes a portable inference node in your AI ecosystem.

MULTI-MODEL SUPPORT Supports a wide range of modern open-source models including Llama series, Google Gemma, Microsoft Phi, Mistral, Qwen, and DeepSeek distilled models. Easily switch between models based on performance and memory requirements.

LITERT-LM & LLAMA.CPP ENGINE Choose between optimized mobile inference engines. LiteRT-LM provides efficient execution on Android hardware, while llama.cpp enables flexible GGUF model support with CPU/GPU acceleration.

HYBRID AI ROUTING When local context limits are reached, smart routing can optionally offload complex queries to cloud models. This ensures uninterrupted long-context reasoning and improved response quality.

PROMPT OPTIMIZATION & COMPRESSION Advanced prompt compression reduces token usage and improves inference efficiency. Long system prompts and tool descriptions are automatically optimized before execution.

DEVICE PERFORMANCE MONITORING Real-time monitoring of GPU usage, memory consumption, token generation speed, and device temperature ensures full transparency of AI workloads.

PRIVACY-FIRST DESIGN All local inference runs entirely on-device. No user data is sent externally unless explicitly configured for hybrid cloud routing.

USE CASES - Offline AI assistant on Android - Local chatbot for private conversations - Developer tool for testing OpenAI-compatible APIs - Edge AI inference for embedded and mobile systems - AI research and model experimentation - Personal AI server replacement for cloud APIs

WHY MOBILE LLM SERVER Unlike traditional cloud-based AI services, Mobile LLM Server gives you full ownership of your AI stack. It combines local inference, API server functionality, and hybrid routing into a single unified mobile platform. This makes it one of the most flexible Android AI runtime environments available today.

Whether you are building AI applications, testing models, or simply want a private offline AI assistant, Mobile LLM Server provides everything you need in one app.

Download now and transform your Android device into a powerful AI inference engine.

  • Run local LLMs with llama.cpp GGUF & LiteRT-LM,API server & LM model test tools.
LATEST RELEASE

What’s new in v1.1.6

  • No changelog was provided for this release.
COMMUNITY RATING

Review for Mobile LM Server - Local AI

7.3/10 from 25 reviews

7.3/10
VERSION HISTORY

Previous versions of Mobile LM Server - Local AI

All versions

KHÁM PHÁ RIÊNG CHO BẠN

Cá nhân hóa có kiểm soát

Lịch sử được lưu trong trình duyệt; server chỉ xử lý tín hiệu đã giới hạn để trả gợi ý và không lưu hồ sơ cá nhân.

Đang nhận diện Android và ABI từ trình duyệt…

Phổ biến tại quốc gia của bạn

Xu hướng theo khu vực đã chọn.

Tiếp tục khám phá

Dựa trên app và danh mục bạn vừa xem.

Discover

Similar apps

View all
01 App
MultiApp Ultra: Dual Space icon

Tools

MultiApp Ultra: Dual Space

waxmoon
8.8 Varies
02 App
Case Connect Mobile icon

Tools

Case Connect Mobile

Corrections Software Solutions
6.7 Varies
03 App
Signage Setup Assistant icon

Tools

Signage Setup Assistant

Samsung Electronics Co., Ltd.
7.0 Varies
04 App
Sky Clock - For Sky COTL icon

Tools

Sky Clock - For Sky COTL

Minidiel
0.0 Varies
05 App
Super Signal: Fast 5G / 4G icon

Tools

Super Signal: Fast 5G / 4G

Technik Apps
0.0 Varies
06 App
SmartWeb icon

Tools

SmartWeb

SmartThings SmartWeb
0.0 Varies
07 App
TV Launcher Premium Smart UI icon

Tools

TV Launcher Premium Smart UI

Smartago
9.0 Varies
08 App
Inventory Rampager icon

Tools

Inventory Rampager

AetherChaser
10.0 Varies