LM Pocket
Developer ToolsUtilities

LM Pocket

Teo Bouancheau · released 2 Jul 2026 · Open in App Store ↗

Revenue 30d not estimated
Downloads 30d not estimated
Rating 5.00 2 reviews

No in-app products collected for this app.

Rating
France100%25.00
United Arab Emirates<1%0
Afghanistan<1%0
Antigua And Barbuda<1%0
Anguilla<1%0
Albania<1%0
Armenia<1%0
Angola<1%0
Argentina<1%0
Austria<1%0

The store shows a different set in every country

Description

Discover, download, and run open-source AI models directly on your device.

No cloud. No subscription. No data leaves your device.

LM Pocket brings the power of local AI models to your device.

Browse thousands of GGUF models on Hugging Face, download them over WiFi, and chat with them entirely on-device.

COMPLETELY PRIVATE

Every conversation stays on your device. No server, no API key, no account required.

Your prompts and responses never leave the device.

DISCOVER MODELS

Browse Hugging Face Hub for quantized GGUF models.

Filter by size, quantization level, and popularity.

See compatibility badges that tell you which models will run well on your device before downloading.

CHAT WITH AI

Stream token-by-token responses with full Markdown rendering.

Manage multiple conversations. Adjust temperature, top-p, top-k, context length, and more.

Supports ChatML, Llama, Mistral, Phi, Gemma, and other prompt formats automatically.

LOCAL API SERVER

Run an OpenAI-compatible HTTP server on your device. Expose /v1/chat/completions, /v1/models, and /v1/embeddings endpoints on your local network.

Connect any app that supports the OpenAI API format.

MANAGE YOUR LIBRARY

Track installed models with storage usage. Import GGUF files from the Files app. Load and unload models with one tap. Monitor real-time performance: tokens per second, memory usage, and thermal state.

BUILT FOR POWER USERS

GPU acceleration for fast inference

Background downloads with pause and resume

Hardware-aware model recommendations

Chat template auto-detection

Conversation export as JSON or Markdown

JSON mode with grammar-constrained output

Embeddings generation

100% FREE

No subscriptions, no in-app purchases, no ads.

Open-source AI should be accessible to everyone.

SUPPORTED MODELS

Run any GGUF-quantized model: Llama, Mistral, Phi, Gemma, CodeLlama, DeepSeek, Qwen, and thousands more from Hugging Face Hub.

Model performance depends on device RAM and model size.

Terms of Service: https://www.apple.com/legal/internet-services/itunes/dev/stdeula/