MLXHub: Local AI & LLM Server
Developer ToolsProductivity

MLXHub: Local AI & LLM Server

Juan Colilla · released 14 Jul 2026 · Open in App Store ↗

Revenue 30d not estimated
Downloads 30d not estimated
Rating 5.00 1 reviews

No in-app products collected for this app.

Rating
Spain100%15.00
United Arab Emirates<1%0
Afghanistan<1%0
Antigua And Barbuda<1%0
Anguilla<1%0
Albania<1%0
Armenia<1%0
Angola<1%0
Argentina<1%0
Austria<1%0

The store shows a different set in every country

Description

Run AI locally. Your device. Your data.

MLXHub brings open-source language models to iPhone and iPad, powered by Apple Silicon and the mlx-swift inference engine. No cloud. No data leaves your device.

Chat with powerful models

Download and run LLMs and vision-language models (VLMs) directly on your device. Send text messages or attach photos — the model sees and responds without touching any server.

Apple Intelligence built in

On supported devices (iPhone 15 Pro / 16 and later with Apple Intelligence enabled), use Apple's on-device model for instant, private responses alongside any downloaded model.

Your personality, your assistant

Set a global Agent Personality to define the AI's tone and style. Override it per-conversation with custom system instructions. No prompt engineering needed — just describe what you want.

Browse and install models

Explore a curated catalog of optimized models — from tiny 0.6B models that fit in under 1 GB to powerful 7B models for deeper reasoning. Color-coded RAM badges tell you at a glance whether a model fits your device. Search HuggingFace directly to install any compatible model.

Local LAN server

Turn your iPhone or iPad into a portable, OpenAI-compatible inference endpoint on your local network. MLXHub's optional LAN server exposes /v1/chat/completions so any app — from a Mac running Continue.dev to a custom script — can use your device's models. No internet required. Auth-protected. Bonjour-discoverable.

Built for Apple Silicon

MLXHub uses mlx-swift, Apple's own machine-learning framework, to run models at full Metal GPU speed. Model weights are quantized (4-bit, 8-bit) for the best quality-per-GB ratio on iPhone and iPad hardware.

Privacy first

No analytics. No tracking. No ads. No account required. Model inference never leaves your device. The optional LAN server only listens on your local network and is off by default.

Terms of Use (EULA): https://www.apple.com/legal/internet-services/itunes/dev/stdeula/