MLX Omni Server is a local inference server for Apple Silicon that exposes OpenAI-compatible HTTP APIs on top of the MLX ecosystem (LLM/VLM + embeddings, with optional image/audio modalities). It is optimized for low-latency, low-concurrency “local/LAN trusted client” usage.
This repository is a fork of the original MLX Omni Server project with significant enhancements and modifications.
- OpenAI Responses API (
/v1/responses) with SSE streaming andinclude=["reasoning.encrypted_content"]. - Robust tool calling (Qwen3 / GLM4 / Minimax M2-focused parsing + recovery).
- Vision/VLM support via
mlx-vlm. - Unified concurrency contract (shared MLX gate + threadpool offload) across endpoints.
The upstream project provided dual API compatibility (OpenAI + Anthropic). This fork focuses on OpenAI-compatible endpoints, with deeper support for Responses, tools, and multimodal workloads.
For details on the original project, please refer to the upstream repository.
- OpenAI-compatible routes and schemas (works with the official OpenAI Python SDK via
base_url=). - Chat + streaming (
/v1/chat/completions) and Responses API (/v1/responses). - Vision (VLM inputs via standard OpenAI “content parts”).
- Structured outputs (JSON Schema) and logprobs where supported.
- Embeddings (
/v1/embeddings). - Optional modalities (install extras; routes remain registered and return
501with an install hint if missing):- Images (
/v1/images/generations) viamflux - Speech-to-text (
/v1/audio/transcriptions) viamlx-whisper - Text-to-speech (
/v1/audio/speech) viaf5-tts-mlx/mlx-audio
- Images (
- Local-first: models run on your Mac; nothing is sent to a hosted API.
The server implements these OpenAI-compatible endpoints (most routes are available with and without the /v1 prefix):
- Chat completions:
/v1/chat/completions - Responses:
/v1/responses - Audio
/v1/audio/speech- Text-to-Speech/v1/audio/transcriptions- Speech-to-Text
- Models
/v1/models- List models/v1/models/{model}- Retrieve model info
- Images
/v1/images/generations- Image generation
- Embeddings
/v1/embeddings- Create embeddings for text
For detailed API documentation and examples, see docs/README.md and docs/apis/.
- macOS with Apple Silicon (M-series)
- Python 3.11+
git(some dependencies are installed from Git URLs)- Internet access for first-time model/dependency downloads (or pre-populate your local caches)
# This repo is not published to PyPI (the PyPI name `mlx-omni-server` points to a different project).
git clone https://github.com/zhutao100/mlx-omni-server.git
cd mlx-omni-server
# Install core + all optional features (images/stt/tts)
python3 -m pip install -e ".[all]"
# Or install a subset (examples)
# python3 -m pip install -e ".[images]"
# python3 -m pip install -e ".[stt]"
# python3 -m pip install -e ".[tts]"mlx-omni-server --host 127.0.0.1 --port 10240The server starts on http://localhost:10240 by default.
curl http://localhost:10240/v1/modelsfrom openai import OpenAI
# Connect to local server
client = OpenAI(
base_url="http://localhost:10240/v1",
api_key="not-needed"
)
# Simple chat request (Chat Completions)
response = client.chat.completions.create(
model="mlx-community/gemma-3-1b-it-4bit-DWQ",
messages=[{"role": "user", "content": "Hello! How are you?"}]
)
print(response.choices[0].message.content)# Custom port
mlx-omni-server --port 8000
# Specific host and port
mlx-omni-server --host 127.0.0.1 --port 8000
# View all options
mlx-omni-server --help| Option | Default | Description |
|---|---|---|
--host |
0.0.0.0 |
Host to bind the server to |
--port |
10240 |
Port to bind the server to |
--workers |
1 |
Number of worker processes |
--log-level |
info |
Logging level (debug, info, warning, error, critical) |
--log-file |
false |
Enable on-disk logging |
--log-dir |
~/Library/Logs/mlx-omni-server |
Directory for on-disk logs (used with --log-file) |
--log-file-format |
jsonl |
On-disk log format: text or jsonl (used with --log-file) |
- Hugging Face downloads are handled by
huggingface-hub(use standard env vars likeHF_HOME,HF_TOKEN, etc. as needed). - Responses reasoning tokens: set
MLX_OMNI_SERVER_REASONING_HMAC_KEYto keepreasoning.encrypted_contentstable across server restarts. - Multi-worker (
--workers > 1) is opt-in and can be unsafe for MLX/unified-memory workloads; prefer--workers 1unless you understand the tradeoffs (seedocs/concurrency_contract.md).
- Not a full OpenAI platform clone (no Assistants, Files, Fine-tuning, etc.).
- Responses are tracked in-memory (TTL) and are not persisted across restarts.
- Designed for trusted localhost/LAN clients; not hardened for untrusted/public internet exposure.
- Start here:
docs/README.md - API reference:
docs/apis/ - Supported models:
docs/supported_models.md - Development guide:
docs/development_guide.md - Development plans:
docs/dev_plans/ - Operations:
docs/operations.md - Concurrency model:
docs/concurrency_contract.md - Architecture + roadmap:
docs/architecture_evaluation.md - Archived notes:
docs/archive/
See docs/architecture_evaluation.md for details. Current priorities:
- Add bounded backpressure around the shared MLX gate (explicit overload behavior).
- Centralize model lifecycle/budgets (cache admission + eviction policies).
- Improve observability around queueing/wait time/execution time.
This project is licensed under the MIT License - see the LICENSE file for details.
This project is a fork of MLX Omni Server by madroidmaq. We acknowledge and appreciate the original work that laid the foundation for this enhanced version.
Core Frameworks:
This project is not affiliated with or endorsed by OpenAI or Apple. It's an independent implementation providing OpenAI-compatible APIs using Apple's MLX framework.
