Skip to content

Repository files navigation

MLX Omni Server

Ask DeepWiki License Python Platform

MLX Omni Server Banner

MLX Omni Server is a local inference server for Apple Silicon that exposes OpenAI-compatible HTTP APIs on top of the MLX ecosystem (LLM/VLM + embeddings, with optional image/audio modalities). It is optimized for low-latency, low-concurrency “local/LAN trusted client” usage.

Fork vs Original Project

This repository is a fork of the original MLX Omni Server project with significant enhancements and modifications.

Key Enhancements in This Fork

  • OpenAI Responses API (/v1/responses) with SSE streaming and include=["reasoning.encrypted_content"].
  • Robust tool calling (Qwen3 / GLM4 / Minimax M2-focused parsing + recovery).
  • Vision/VLM support via mlx-vlm.
  • Unified concurrency contract (shared MLX gate + threadpool offload) across endpoints.

Differences from Original

The upstream project provided dual API compatibility (OpenAI + Anthropic). This fork focuses on OpenAI-compatible endpoints, with deeper support for Responses, tools, and multimodal workloads.

For details on the original project, please refer to the upstream repository.

Features

  • OpenAI-compatible routes and schemas (works with the official OpenAI Python SDK via base_url=).
  • Chat + streaming (/v1/chat/completions) and Responses API (/v1/responses).
  • Vision (VLM inputs via standard OpenAI “content parts”).
  • Structured outputs (JSON Schema) and logprobs where supported.
  • Embeddings (/v1/embeddings).
  • Optional modalities (install extras; routes remain registered and return 501 with an install hint if missing):
    • Images (/v1/images/generations) via mflux
    • Speech-to-text (/v1/audio/transcriptions) via mlx-whisper
    • Text-to-speech (/v1/audio/speech) via f5-tts-mlx / mlx-audio
  • Local-first: models run on your Mac; nothing is sent to a hosted API.

Supported API Endpoints

The server implements these OpenAI-compatible endpoints (most routes are available with and without the /v1 prefix):

  • Chat completions: /v1/chat/completions
  • Responses: /v1/responses
  • Audio
    • /v1/audio/speech - Text-to-Speech
    • /v1/audio/transcriptions - Speech-to-Text
  • Models
    • /v1/models - List models
    • /v1/models/{model} - Retrieve model info
  • Images
    • /v1/images/generations - Image generation
  • Embeddings
    • /v1/embeddings - Create embeddings for text

For detailed API documentation and examples, see docs/README.md and docs/apis/.

Quick Start

Prerequisites

  • macOS with Apple Silicon (M-series)
  • Python 3.11+
  • git (some dependencies are installed from Git URLs)
  • Internet access for first-time model/dependency downloads (or pre-populate your local caches)

Installation

# This repo is not published to PyPI (the PyPI name `mlx-omni-server` points to a different project).
git clone https://github.com/zhutao100/mlx-omni-server.git
cd mlx-omni-server

# Install core + all optional features (images/stt/tts)
python3 -m pip install -e ".[all]"

# Or install a subset (examples)
# python3 -m pip install -e ".[images]"
# python3 -m pip install -e ".[stt]"
# python3 -m pip install -e ".[tts]"

Start the Server

mlx-omni-server --host 127.0.0.1 --port 10240

The server starts on http://localhost:10240 by default.

Sanity Check (cURL)

curl http://localhost:10240/v1/models

Python Client Example

from openai import OpenAI

# Connect to local server
client = OpenAI(
    base_url="http://localhost:10240/v1",
    api_key="not-needed"
)

# Simple chat request (Chat Completions)
response = client.chat.completions.create(
    model="mlx-community/gemma-3-1b-it-4bit-DWQ",
    messages=[{"role": "user", "content": "Hello! How are you?"}]
)
print(response.choices[0].message.content)

Server Configuration

# Custom port
mlx-omni-server --port 8000

# Specific host and port
mlx-omni-server --host 127.0.0.1 --port 8000

# View all options
mlx-omni-server --help
Option Default Description
--host 0.0.0.0 Host to bind the server to
--port 10240 Port to bind the server to
--workers 1 Number of worker processes
--log-level info Logging level (debug, info, warning, error, critical)
--log-file false Enable on-disk logging
--log-dir ~/Library/Logs/mlx-omni-server Directory for on-disk logs (used with --log-file)
--log-file-format jsonl On-disk log format: text or jsonl (used with --log-file)

Configuration Notes

  • Hugging Face downloads are handled by huggingface-hub (use standard env vars like HF_HOME, HF_TOKEN, etc. as needed).
  • Responses reasoning tokens: set MLX_OMNI_SERVER_REASONING_HMAC_KEY to keep reasoning.encrypted_content stable across server restarts.
  • Multi-worker (--workers > 1) is opt-in and can be unsafe for MLX/unified-memory workloads; prefer --workers 1 unless you understand the tradeoffs (see docs/concurrency_contract.md).

Limitations / Non-goals

  • Not a full OpenAI platform clone (no Assistants, Files, Fine-tuning, etc.).
  • Responses are tracked in-memory (TTL) and are not persisted across restarts.
  • Designed for trusted localhost/LAN clients; not hardened for untrusted/public internet exposure.

Documentation

Next Steps (Roadmap)

See docs/architecture_evaluation.md for details. Current priorities:

  • Add bounded backpressure around the shared MLX gate (explicit overload behavior).
  • Centralize model lifecycle/budgets (cache admission + eviction policies).
  • Improve observability around queueing/wait time/execution time.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Acknowledgments

This project is a fork of MLX Omni Server by madroidmaq. We acknowledge and appreciate the original work that laid the foundation for this enhanced version.

Core Frameworks:

  • Built with MLX by Apple
  • API design inspired by OpenAI
  • Server implementation with FastAPI

Disclaimer

This project is not affiliated with or endorsed by OpenAI or Apple. It's an independent implementation providing OpenAI-compatible APIs using Apple's MLX framework.

About

A high-performance local inference server built on Apple's MLX framework, optimized for Apple Silicon (M-series) chips. It provides OpenAI-compatible API endpoints, enabling seamless integration with existing OpenAI SDK clients while delivering fast, private AI processing directly on your Mac.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Contributors

Languages