mlx-serve
Run MLX and GGUF models locally on Apple Silicon Macs.
- Stars
- 276
- Forks
- 21
- Updated
- Updated Jul 14, 2026
Run MLX and GGUF models locally on Apple Silicon Macs.
mlx-serve combines a native Zig inference server for Apple Silicon with the MLX Core menu-bar app. Users can download and manage MLX or GGUF models, chat from the terminal or desktop UI, use tool-calling and agent features, and connect existing clients through OpenAI-, Anthropic-, or Ollama-compatible APIs. It does not require Python or cloud inference, but it is specifically designed for Apple Silicon Macs and remains limited by the machine's available memory and compute.
Resource types
Use cases
Platforms
Route coding agents across many model providers through one gateway.
Combine DeepSeek R1 reasoning with Claude responses through one streaming API.
Operating systems
Runtime
Protocols & integrations
Capabilities
Audience
Public GitHub facts last synced Jul 14, 2026.
Call no-auth utility APIs from scripts, automation workflows, or LLM tools.
Route prompts by complexity to control LLM API spending.