Back to list
High-Potential
Swift
⚡ Slotstream: Stream 105GB Models on Mac
413 stars28 forksSwift
apple-siliconclaude-codeinference-enginellmllm-inferencelocal-ailocal-llmmacosmetalmixture-of-expertsmlxmoe
The technical direction here is quite clever: it attempts to run a 105 GB AI model on Macs with limited RAM. Specifically, it streams a 125B Mixture of Experts model (Qwen3.8-Flash-Next) directly from the SSD, caching only the most active experts in memory. This allows Macs with 16 to 64 GB of RAM to handle the inference.
Built as a native Swift binary using MLX and Metal, it requires no Python and runs entirely offline. It also maintains compatibility with Claude Code, Codex, and OpenAI clients. This engineering approach of trading storage I/O for memory offers an interesting workaround for running massive models locally.