
🛠️ Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Summary
Slotstream is a Mac-native tool built with MLX and Swift that runs a large 125B parameter model on low-memory Mac hardware using expert-offloading and SSD streaming. It features an auto-mode for balancing memory usage and speed.
Why it’s interesting
It enables running a massive 125B parameter model on a 48GB Mac at roughly 12 tokens per second.
Target user
Mac users looking to run large language models on lower-memory hardware.
Source metrics: Points 10 · Comments 0
HN discussion · Project
Source: #HackerNews / Show HN
Summary
Slotstream is a Mac-native tool built with MLX and Swift that runs a large 125B parameter model on low-memory Mac hardware using expert-offloading and SSD streaming. It features an auto-mode for balancing memory usage and speed.
Why it’s interesting
It enables running a massive 125B parameter model on a 48GB Mac at roughly 12 tokens per second.
Target user
Mac users looking to run large language models on lower-memory hardware.
Source metrics: Points 10 · Comments 0
HN discussion · Project
Source: #HackerNews / Show HN