
🛠️ Slotstream, run Qwen3.8-Flash-Next 4-bit on a low-memory Mac
Summary
Slotstream is a Mac-native tool built using MLX and Swift that allows running the large Qwen3.8-Flash-Next 4-bit model on low-memory Macs starting from 16GB. It achieves this by using expert-offloading and SSD-streaming to handle a 125B parameter model that typically requires over 100GB of RAM.
Why it’s interesting
It enables running a massive 125B parameter model on a low-memory 16GB Mac using expert-offloading and SSD-streaming.
Target user
Mac users who want to run large local language models on lower-memory hardware
Source metrics: Points 2 · Comments 0
HN discussion · Project
Source: #HackerNews / Show HN
Summary
Slotstream is a Mac-native tool built using MLX and Swift that allows running the large Qwen3.8-Flash-Next 4-bit model on low-memory Macs starting from 16GB. It achieves this by using expert-offloading and SSD-streaming to handle a 125B parameter model that typically requires over 100GB of RAM.
Why it’s interesting
It enables running a massive 125B parameter model on a low-memory 16GB Mac using expert-offloading and SSD-streaming.
Target user
Mac users who want to run large local language models on lower-memory hardware
Source metrics: Points 2 · Comments 0
HN discussion · Project
Source: #HackerNews / Show HN