> That's small enough to run well on ~$5,000 of hardware... Honestly curious whe...

simonw · 2025-11-07T21:24:06 1762550646

Yes, I mean a Mac Studio with MLX.

An M3 Ultra with 256GB of RAM is $5599. That should just about be enough to fit MiniMax M2 at 8bit for MLX: https://huggingface.co/mlx-community/MiniMax-M2-8bit

Or maybe run a smaller quantized one to leave more memory for other apps!

Here are performance numbers for the 4bit MLX one: https://x.com/ivanfioravanti/status/1983590151910781298 - 30+ tokens per second.

zht · 2025-11-08T13:11:27 1762607487

It’s kinda misleading to omit the generally terrible prompt processing speed on Macs

30 tokens per second looks good until you have to wait minutes for the first token

simonw · 2025-11-08T13:28:12 1762608492

The tweet I linked to includes that information in the chart.

oxcidized · 2025-11-08T03:10:48 1762571448

Thanks for the info! Definitely much better than I expected.

fzzzy · 2025-11-08T00:39:10 1762562350

Running in cpu ram works fine. It’s not hard to build a machine with a terabyte of RAM.

oxcidized · 2025-11-08T03:10:22 1762571422

Admittedly I've not tried running on system RAM often, but every time I've tried it's been abysmally slow (< 1 T/s) when I've tried on something like KoboldCPP or ollama. Is there any particular method required to run them faster? Or is it just "get faster RAM"? I fully admit my DDR3 system has quite slow RAM...