Running MiniMax-M3 — a 400B model — on two desktops
How we serve a ~400B mixture-of-experts model across two NVIDIA GB10 desktops with llama.cpp's RPC backend — stats, gotchas, and a downloadable config.
ARCHIVE // TAG · "llama.cpp"
How we serve a ~400B mixture-of-experts model across two NVIDIA GB10 desktops with llama.cpp's RPC backend — stats, gotchas, and a downloadable config.