How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs
DEV Community
How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs
Mixture-of-experts models like Qwen3.6 35B-A3B, gpt-oss or GLM Flash are too big for most consumer...
0 comments
No comments yet.