โ‰ˆ the.bay.news

How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs

DEV Community
How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs
Mixture-of-experts models like Qwen3.6 35B-A3B, gpt-oss or GLM Flash are too big for most consumer...

0 comments

Sign in to join the discussion โ€” your thebay.events account works here.

No comments yet.