β‰ˆ the.bay.news

Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

MichaΕ‚ Piszczek
Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP
How I built a 5.01 BPW Qwen3.8 27B hybrid, filled its 256K context, tuned MTP and gained 21.97% with a measured llama.cpp runtime.

0 comments

Sign in to join the discussion β€” your thebay.events account works here.

No comments yet.