โ‰ˆ the.bay.news

Did FP8 make the model dumber? A per-prompt regression check for quantized serving

DEV Community
Did FP8 make the model dumber? A per-prompt regression check for quantized serving
FP8 gave us a clean 1.5x on Qwen3-8B serving throughput on an RTX PRO 6000 Blackwell (1,725 to 2,597...

0 comments

Sign in to join the discussion โ€” your thebay.events account works here.

No comments yet.