โ‰ˆ the.bay.news

Prefill vs Decode: The Two Halves of Inference

DEV Community
Prefill vs Decode: The Two Halves of Inference
Why reading a prompt and writing an answer are different computations on different hardware bottlenecks, derived from arithmetic intensity.

0 comments

Sign in to join the discussion โ€” your thebay.events account works here.

No comments yet.