the.bay.news

Diffusion-based drafting LLM breakthroughs, and maybe we should work on fine-tuning for Julia; and inference rather than pretraining?

Julia Programming Language
Diffusion-based drafting LLM breakthroughs, and maybe we should work on fine-tuning for Julia; and inference rather than pretraining?
Continuing the discussion from Proof of concept LLM chatbot built with Julia: KeemenaLM.jl: JetSpec reaches 9.64x on MATH-500 and 4.58x on open-ended chat, and these gains carry into real single-stream serving on JetSpec’s own engine with an average of around 1000 TPS throughput on MATH-500 using a single B200 GPU. That’s up to 9.64x faster inference vs “only” (still great) up to 6.12x faster with the mainstream DFlash (and DDTree also beats it). @mantzaris I really appreciate the work yo...

0 comments

Sign in to join the discussion — your thebay.events account works here.

No comments yet.