Diffusion-based drafting LLM breakthroughs, and maybe we should work on fine-tuning for Julia; and inference rather than pretraining?
Julia Programming Language
Diffusion-based drafting LLM breakthroughs, and maybe we should work on fine-tuning for Julia; and inference rather than pretraining?
Continuing the discussion from Proof of concept LLM chatbot built with Julia: KeemenaLM.jl: JetSpec reaches 9.64x on MATH-500 and 4.58x on open-ended chat, and these gains carry into real single-stream serving on JetSpec’s own engine with an average of around 1000 TPS throughput on MATH-500 using a single B200 GPU. That’s up to 9.64x faster inference vs “only” (still great) up to 6.12x faster with the mainstream DFlash (and DDTree also beats it). @mantzaris I really appreciate the work yo...
0 comments
No comments yet.