FreeToken claims 39.3 tok/s for Qwen3.6-35B on an 8GB RTX 4060 laptop GPU
programming.dev
FreeToken claims 39.3 tok/s for Qwen3.6-35B on an 8GB RTX 4060 laptop GPU
>[https://lemmy.ml/api/v3/image_proxy?url=https%3A%2F%2Flemmy.world%2Fpictrs%2Fimage%2F77d45aa2-fb79-4737-acb2-fca56300d7af.jpeg] > >The project is an MoE-native serving engine that treats GPU, CPU, host RAM, and PCIe bandwidth as one inference platform. ⚡ > >Published paper results include: > >text >Qwen3.6-35B-A3B >RTX 4060 Laptop 8GB >39.3 tok/s > >DeepSeek-V4-Flash 284B >RTX 5090 >22-25 tok/s > > >The full expert pool lives in system RAM and VRAM acts as an expert cache. > >On cache misses, FreeToken can either transfer an expert to the GPU or execute it directly on the CPU, with the split chosen from measured bandwidth. > >Important caveat: low VRAM does not mean low total memory. The host RAM still has to hold the expert weights. > >https://github.com/FlashML-org/FreeToken [https://github.com/FlashML-org/FreeToken] > >Current support is mainly x86_64 + NVIDIA RTX 30/40/50-series hardware. > >Has anyone here benchmarked it against llama.cpp/Ollama on the same checkpoint and hardware? I would be interested in real-world agent workloads rather than short synthetic decode tests.
0 comments
No comments yet.