the.bay.news

Ferrum 0.8.4 — a single-binary Rust LLM runtime for Metal and CUDA

The Rust Programming Language Forum
Ferrum 0.8.4 — a single-binary Rust LLM runtime for Metal and CUDA
Hi Rustaceans, I’m the maintainer of Ferrum, an MIT-licensed local LLM inference runtime written in Rust. Version 0.8.4 is now available. Ferrum aims to make running and serving local LLMs feel like installing a normal CLI: One Rust binary No Python, PyTorch, or vLLM runtime Metal acceleration on Apple Silicon CUDA support on NVIDIA GPUs An OpenAI-compatible chat completions API Apple Silicon quick start: brew tap sizzlecar/ferrum brew install ferrum ferrum doctor qwen3.5:4b-q4_k_m...

0 comments

Sign in to join the discussion — your thebay.events account works here.

No comments yet.