Magnitude, an Apache-2.0 inference engine that tunes itself to your machine
TL;DR
- On 30 September Magnitude (YC S25) launched on Hacker News: an open-source, Apache-2.0 inference engine in Rust for running coding agents on local models, on macOS, Linux and Windows.
- The founders say it tunes its kernels on your device for about a minute per model and, on an M4 Pro with Qwen 3.6 35B at 4-bit, decodes 92% faster than llama.cpp (30 to 57 tok/s).
- Several Hacker News commenters on an M5 Max reported llama.cpp or MLX engines running faster; a founder said there may be a Metal 4 matmul issue and that he'd look into it.
Read the full story
Sign in with your email to read AI News. It’s free.