Magnitude, an Apache-2.0 inference engine that tunes itself to your machine

TL;DR

  • On 30 September Magnitude (YC S25) launched on Hacker News: an open-source, Apache-2.0 inference engine in Rust for running coding agents on local models, on macOS, Linux and Windows.
  • The founders say it tunes its kernels on your device for about a minute per model and, on an M4 Pro with Qwen 3.6 35B at 4-bit, decodes 92% faster than llama.cpp (30 to 57 tok/s).
  • Several Hacker News commenters on an M5 Max reported llama.cpp or MLX engines running faster; a founder said there may be a Metal 4 matmul issue and that he'd look into it.

Read the full story

Sign in with your email to read AI News. It’s free.

Share this story

Explain like I’m 15