Strata runs Qwen's 125B open model on a gaming PC with a 12 GB graphics card
TL;DR
- Strata, an MIT-licensed setup and inference engine on GitHub, runs Qwen's 125-billion-parameter Qwen3.8-Flash-Next on a PC with a 12 GB graphics card and 32 GB or more of RAM.
- Its makers measured 94 tokens per second on an RTX 5070, and it serves OpenAI- and Anthropic-style APIs on your own machine.
- One Hacker News commenter's picture test found larger errors through Strata than through llama.cpp with the same weights, and Strata runs on Windows and Linux, not on a Mac.
Read the full story
Sign in with your email to read AI News. It’s free.