An open source inference engine optimized for consumer hardware. The desktop app profiles your machine, recommends the best models for it, then downloads, tunes, and runs them. One click connects the agent you already use.
Magnitude profiles your hardware and estimates tok/s for every model in the catalog before you download anything. It ranks them by speed, accuracy, intelligence, and memory so you can pick.
They run whatever model you pick. Magnitude helps you pick. It estimates how every model and quant will perform on your machine before you download, then tunes the one you choose for your exact hardware, from context size to speculative decoding.
The desktop app is native on macOS, Linux, and Windows. It runs on Apple Silicon, NVIDIA and AMD GPUs, and CPU-only machines, including unified-memory boxes like DGX Spark and Strix Halo.
Can I use Magnitude from another computer, a container, or WSL?
Yes. Enable Network access for direct connections and use the server address and API key. SSH tunnels and mirrored WSL networking can use the local endpoint without enabling it. Only inference is available from other devices. See Network access and Windows and WSL.
Why is the first response slow and later ones fast?
Models load on demand and unload when idle, so the first request after a pause includes loading. Later turns reuse cached context, and supported models use speculative decoding. Click Load model in My Models to load ahead of time.