Skip to main content
Magnitude is an open source inference server that runs the best local models for your hardware, plugged into the agent you already use. It profiles your machine, recommends the models that fit, then downloads, tunes, and runs them. Works with Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, or use the built-in harness.

Why Magnitude?

  • Free to run: no token costs, API keys, or rate limits
  • Fully private and offline: models, prompts, and files stay on your machine
  • Agent-first setup: one prompt and your agent walks you through the rest
  • Knows your hardware: profiles your chip, memory, and bandwidth
  • Recommends what fits: the best models for your machine, with estimated tok/s
  • Tuned end to end: speculative decoding, concurrency, all set for your machine
  • Models on demand: loaded on request, unloaded when idle or memory fills
  • Open source: Apache 2.0, yours to modify
Magnitude works with Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, or you can use Magnitude’s built-in harness. Get started by sending one prompt to your agent. After setup, Magnitude runs in the background and loads the selected model when your agent needs it.