Inference
On-demand local inference built for agent workloads
Magnitude runs a headless local inference server. Your harness sends requests to it, and Magnitude manages the model process, memory, context, and concurrency.
⌘I
Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
On-demand local inference built for agent workloads