Inference
On-demand local inference built for agent workloads
Magnitude runs a headless local inference server. Your harness sends requests to it, and Magnitude manages the model process, memory, context, and concurrency.
On-demand local inference built for agent workloads