Self-hosted control plane that fits local AI to your hardware and real workload: it measures, tests changes in idle windows, verifies them on live traffic and rolls back regressions. llama.cpp, image, audio and ONNX/NPU engines, GPU tuning, multi-user roles, an OpenAI-compatible gateway with quotas, fleet serving and MCP. Stdlib Python, no Docker.