A high-throughput and memory-efficient inference and serving engine for LLMs
A context engine for DeepSeek Harness — reversible, token-budgeted compression of the live context window.
Qwen3.8-27B full-NVFP4 deployment on 2x NVIDIA DGX Spark with SGLang, RoCE, NGRAM/MTP, and benchmarks
Local-first, open-source experiment tracking for machine learning and agents
Connect your ChatGPT, Claude, GitHub Copilot, Grok, and OpenCode subscriptions and use them securely from local apps, APIs, SDKs, and a see them via dashboard.
Qwen3.8-27B at 34-38 tok/s on DGX Spark (GB10) — one-command SGLang + NVFP4 + DSpark setup, systemd, Claude Code ready
Qwen3.8-27B on a single RTX 3090: 416 tok/s batched or 25 ms/token single-user, 150k context, vLLM
Official Cursor SDK to Anthropic-compatible HTTP gateway with native MCP tool continuation.
Self-hosted GPU cluster orchestration: deploy and serve AI models on your own hardware
Serving Qwen3.8-27B (Unsloth UD-Q4_K_XL GGUF) at 23 tok/s over a 104,192-token context on a single NVIDIA L4 24GB, llama.cpp + MTP speculative decoding. One-command GCP provisioning, tuning sweep, quality verification, and a binary-searched context ceiling.
use glm/minimax/openai/claude api in your deepseek harness
Logical KV cache memory profiler, zombie leak detector, and fragmentation inspector for PagedAttention (vLLM, SGLang)
Reproducible Docker deployment for Qwen3.8-27B NVFP4 on NVIDIA DGX Spark (GB10, SM121) via OpenAI-compatible vLLM. Known-good baseline + guarded redeploy/rollback, diagnostics, and benchmark. Unofficial community project.
🔥 FTRAIN V1.0 — High-performance AI training framework designed for maximum speed, scalability, and memory efficiency. Streamline deep learning workflows, model pre-training, and fine-tuning with low-overhead distributed computing and optimized data pipelines. Built for high-throughput ML workloads.
Safety protocol framework for LLM agents. Binding, scope, budget, approval, monitoring, audit, kill switch. Reference architecture with on-chain binding, on-chain audit, and insurance interface.
Gemma 4 31B FP8 MTP5 vLLM profile for two Intel Arc Pro B70 GPUs
Curated community plugin directory and live marketplace for DeepSeek Harness.
Native Muse Glimmer 30B runtime for two Intel Arc Pro B70 GPUs