🧭 Stop paying for a bloated AGENTS.md on every task — split agent instructions into a small kernel + task-routed wiki topics loaded on demand, self-maintaining with byte budgets and orphan checks in CI.
Running Qwen3.8-Flash-Next on 4x RTX 5060 Ti (16GB): ~56 tok/s decode, 1.7-2.1k tok/s prefill at 495K context on a 125B-class MoE, via one expert tier with two access paths (contiguous CUDA-VMM prefill view + dynamic LRU decode mirror sharing the same VRAM rows).
RTX 3090/SM86 NInfer fork with 32% faster Qwen3.8-27B 32K prefill, RK8V4 long-context serving, and UTF-8 recovery.
Open-source cloud desktops for AI agents. Each agent gets a Linux desktop with a dev server, tests and a signed-in Chrome that pauses when idle, so your laptop stays fast.
An Open, Zero-Rent Protocol for Distributed Edge Compute and Authorized Bandwidth Relaying
Ultra-fast, sub-100ms API & webhook guardrail and triage gateway powered by TypeSafe AI System One (Jev). Parallel 7-dimension speculative evaluation, deterministic policy router, zero-dependency SQLite audit trail, and automated outbound dispatch.
Build an LLM inference engine in Rust, one tested day at a time — days 0–5 free
Rust weight-streaming engine for large diffusion transformers (FLUX.2, Qwen-Image) on consumer GPUs — pinned-memory + VRAM ring buffer, DLPack zero-copy CUDA tensors, CUDA-event prefetch. No quantization required
High-performance HTTP reverse proxy, secret redaction firewall, and two-tier caching engine for AI coding agents and LLM applications.
High-performance single-GPU inference for selected model checkpoints and GPUs.
MiMo V2.6 Pro RL on eight DGX Sparks: vLLM patches, DFlash recipe, and reproducible results.
A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops.
Local proxy that learns your app's typed LLM decisions and answers them with a Laya head. Jev and OpenAI compatible.
Native, lock-free continuous batching scheduler and KV-block table manager for dense LLMs
Anonymous fork of FreeToken — edge-native MoE serving with DeepSeek-V4.1, vision input, speculative decoding and the FTW format.
Provider-neutral System One runtime for TypeScript and Pi
An open-source, high-performance System One Model for one-pass typed decisions, with an embeddable C runtime and WebAssembly support.