Qwen3.8-Flash-Next and Qwen3.8-27B on one Radeon AI PRO R9700 (32 GB): llama.cpp settings, speed and tool-calling results, and the gotchas
Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
Widget iCUE responsive pour le monitoring de llama.cpp et des capteurs matériels.
Strata on 4 GPUs: 1 main + 3 computing expert tiers (--tier-gpus), N-way head split, N-way prompt offload, working --dump-routing. Measured ladder from a 4x RTX 5060 Ti box.
A very simple Flask app to chat with an Ollama Model from your own UI.
Offline iPhone game: kids practice Argentine Spanish with on-device AI shopkeepers (Gemma 4 + Whisper + Piper). DEV Hacktoberfest Weekend Challenge 2026.
Pure x86-64 assembly LLM engine for Gemma-2B in 5KB flat machine code. Zero runtimes, memory-saturating AVX2+F16C execution.
Local job agent: searches Workday career sites, scores postings against your resume, and tailors it with Qwen2.5-0.5B running in your browser.
Run the official Claude Code CLI on your own local models (Ollama / LM Studio / Hugging Face) at zero API cost — with model-switch VRAM eviction, auto context variants, a live status line and an LM-Studio-style tune panel. macOS / Apple Silicon.
Server di inferenza dedicato a Qwen3.6-35B-A3B (GGUF Unsloth) con MoE-expansion, basato su llama-server — build statica CUDA 12+ per Linux e Windows
Strata for a PC with two NVIDIA GPUs and 32 GB of RAM: the low-RAM mode on a layer split, both cards working at the same time. 143 tok/s on an RTX 5080 + 4060 Ti.
A measurement bench for local LLM inference. One variable at a time, raw data published, limits declared. Every number comes with the command that produced it
Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
Qwen3.8-27B-GSQ3 on one RTX 4080: 100k context with rk4v4-e8, up to 2720 tok/s prefill and 262 tok/s generation
WHIRL (Windows HIP Inference for RDNA LLMs): native Windows C++/HIP LLM inference engine for the AMD Radeon AI PRO R9700
3000 tok/s prefill; 120 tok/s decode on Strix Halo. This is a fork of gufo to support Qwen 3.6 35B A3B for workloads that prioritizes speed over intelligence.
A full, unpruned 177B MoE model at 11.5 tok/s on one 12 GB GPU. Expert streaming for llama.cpp, with the measurement harness and every dead end.
A private, local AI chat for Ollama — easy to install, runs entirely on your own computer. No account, no cloud, no subscription. Open source (GPLv3), Linux only.
The fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090: up to 500 tokens/s for one agent and up to 2,000 tokens/s for many, contexts up to 1M tokens, and a launcher that shows what every setting costs. Windows and Linux.
Kyojin: the Yamz inference engine for AMD Strix Halo (ROCm, gfx1151), built on ExLlamaV3. Runs 300B-class MoE models on one 128 GB mini PC.