Serving Qwen3.8-27B (Unsloth UD-Q4_K_XL GGUF) at 23 tok/s over a 104,192-token context on a single NVIDIA L4 24GB, llama.cpp + MTP speculative decoding. One-command GCP provisioning, tuning sweep, quality verification, and a binary-searched context ceiling.