Signal
Advertise
Signal
Advertise
Sign in
Submit
Discover trends that matter
Trending repositories
Daily
Weekly
Monthly
Yearly
Live mentions
Topics
GitHub trending
Repositories
Developers
Insights
Stats
Log in
sztlink/turboquant-cuda-bench — GitHub trending stats & insights | Trendshift
Featured
open-connector
sztlink/turboquant-cuda-bench
#
AI infrastructure
Long-context quality probes and KV-cache research on local GPUs: retrieval is not utilization.
Visit GitHub
JavaScript
2
2 contributors
MIT License
website
Social mentions
Recent discussions about this repository across the web
KVarN (Huawei) vs TurboQuant, same 4-bit K / 2-bit V, Qwen3-4B. KVarN edges it on gsm8k and code (HumanEval near fp16). but MATH collapses: 57% of long outputs degenerate into repetition vs 10% for…
@sztlink · x.com
The useful split: Action can survive after target identity collapses. Source-rank can survive after target fails. Exact trace identity is the strictest layer. Fidelity is not one number.
@sztlink · x.com
3/4 The point is not "turbo3 is bad". The point is metric coverage. KLD/token-match can miss path drift. "Lossless" needs a trajectory axis. REFRACT is @no_stp_on_snek's lens. My receipt is CUDA…
@sztlink · x.com
published the first public cut of turboquant-cuda-bench: retrieved != used long-context / KV-cache receipts up to 192K on local RTX 4090: Qwen, llama.cpp, vLLM, TurboQuant, CASK, KVFidelity.
@sztlink · x.com
Small fixture. Local result. Not a benchmark claim. The diagnostic distinction is the point: retrieval = evidence enters context utilization = evidence participates in answer closure
@sztlink · x.com
At 7B / 16K / exact-match retrieval: TurboQuant K8V4 holds. Calibrated FP8 emits near-miss precision errors on the digits it should hit.
@sztlink · x.com
Repository activities
repository's daily and monthly activities across stars, forks, merged PRs, issues, and closed issues