Signal
Advertise
Signal
Advertise
Sign in
Submit
Discover trends that matter
Trending repositories
Daily
Weekly
Monthly
Yearly
Live mentions
Topics
GitHub trending
Repositories
Developers
Insights
Stats
Log in
notwitcheer/llm-bench-rig — GitHub trending stats & insights | Trendshift
Featured
TrueForge
open-connector
notwitcheer/llm-bench-rig
#
Local LLM
Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publishable cards.
Visit GitHub
Like notwitcheer/llm-bench-rig, 0 likes
0
Python
32
3
2 contributors
Social mentions
Recent discussions about this repository across the web
ran the full-precision BF16 of Qwen3.8-27B through the same five-task board as all seven GGUF quants. it scored 93.5. the 21GB Q6_K quant scored 93.7. the BF16 is a 54.7GB split GGUF, 1.7x more than…
@witcheer · x.com
I ran the same eight cuts of Qwen3.8-27B, from an 8.4GB file to 27GB, through GPQA-diamond on my RTX 5090: 198 graduate-level science questions designed to be google-proof. on the standard benchmarks…
@witcheer · x.com
the quant tax on Qwen3.8-27B is nearly zero. seven GGUF rungs on one RTX 5090: Q8_0 down to 2-bit spans 2.9 points, with no cliff at any step. Q6_K ties Q8_0 to the second decimal, so the biggest…
@witcheer · x.com
NVIDIA ships two speculative-decode drafters with Nemotron 3.5 Lightning. I benched both on one RTX 5090, and both make the model slower. best case 0.78x, worst 0.53x, negative on all four workloads…
@witcheer · x.com
I ran all seven GGUF quants of Nemotron 3.5 Lightning through the full benchmark board on one RTX 5090. the quant tax is nearly zero: six rungs sit within 1.1 points of each other while decode runs…
@witcheer · x.com
a 2-bit quant of Qwen3.6-35B-A3B scored 92.4 on my five-task board, 0.9 under the Q4 GGUF of the same model, while decoding 6% faster: 285.7 tok/s vs 270.6 on the same RTX 5090. that is EschaLabs'…
@witcheer · x.com
Meta is back in open weights: Muse Glimmer, a dense 30B VLM, Apache 2.0, released yesterday. I benchmarked it on my RTX 5090 28B text decoder plus a 2B vision encoder. hybrid attention, three…
@witcheer · x.com
I spent yesterday mapping --n-cpu-moe on Ling-3.0-flash Q3_K_M (one 5090, 64GB DDR5): all experts in RAM is 42.9 tok/s at 3.6GB VRAM, feeding experts to the card climbs to 66 tok/s at 27.5GB, then 22…
@witcheer · x.com
Ling-3.0-flash is a week old and llama.cpp support hasn't merged yet. I benchmarked it anyway: a 127B MoE running on one RTX 5090, before upstream can load it. the numbers (Q3_K_M, 58GB file,…
@witcheer · x.com
LFM2.5-2.6B on tool-use, reasoning on vs off, 30 tasks on one 5090. result: - reasoning on: 96.7% - reasoning off: 70.0% - a 26.7 point drop, p=0.0078. 8 of 30 tasks fell over, 0 went the other way.…
@witcheer · x.com
Load more
Repository activities
repository's daily and monthly activities across stars, forks, merged PRs, issues, and closed issues