Signal
Advertise
Signal
Advertise
Sign in
Submit
Discover trends that matter
Trending repositories
Daily
Weekly
Monthly
Yearly
Live mentions
Topics
GitHub trending
Repositories
Developers
Insights
Stats
Log in
neko-legends/spark-bench — GitHub trending stats & insights | Trendshift
Bifrost
Omnigraph
Busbar
VModal Mobile SDK
neko-legends/spark-bench
#
AI infrastructure
4x DGX Spark TP=4 serving recipe plus shadow KV prefill catch-up. Harness-agnostic (Pi, Eva, Hermes).
Visit GitHub
Like neko-legends/spark-bench, 0 likes
0
Bookmark neko-legends/spark-bench, 0 bookmarks
0
Shell
10
1
1 contributors
MIT License
Social mentions
Recent discussions about this repository across the web
That's cause dude used DS4.1 Flash on SGLANG, which corrupts outputs. That's why I used vllm for this one. Jun hey remember when we were testing SGLANG on deepseek v4.1 and it was no go because you…
@softpoo · x.com
3/ Three things worth knowing:· DSpark k=5 with a greedy draft: +30% single stream (4.36 tokens accepted/step, 67% per-position) · SGLang's DSpark path corrupts output on GB10; vLLM's is clean on the…
@softpoo · x.com
Proof of Agent work. Built by Depths, tested and approved for our infrastructure work by Eva. But need to test for stability... we shall see in the next 24 hours.
@softpoo · x.com
Used Fable 5.1 repeatedly to squeeze more performance on 4x DGX Spark + GLM 5.3 Flash EXL3. Special thanks to MiaAI_lab and Reederey for their recent updates, which my repo uses in various upgrades…
@softpoo · x.com
The bad guys already have hacking tools. The good guys need AI that will actually help defend their code, not refuse the job. Grab GLM 5.3 DFlash2 Uncensored in EXL3. It runs great on DGX Sparks.
@softpoo · x.com
100 tok/s achieved on c1 structured GLM 5.3 Flash on 4x spark. EXL3 4bpw C4 performance 253 tok/s. Bench at 420k holds: structured 93.7 / code 37.8 / math 74.5 / prose 36.2 tok/s — no regression vs…
@softpoo · x.com
OK finally got it, my contribution: the best combo for GLM 5.3 Flash on 4x sparks: **SGLang + NVFP4 + DFlash2, TP4, NCCL over RoCE**: 88 tok/s for c1, 90 for c4. Was 10-38 tok/s before optimization.…
@softpoo · x.com
Reached a new peak 159 tok/s and record 290 tok/s on good o Deepseek v4 Flash 0731 abliterated, and it was still doing the same memory work. I'm going back to this setup on the 4x sparks (of the…
@softpoo · x.com
lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL Why it matters: this is a compressed-tensors NVFP4 W4A4 Qwen3.8 derivative with the BF16 vision/video tower and all 15 native BF16 MTP tensors retained. It is…
@bussyjd · x.com
quick topology connection - 1st (left) 400g connector splits to, - bottom left spark + top left spark - 2nd (right) 400g connector splits to, - bottom right spark + top right spark The github repo…
@softpoo · x.com
Repository activities
repository's daily and monthly activities across stars, forks, merged PRs, issues, and closed issues