Signal
Advertise
Signal
Advertise
Sign in
Submit
Discover trends that matter
Trending repositories
Daily
Weekly
Monthly
Yearly
Live mentions
Topics
GitHub trending
Repositories
Developers
Insights
Stats
Log in
sudoingX/qwen38-mtp — GitHub trending stats & insights | Trendshift
Featured
Busbar
open-connector
sudoingX/qwen38-mtp
#
NLP
#
Local LLM
One llama.cpp flag unlocks +33-39% decode speed for Qwen3.8-27B on consumer GPUs. The MTP head already ships inside your GGUF. Recipe, paired benchmarks, probe tool.
Visit GitHub
Like sudoingX/qwen38-mtp, 0 likes
0
Bookmark sudoingX/qwen38-mtp, 0 bookmarks
0
Python
213
66
6 contributors
Apache License 2.0
Social mentions
Recent discussions about this repository across the web
a day ago i open sourced a repo for qwen 3.8 27b dense with two speed rows in it. it now holds numbers measured by 28 people on their own hardware, pascal to blackwell. if you run local ai on…
@sudoingX · x.com
31.0 to 41.3 tok/s on an RTX 3090, just by enabling Qwen3.8-27B's built-in speculative head. It was already sitting inside the 17GB GGUF, unused. My 3090 is getting this tonight. Free speed hidden in…
@sakurayukiai · x.com
the qwen 3.8 27b dense speed table, every number a paired a/b measured by the person who owns the card: > rtx 5090 desktop: 66 → 144 tok/s, the current crown > rtx 5090 desktop: 61 → 135, +120% > rtx…
@sudoingX · x.com
Repository activities
repository's daily and monthly activities across stars, forks, merged PRs, issues, and closed issues