Signal
Advertise
Signal
Advertise
Sign in
Submit
Discover trends that matter
Trending repositories
Daily
Weekly
Monthly
Yearly
Live mentions
Topics
GitHub trending
Repositories
Developers
Insights
Stats
Log in
noonghunna/club-3090 — GitHub trending stats & insights | Trendshift
Featured
Busbar
open-connector
noonghunna/club-3090
#
Local LLM
#
Self-hosted
Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
Visit GitHub
Like noonghunna/club-3090, 0 likes
0
Bookmark noonghunna/club-3090, 0 bookmarks
0
Python
2.1k
135
14 contributors
Apache License 2.0
Social mentions
Recent discussions about this repository across the web
Ultra-fast configuration: Main model: Frozenlock/Qwen3.8-27B-int4-AutoRound Draft model: syvai/Qwen3.8-27B-DFlash2-W4A16 GPUs: 2× RTX 3090 Tensor Parallelism: TP=2 Speculative decoding: DFlash2 KV…
@geldeki · x.com
In our quality testing low and xhigh gave similar results (apart from time spent), medium was just slightly behind.
@malikwas1f · x.com
3/ Bonus: this wasn't just a fun experiment to see what kind of capabilities you can get out of a local model, it was also a load test of the new subagent system in gmux
@mgabor_ · x.com
My two best local-AI profiles on a single RTX 5090: 27B dense: • NVFP4 KV • MTP = 3 • ~200 tok/s • 312,881 KV tokens 35B MoE: • maxed KV • ~230 tok/s • ~1,000 tok/s aggregate • 1.2M-token KV pool…
@seanhighness · x.com
I got a 35B MoE running on one RTX 5090 with: • 1.2M tokens in the KV pool • ~230 tok/s single-stream decode • ~1,000 tok/s aggregate throughput The interesting part isn't just speed. It's how much…
@seanhighness · x.com
I found the solution what had to be done and shared on the forum before, let me dig it again.
@malikwas1f · x.com
No matey! MTP is broken/misaligned with 0% acceptance rate. With MTP fixed this is going to fly.
@malikwas1f · x.com
Ran Qwen3.6-27B with vision on a single RTX 3090 24GB using ik_llama.cpp — here's the full breakdown.
@nawariokra · x.com
I was prepared for Fable fumble, and you?
@malikwas1f · x.com
Upto 1100 tps on RTX 3090x2 for Diffusion Gemma 4 26B. Unleash this mini monster on your gpus now! If you are running nvidia gpus locally, come grab the recipe at club-3090. P.S. a ⭐️ on Github is…
@malikwas1f · x.com
Load more
Repository activities
repository's daily and monthly activities across stars, forks, merged PRs, issues, and closed issues