Signal
Advertise
Signal
Advertise
Sign in
Discover trends that matter
Trending repositories
Daily
Weekly
Monthly
Yearly
Live mentions
Topics
GitHub trending
Repositories
Developers
Insights
Stats
Log in
giannisanni/pulsar — GitHub trending stats & insights | Trendshift
Featured
open-connector
giannisanni/pulsar
#
AI infrastructure
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
Visit GitHub
Rust
107
9
1 contributors
Custom license
Social mentions
Recent discussions about this repository across the web
It a really interesting concept though, since a lot of LLMs nowadays are MoE, the actual active size is small, and since experts as a units are even smaller, how optimized can we make streaming…
@misaalanshori03 · x.com
Running Tencent Hy3, a 295B-parameter MoE, on one RTX 4060 Ti 16GB. How: NeutronStar (my CUDA fork of @antirez's ds4) keeps attention + shared experts resident and streams the routed experts off an…
@Giannisanii · x.com
this tweet sent me down a rabbit hole: 284B DeepSeek V4 Flash running on a single RTX 4060 Ti, experts streaming from a cheap NVMe via antirez's ds4. 6x speedup overnight + found and fixed the bug…
@Giannisanii · x.com
Repository activities
repository's daily and monthly activities across stars, forks, merged PRs, issues, and closed issues