Signal
Advertise
Signal
Advertise
Sign in
Submit
Discover trends that matter
Trending repositories
Daily
Weekly
Monthly
Yearly
Live mentions
Topics
GitHub trending
Repositories
Developers
Insights
Stats
Log in
Libertai/glm53-flash-vllm-gb10 — GitHub trending stats & insights | Trendshift
Bifrost
Busbar
open-flow
open-connector
Libertai/glm53-flash-vllm-gb10
#
NLP
#
AI infrastructure
LibertAI Labs: running GLM-5.3-Flash-NVFP4 under vLLM on 2x GB10 (sm_121). Hand-written sparse-MLA CUDA kernel for NoPE MLA, plus a fix for vLLM's uninitialised NVFP4 MoE activation scale.
Visit GitHub
Like Libertai/glm53-flash-vllm-gb10, 0 likes
0
Bookmark Libertai/glm53-flash-vllm-gb10, 0 bookmarks
0
Python
6
1 contributors
Apache License 2.0
Social mentions
Recent discussions about this repository across the web
It ships as two vLLM plugin entry points. Zero patched vLLM files: VLLM_GLM53_CUDA_SPARSE_MLA=1 VLLM_GLM53_MOE_INPUT_SCALE=1.0
@MosheMalawach · x.com
Repository activities
repository's daily and monthly activities across stars, forks, merged PRs, issues, and closed issues