-
Notifications
You must be signed in to change notification settings - Fork 291
Pull requests: lightseekorg/tokenspeed
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
ci(amd-kernel): Add Kimi K3 MoE Kernel Benchmarks
#1725
opened Sep 22, 2026 by
Max191
Contributor
•
4/4
Loading…
perf(moe): skip padded stage2 reduction tiles
#1724
opened Sep 22, 2026 by
qedawkins
Contributor
Loading…
ci(amd-kernel): Add GLM 5.3 Flash MoE Kernel Benchmarks
#1723
opened Sep 22, 2026 by
Max191
Contributor
•
3/4
Loading…
ci(amd-kernel): Add GLM 5.3 Flash DSA Kernel Benchmarks
#1721
opened Sep 22, 2026 by
Max191
Contributor
•
2/4
Loading…
perf(kimi3): cut extra passes and undersized launches from gfx1250 prefill
#1720
opened Sep 22, 2026 by
Yu-Zhewen
Contributor
Loading…
feat(dcp): support GPU MLA and DSA decode context parallelism
#1718
opened Sep 22, 2026 by
OftenDream
Collaborator
•
Draft
feat(amd): add Kimi K3 prefill projection kernel
#1711
opened Sep 22, 2026 by
antiagainst
Member
Loading…
refactor(kernel): move top-k related kernel after moe_topk api
#1706
opened Sep 22, 2026 by
borontion
Contributor
Loading…
refactor(kernel): clean up dsv4 mega moe apis
#1701
opened Sep 21, 2026 by
borontion
Contributor
•
1/2
Loading…
feat (kernel): Support block-interleaved DCP MLA decode
#1694
opened Sep 21, 2026 by
wzhao18
Loading…
perf(kernel): Add an optional split-KV floor to MLA decode
#1693
opened Sep 21, 2026 by
wzhao18
Loading…
perf(kernel): wire PDL into the activation Triton kernels
#1691
opened Sep 21, 2026 by
tuanzhangCS
Contributor
Loading…
perf(engram): Fuse input preparation and n-gram hashing
#1687
opened Sep 21, 2026 by
chenht2022
Member
Loading…
fix(sampling): acquire PDL inputs before softmax reads
#1682
opened Sep 21, 2026 by
yechank-nvidia
Collaborator
•
Draft
perf(hyperconnection): fuse batched HC mix up to 256 rows
#1679
opened Sep 21, 2026 by
tuanzhangCS
Contributor
•
Draft
perf(attention): tune CuTe QSA tiling and asynchronous KV staging
#1677
opened Sep 21, 2026 by
tuanzhangCS
Contributor
Loading…
feat(kda): Add optional BF16 KDA state and FlashInfer decode/verify support
#1666
opened Sep 20, 2026 by
chenht2022
Member
Loading…
feat: support tp shard on MoE shared expert
#1652
opened Sep 20, 2026 by
byshiue
Collaborator
Loading…
perf(moe): fuse FlashInfer NVFP4 routing-map padding initialization
#1644
opened Sep 19, 2026 by
byshiue
Collaborator
Loading…
perf(amd): give gfx1250 MXFP4 decode's spare warp bit to N
#1641
opened Sep 18, 2026 by
jerryyin
Member
Loading…
refactor(kernel): unify PyTorch references under numerics/reference
#1626
opened Sep 17, 2026 by
antiagainst
Member
•
Draft
[Don't Merge][Draft] feat(kimi-k3) DEP + TP o projection on attention
#1610
opened Sep 17, 2026 by
byshiue
Collaborator
Loading…
fix(runtime): aggregate cache flush results across data parallel schedulers
#1608
opened Sep 17, 2026 by
yechank-nvidia
Collaborator
Loading…
perf(sampling): skip unused fallback-index reduction in target-only v…
#1607
opened Sep 17, 2026 by
yechank-nvidia
Collaborator
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.