-
Notifications
You must be signed in to change notification settings - Fork 814
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[PyTorch] Deprecate unused is_cg_capturable in parallel_cross_entropy
2.19
#3455
opened Sep 1, 2026 by
pggPL
Collaborator
Loading…
5 of 13 tasks
[Common] row-scaled nvfp4 path: fuse row/col amax into a single TMA-tiled kernel
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3454
opened Sep 1, 2026 by
cael-ling
Contributor
Loading…
2 of 13 tasks
Attribute and retry transient CP pool NaNs
2.19
#3453
opened Sep 1, 2026 by
sudhakarsingh27
Member
Loading…
7 of 12 tasks
[JAX] Add sqrtsoftplus router score function
#3448
opened Aug 31, 2026 by
jberchtold-nvidia
Collaborator
•
Draft
8 of 13 tasks
[JAX] Attention support for explicit sharding
#3446
opened Aug 31, 2026 by
jberchtold-nvidia
Collaborator
•
Draft
8 of 13 tasks
[Common] Reduce default cfg's SMEM usage for mhc triton kernel
2.19
#3442
opened Aug 28, 2026 by
kainzhong
Collaborator
Loading…
8 of 13 tasks
[Core] Fix L40 cublas heuristics
2.20
#3440
opened Aug 28, 2026 by
jberchtold-nvidia
Collaborator
Loading…
13 tasks done
[Common] Let MXFP8 quantization always reduce dbias in colwise
#3439
opened Aug 28, 2026 by
kainzhong
Collaborator
Loading…
8 of 13 tasks
Relax runtime checks for activation recompute into Warnings
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3436
opened Aug 28, 2026 by
ghadiaravi13
Contributor
•
Draft
1 of 13 tasks
[Common] Split grouped activation build
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
org-contribution
#3430
opened Aug 27, 2026 by
harryzhou2000
Contributor
Loading…
[PyTorch] Reduce CUDA graph memory retention
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3427
opened Aug 26, 2026 by
buptzyb
Contributor
Loading…
[PyTorch] Fix mutable QB bounds in CUDA graphs
org-contribution
#3426
opened Aug 26, 2026 by
harryzhou2000
Contributor
Loading…
Support paged stashing for GroupedLinear activations
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3423
opened Aug 25, 2026 by
lhb8125
Contributor
Loading…
Fix: fused QUproj + RoPE + Quant
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3421
opened Aug 24, 2026 by
ghadiaravi13
Contributor
Loading…
13 tasks
[PyTorch] Allow CP P2P transport group overrides
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3420
opened Aug 24, 2026 by
xiaoyao0115
•
Draft
[PyTorch] Bound FusedAdam tensor-handle usage
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3418
opened Aug 23, 2026 by
zupengwang
Loading…
[Common] Fix TMA synchronization in quantization kernels
#3417
opened Aug 22, 2026 by
Oleg-Goncharov
Collaborator
Loading…
6 of 13 tasks
Ring Attention: free unused kv comm buffers
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3411
opened Aug 21, 2026 by
francesco-bertolotti
Contributor
Loading…
6 of 13 tasks
Reduce Grouped MLP Fuser CPU Overhead
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3410
opened Aug 20, 2026 by
zhongbozhu
Collaborator
Loading…
13 tasks
Previous Next
ProTip!
Adding no:label will show everything without a label.