npx skills add nvidia/skills --skill nemo-mbridge-perf-tp-dp-comm-overlap
promptingcompany/nv-skills
nemo-mbridge-perf-tp-dp-comm-overlap
Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
Installation
npx skills add promptingcompany/nv-skills --skill nemo-mbridge-perf-tp-dp-comm-overlap
Similar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, includi…
1.8K installsGuidance for distinctive, intentional visual design when building new UI or reshaping an existi…
866.4K installsBrowser automation CLI for AI agents. Use when the user needs to interact with websites, includ…
810.4K installsReview UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "chec…
617.3K installsBuild, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and …
576.5K installsAlso in this package
Other skills from promptingcompany/nv-skills · top by installs.
npx skills add promptingcompany/nv-skills
More details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Also listed on
Alternate registries and mirrors of this skill.
Repository health
main
Skill metadata
Parsed from SKILL.md frontmatter.
Package contents
Files included with this skill beyond the listing page.
-
skill md
SKILL.md3,694 B -
docs
SUMMARY.md197 B
History
- First seen on skills.sh
- First recorded snapshot · 39 installs
SKILL.md
TP / DP / PP Communication Overlap Skill
For stable background and recommendation level, see:
- @docs/training/communication-overlap.md
Enablement
Minimal Bridge override:
from megatron.bridge.training.comm_overlap import CommOverlapConfig
cfg.model.tensor_model_parallel_size = 4
cfg.model.sequence_parallel = True
cfg.model.pipeline_model_parallel_size = 4
cfg.model.virtual_pipeline_model_parallel_size = 2
cfg.comm_overlap = CommOverlapConfig(
tp_comm_overlap=True,
)
cfg.ddp.use_distributed_optimizer = True
cfg.ddp.overlap_grad_reduce = True
cfg.ddp.overlap_param_gather = True
Optional TP preset:
from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
Precision knobs belong to mixed precision:
cfg.mixed_precision.grad_reduce_in_fp32 = False
cfg.mixed_precision.fp8_param_gather = False
Code Anchors
Bridge overlap gating:
```439:449:src/megatron/bridge/training/commoverlap.py if self.usercommoverlapcfg.tpcommoverlap is True: if modelcfg.tensormodelparallelsize < 2: ... elif not modelcfg.sequenceparallel: ... elif not HAVE_TE: ...
PP overlap selection:
```451:458:src/megatron/bridge/training/comm_overlap.py
if model_cfg.pipeline_model_parallel_size > 1:
if vp_size > 1:
comm_overlap_cfg.overlap_p2p_comm = True
comm_overlap_cfg.batch_p2p_comm = False
else:
comm_overlap_cfg.overlap_p2p_comm = False
comm_overlap_cfg.batch_p2p_comm = True
DP overlap defaults:
```572:579:src/megatron/bridge/training/commoverlap.py if self.dataparallelsize > 1: commoverlapcfg.bucketsize = 128 1024 1024 commoverlapcfg.overlapgradreduce = True commoverlapcfg.overlapparamgather = True
Launch-time env tuning:
```570:609:src/megatron/bridge/recipes/run_plugins.py
executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)
...
executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
Pitfalls
- TP overlap silently disables itself if
sequence_parallel=Falseor Transformer Engine is unavailable. - PP overlap is not enabled for all PP cases. Bridge only auto-selects
overlapp2pcomm=TruewhenPP > 1andVPP > 1. bucket_sizeis a parameter-count knob, not a byte-size knob.gradreduceinfp32andfp8param_gathershould be set through mixed precision, not as standalone DDP tuning first.CUDADEVICEMAX_CONNECTIONSand LayerNorm SM margin are launch-time plugin settings, notCommOverlapConfigfields.
Verification
Use the checked-in overlap unit coverage first:
uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -q
Optional second check if nemo_run is available:
uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -q
Success criteria:
- first command reports
26 passed - second command validates plugin-owned env wiring when not skipped