Skip to content

[distributed] Default MoE ep_plans to token dispatch - #49160

Merged
3outeille merged 119 commits into
mainfrom
ep-dispatch-default
Oct 6, 2026
Merged

3outeille merged 119 commits into
mainfrom
ep-dispatch-default

Conversation

@3outeille

@3outeille 3outeille commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

CPU CI GPU run-slow

Converted (44 models, 45 config classes): afmoe, axk1, axk2, cohere2_moe, deepseek_ocr2, deepseek_v2, deepseek_v3, deepseek_v32, deepseek_v4, dots1, ernie4_5_moe, ernie4_5_vl_moe, exaone_moe, flex_olmo, glm4_moe, glm4_moe_lite, glm4v_moe, glm5_next, glm_moe_dsa, gpt_oss, hunyuan_v1_moe, hy_v3, kimi_linear, laguna, lfm2_moe, mellum, minimax, minimax_m2, minimax_m3_vl, mistral4, mixtral, olmoe, openai_privacy_filter, phimoe, qwen2_moe, qwen3_5_moe, qwen3_moe, qwen3_next, qwen3_omni_moe, qwen3_vl_moe, qwen4_exp, solar_open, step3p7, zaya.

Still on masking (6 models): gemma4, diffusion_gemma, granitemoe_swa, hy_v4, inkling, mimo_v2_flash. To be fix as they were failing already before the stack of PR

@3outeille 3outeille changed the title [distributed] Default MoE ep_plans to token dispatch [distributed] Default MoE ep_plans to token dispatch Sep 28, 2026
@3outeille

Copy link
Copy Markdown
Member Author

run-slow: afmoe, axk1, axk2, cohere2_moe, deepseek_ocr2, deepseek_v2, deepseek_v3, deepseek_v32, deepseek_v4, dots1, ernie4_5_moe, ernie4_5_vl_moe, exaone_moe, flex_olmo, glm4_moe, glm4_moe_lite

@HuggingFaceDocBuilderDev

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

@3outeille
3outeille marked this pull request as draft September 28, 2026 18:15
@3outeille
3outeille changed the base branch from ep-dispatch to ep-trainer September 28, 2026 18:23
@3outeille
3outeille added this pull request to stack #48878 September 28, 2026 18:23
@3outeille
3outeille force-pushed the ep-dispatch-default branch 2 times, most recently from 66105e6 to 1cf28ef Compare September 28, 2026 19:18
@3outeille

Copy link
Copy Markdown
Member Author

run-slow: afmoe, axk1, axk2, cohere2_moe, deepseek_ocr2, deepseek_v2, deepseek_v3, deepseek_v32, deepseek_v4, dots1, ernie4_5_moe, ernie4_5_vl_moe, exaone_moe, flex_olmo, glm4_moe, glm4_moe_lite

@3outeille
3outeille marked this pull request as ready for review September 28, 2026 19:31
@3outeille
3outeille removed the request for review from zucchini-nlp September 28, 2026 19:32
@github-actions

Copy link
Copy Markdown
Contributor

Workflow Run ⚙️💔 This comment contains run-slow, but unknown error occurred and the workflow run aborted! (AMD CI)

"layers": (["hidden_states", "attention_mask"], ["hidden_states"]),
"norm": (["hidden_states"], ["hidden_states"]),
}
base_model_ep_plan = {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is a dense model, should not have ep_plan

@ArthurZucker ArthurZucker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SGTM, keeping experts in tp plan would IMO mean we are doing TP over EP. Which we are not. I think it woul dbe nice to have that written somewhere?

IDK if we plan to support it, maybenot? but that would clearup potential missunderstanding + the checks in PR stack 2 I think

@3outeille

Copy link
Copy Markdown
Member Author

SGTM, keeping experts in tp plan would IMO mean we are doing TP over EP. Which we are not. I think it woul dbe nice to have that written somewhere?

We are not doing ETP (expert tensor parallel) by keeping the experts in tp_plan, we are just keeping the masked EP/TP all_reduce for backward compatibility. I added a comment in the configuration of qwen3_moe (cf src/transformers/models/qwen3_moe/configuration_qwen3_moe.py), i'll add that to every other models

As for ETP, I dont think it's a priority to support it for now.

@3outeille
3outeille force-pushed the ep-dispatch-default branch 2 times, most recently from 87bc190 to 0161264 Compare September 30, 2026 22:57
@3outeille
3outeille removed this pull request from stack #48878 September 30, 2026 23:12
@3outeille
3outeille added this pull request to stack #49217 September 30, 2026 23:13
`initialize_distributed_mesh` now builds two named views of the same ranks:
`(pp, fsdp, tp)` for dense layers and `(pp, efsdp, ep)` for experts, both
keeping size-one axes so callers select dimensions by name. `MeshManager`
routes `ep`/`efsdp` lookups to the expert view and everything else to the
dense view. `DistributedConfig` gains `ep_size` (defaults to `tp_size` when
`enable_expert_parallel=True`) and `efsdp_size`, with size validation.

Model execution is unchanged: expert sharding and FSDP still use the `tp`
and `fsdp` axes, and loading rejects `ep_size != tp_size` until the
all-to-all dispatcher lands.
@3outeille
3outeille changed the base branch from ep-trainer to ep-dispatch October 6, 2026 08:59
Base automatically changed from ep-dispatch to main October 6, 2026 09:36
@3outeille
3outeille enabled auto-merge October 6, 2026 09:38
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

[For maintainers] Suggested jobs to run (before merge)

run-slow: afmoe, axk1, axk2, cohere2_moe, deepseek_ocr2, deepseek_v2, deepseek_v3, deepseek_v32, deepseek_v4, dots1, ernie4_5_moe, ernie4_5_vl_moe, exaone_moe, flex_olmo, glm4_moe, glm4_moe_lite

@3outeille
3outeille disabled auto-merge October 6, 2026 09:52
@3outeille

Copy link
Copy Markdown
Member Author

run-slow: afmoe, axk1, axk2, cohere2_moe, deepseek_ocr2, deepseek_v2, deepseek_v3, deepseek_v32, deepseek_v4, dots1, ernie4_5_moe, ernie4_5_vl_moe, exaone_moe, flex_olmo, glm4_moe, glm4_moe_lite

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

CI recap

Dashboard: View test results in Grafana
Latest run: 37438245206:1
Result: success | Jobs: 16 | Tests: 197,626 | Failures: 1 | Duration: 14h 39m

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

AMD CI

Workflow Run ⚙️

This comment contains run-slow, running the specified jobs on AMD:

models: ["models/afmoe", "models/axk1", "models/axk2", "models/cohere2_moe", "models/deepseek_ocr2", "models/deepseek_v2", "models/deepseek_v3", "models/deepseek_v32", "models/deepseek_v4", "models/dots1", "models/ernie4_5_moe", "models/ernie4_5_vl_moe", "models/exaone_moe", "models/flex_olmo", "models/glm4_moe", "models/glm4_moe_lite"]

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Nvidia CI

Workflow Run ⚙️

This comment contains run-slow, running the specified jobs on Nvidia:

models: ["models/afmoe", "models/axk1", "models/axk2", "models/cohere2_moe", "models/deepseek_ocr2", "models/deepseek_v2", "models/deepseek_v3", "models/deepseek_v32", "models/deepseek_v4", "models/dots1", "models/ernie4_5_moe", "models/ernie4_5_vl_moe", "models/exaone_moe", "models/flex_olmo", "models/glm4_moe", "models/glm4_moe_lite"]
quantizations: []

@3outeille
3outeille added this pull request to the merge queue Oct 6, 2026
Merged via the queue into main with commit 9921402 Oct 6, 2026
114 of 116 checks passed
@3outeille
3outeille deleted the ep-dispatch-default branch October 6, 2026 10:17
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

CI Results (Nvidia)

Workflow Run ⚙️

Commit Info

Context Commit Description
RUN 04dc9933 workflow commit (merge commit)
PR 114cc5c0 branch commit (from PR)
main d28b7937 base commit (on main)

✅ No failing test specific to this PR 🎉 👏 !

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

CI Results (AMD)

Workflow Run ⚙️

Commit Info

Context Commit Description
RUN 04dc9933 workflow commit (merge commit)
PR 114cc5c0 branch commit (from PR)
main d28b7937 base commit (on main)

✅ No failing test specific to this PR 🎉 👏 !

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants