add B300 vLLM AgentX single-node MiniMax-M3 FP4 EAGLE3-GQA MTP / 新增 B300 vLLM AgentX 单节点 MiniMax-M3 FP4 EAGLE3-GQA MTP - #2328
Conversation
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
2 similar comments
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30120171478 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30121331788 |
df796e1 to
a71fd2b
Compare
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30309054681 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30326099097 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30326099097 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30375457395 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30375457395 |
…AL for eval Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
f52ea00 to
3e84e61
Compare
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
… minimaxm3-fp4-b300-vllm-agentic-mtp 条目 Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30394805725 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30566516423 |
2 similar comments
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30566516423 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30566516423 |
中文:将 MiniMax-M3 AgentX 扫描固定到 AIPerf PR #31,并同步当前 main。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30655420064 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30655420064 |
…r profiling handoff) 中文:将 MiniMax-M3 AgentX 扫描固定到 AIPerf ed05782(全局锚定性能分析切换点) Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30666322654 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30673796617 |
|
/stage-results 30673796617 |
|
@xinli-sw staged run 30673796617: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-31~r30673796617 This run remains available across future @xinli-sw 已将运行 30673796617 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-31~r30673796617 后续的 |
|
/stage-results 30514549263 |
|
@xinli-sw staged run 30514549263: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-30~r30514549263 This run remains available across future @xinli-sw 已将运行 30514549263 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-30~r30514549263 后续的 |
|
/stage-results 30506325008 |
|
@xinli-sw staged run 30506325008: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-30~r30506325008 This run remains available across future @xinli-sw 已将运行 30506325008 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-30~r30506325008 后续的 |
…hdogs active across barriers) 中文:将 MiniMax-M3 AgentX 扫描固定到 AIPerf abf55f9(跨 barrier 保持空闲看门狗活跃) Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30730035219 |
|
/stage-results 30730035219 |
|
@xinli-sw staged run 30730035219: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-02~r30730035219 This run remains available across future @xinli-sw 已将运行 30730035219 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-02~r30730035219 后续的 |
…DRAM arms TP4 GPU-resident: [1,2,5,10,15,20]; TP2 GPU-resident: [1,2,5,10,15,20]; TP4 DRAM offload: [20,30,40,50,60] 中文:扩展 MiniMax-M3 B300 MTP AgentX 扫描空间:TP4 GPU 驻留 [1,2,5,10,15,20];TP2 GPU 驻留 [1,2,5,10,15,20];TP4 DRAM 卸载 [20,30,40,50,60] Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30733110281 |
|
/stage-results 30733110281 |
|
@xinli-sw staged run 30733110281: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-02~r30733110281 This run remains available across future @xinli-sw 已将运行 30733110281 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-02~r30733110281 后续的 |
…offload TP2: [1,2,5] (drop c10,c15,c20); TP4 DRAM: [30,40,50,60,80,90,100] (drop c20, add c80/90/100) 中文:精简扫描空间——TP2 缩减至 [1,2,5],TP4 DRAM 卸载调整为 [30,40,50,60,80,90,100] Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30751989616 |
|
/stage-results 30751989616 |
|
@xinli-sw staged run 30751989616: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-02~r30751989616 This run remains available across future @xinli-sw 已将运行 30751989616 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-02~r30751989616 后续的 |
Day-zero single-node agentic-coding recipe for MiniMax-M3 (NVFP4, MoE) on B300 with vLLM and EAGLE3-GQA speculative decoding. GPU-resident KV only — no DRAM offload arm.
Contents
benchmarks/single_node/agentic/minimaxm3_fp4_b300_mtp.shconfigs/nvidia-master.yaml→minimaxm3-fp4-b300-vllm-agentic-mtp(5 configs)perf-changelog.yamlentryImage:
vllm/vllm-openai:nightly-387189c42997b27e2c04b5d97ef8190ffa2bf909Draft model:
Inferact/MiniMax-M3-EAGLE3-GQA— 3 speculative tokens, synthetic rejection sampling, thinking-on acceptance length 2.83 (from the canonical MiniMax-M3 EAGLE3 AL distribution, action 28061204145).Serve flags
use_trtllm_attention=true) with FP8 indexer KVFLASH_ATTN--enable-prefix-caching,--block-size 128,--gpu-memory-utilization 0.9--max-cudagraph-capture-size 512,--max-num-batched-tokens 16384,--stream-interval 20--reasoning-parser minimax_m3,--default-chat-template-kwargs '{"thinking_mode":"enabled"}'--all2all-backend flashinfer_nvlink_one_sidedSearch space
Five configs, GPU-resident KV only:
dram-utilization: 0.80is set in the matrix but only applies to DRAM-offload arms; all points here havekv-offloading: none.Eval logic
Added
EVAL_ONLYbranch: whenEVAL_ONLY=truethe script callsrun_eval --port $PORTinstead of the benchmark replay path, so lm-eval can be run against a live server without re-running the full agentic sweep.中文说明
MiniMax-M3(NVFP4,MoE)在 B300 单节点上的首日 agentic-coding 配方,使用 vLLM 搭配 EAGLE3-GQA 投机解码,仅含 GPU 驻留 KV 配置(无 DRAM 卸载分支)。
镜像:
vllm/vllm-openai:nightly-387189c42997b27e2c04b5d97ef8190ffa2bf909草稿模型:
Inferact/MiniMax-M3-EAGLE3-GQA,投机 token 数 3,合成拒绝采样,thinking-on 合成接受长度 2.83(来自 MiniMax-M3 EAGLE3 AL 分布,action 28061204145)。共 5 个配置:TP8 并发 1;TP4 并发 1、2、16;TP2 并发 2。新增
EVAL_ONLY分支,支持对运行中的服务端单独执行 lm-eval 评估。