Skip to content

[autoresearch] Goal: fold block-scaled FP8 into DeepGEMM's MegaMoE (the H200-now subset)#4

Draft
wseaton wants to merge 7 commits into
autoresearch/20260722T150510Z-goal-fold-block-scaled-fp8-into-deepgemm-basefrom
autoresearch/20260722T150510Z-goal-fold-block-scaled-fp8-into-deepgemm
Draft

[autoresearch] Goal: fold block-scaled FP8 into DeepGEMM's MegaMoE (the H200-now subset)#4
wseaton wants to merge 7 commits into
autoresearch/20260722T150510Z-goal-fold-block-scaled-fp8-into-deepgemm-basefrom
autoresearch/20260722T150510Z-goal-fold-block-scaled-fp8-into-deepgemm

Conversation

@wseaton

@wseaton wseaton commented Jul 23, 2026

Copy link
Copy Markdown

Autoresearch run 20260722T150510Z-goal-fold-block-scaled-fp8-into-deepgemm.

Goal: Goal: fold block-scaled FP8 into DeepGEMM's MegaMoE (the H200-now subset)

Gate: correctness

Kept candidates:

  • iter 1: # Candidate: Block-scaled FP8 MegaMoE on SM90 (H200) ## Summary Added an SM90 (Hopper/H200) execution path for `fp8_fp
  • iter 2: # Candidate: Block-scaled FP8 MegaMoE on SM90 — grouped GEMM path ## Summary Replaced the pure-Python per-expert dequa
  • iter 3: # Candidate: Block-scaled FP8 MegaMoE on SM90 — SwiGLU path optimization ## Summary Optimized the SM90 FP8 MegaMoE pat
  • iter 4: # Candidate: Block-scaled FP8 MegaMoE on SM90 — fused SwiGLU+requant ## Summary Optimized the SM90 FP8 MegaMoE path by
  • iter 7: # Candidate: Block-scaled FP8 MegaMoE on SM90 — direct FP8 gather with per-32→per-128 rescaling ## Summary Optimized t
  • iter 8: # Candidate: Block-scaled FP8 MegaMoE on SM90 — BF16 SwiGLU to reduce memory traffic ## Summary Optimized the SM90 FP8
  • iter 10: # Candidate: Block-scaled FP8 MegaMoE on SM90 — eliminate redundant float casts in SwiGLU+requant ## Summary Optimized

Model claude-opus-4-6 · cost $69.50 · 40203s.

🤖 opened by crucible autoresearch (draft — review and steer).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant