szha.ai
Compute-optimal is not cluster-optimal
https://arxiv.org/pdf/2608.10605: Our new paper folds the systems stage into the scaling-law stage. Price every candidate architecture on what the cluster actually delivers, and the answer changes: the sparsity an MoE 'should' have depends on the cluster you train it on.
Comments
1 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
[deleted]