On the sensitivity of machine-learned probabilistic weather forecast models to scale-aware scoring rules
Decision letter
Decision Letter
VERDICT: major
Summary of Evaluation
This note compares training objectives for a global machine-learned probabilistic weather model: an almost-fair CRPS baseline, a fair global energy score, a graph-localized energy score with a weak global anchor, and — in a second, cheaper set of twelve experiments — a family of scale-aware constructions (multi-scale residual bands, graph variogram and edge scores, spectral energy score, spectral magnitude CRPS, with and without band/variable weighting). The mathematical definitions in Section 2 are clear, self-contained and implementable, the design holds architecture, data and schedule fixed across objectives, and the authors are unusually candid about the study's preliminary character, the ad hoc nature of the weighting choices, and the possibility that greater capacity or resolution would change the ranking. The panel agreed that the questions asked are worthwhile, the design is not unsound, and the ethics/scope position is unproblematic. The citation audit found no integrity violations and judged the self-citations transparent and load-bearing.
The problem is not the design but the evidentiary layer sitting on top of it. Both central claims — that multivariate objectives give "broadly similar" skill to CRPS, and that scale-aware losses "substantially improve" small-scale variability — are supported exclusively by visual comparison of curves with no uncertainty quantification whatsoever: no confidence intervals, no paired-difference statistics, no seed replication statement, and, for the spectral results, no summary metric of any kind. All four scored reviewers arrived at this independently; I treat it as one observation rather than four votes, but it is the same observation each time and it is correct. Eighty-four initialization dates are already available, so a block bootstrap over dates costs no additional training compute. The reason this pushes the decision to major rather than minor is that the outcome of that reanalysis could change what the paper concludes: the tropical separation (graph energy best, global energy degraded) is the only differentiating skill result in the paper and the only concrete evidence offered for the closing statement that localized multivariate objectives are "promising" while the purely global energy score "appears less robust." If the paired differences over 84 dates do not clear sampling noise from a single run each, that sentence must be withdrawn, and the abstract's tropical statements with it. I am not asking the authors to obtain a particular answer; I am asking them to compute the quantity that determines which sentence is defensible.
The same logic applies to the spectral results. The manuscript's own text ("comparatively small" differences among scale-aware objectives, "slight overcompensation," "seem to be slightly more successful") reads as an accurate description of figures I cannot check, because Figures 3–14 carry one-line captions that do not state what is plotted, how it is normalized, what the reference is, or in which direction the ratios are taken, and the text does not describe their content. A single integrated diagnostic per configuration — for example, a mean absolute log-ratio to ERA5 over defined wavenumber bands, with bootstrap intervals across the 84 dates — would let a reader confirm or refute "substantially improves" and would also test the paper's most interesting assertion, that weighting matters as much as the mechanism. That assertion is currently confounded: the weights are described as untuned and their numerical values are not reported, so the weighted-versus-unweighted contrast cannot separate "weighting matters" from "these particular weights happened to help."
I differ from the more severe reviewers on two points. First, the graph-versus-anchor confound (Issue 2 in the debate synthesis) does not, in my judgement, require a new training run for this paper to be publishable. The composite objective is documented in Table 1 and run as specified; the fair remedy is to state explicitly, in the abstract and conclusion, that the tropical advantage cannot be attributed to the graph localization as opposed to the anchor or their interaction, and to drop any wording that implies otherwise. The ablation is a good next experiment, not a precondition. Second, the absence of a comparison against Pacchiardi et al.'s patched energy score is a positioning matter, not a validity failure; discuss the relationship, do not run it.
The methods audit identifies several genuine HARD gaps — random seeds and single-versus-multiple-run status, library and Anemoi versions, hardware and compute, ERA5 product variant and train/validation/test dates, early-stopping behaviour, custom code availability, data availability, and the exact ensemble-generation mechanism at inference. These are all text-and-files items and none of them individually would have changed the verdict, but they matter here more than usual because the paper's claims rest on differences between single runs: without a seed statement a reader cannot tell whether the tropical separation is a loss-function effect or run-to-run variation.
None of this is a judgement against the work's direction. A systematic, honestly reported comparison of scoring rules on a fixed architecture is useful to the field, the graph-based localization is a sensible and genuinely more general construction than fixed patches, and the finding that per-scale weighting may dominate the choice of multivariate score — if it survives quantification — is the kind of unglamorous result the field needs. What is required is that the quantitative backing catch up with the prose.
Required Revisions
-
Quantify the skill comparison in Figure 2. Report the actual fair-CRPS values (not only relative curves) for each of the three experiments, each region and each lead time shown, with confidence intervals obtained by bootstrap over the 84 initialization dates, and present the pairwise differences (CRPS vs. energy score; CRPS vs. graph energy) as paired quantities with intervals. State the margin, in CRPS units, that you regard as "broadly similar." Then reconcile the text with the result: if the tropical differences do not clear sampling uncertainty, remove the claims that the graph energy score "performs best" and that the global energy score "shows some degradation" from the abstract, Section 4.1 and Section 5, or restate them as not distinguishable from noise. Also show, rather than assert, the statement that "conclusions are also consistent when the fair energy score or fair graph energy score is used for verification."
-
Provide a quantitative spectral diagnostic for the twelve configurations. Add at least one scalar summary per experiment, variable, lead time and wavenumber band (for example mean absolute log tendency-spectrum ratio to ERA5 over the bands used in training, or an equivalent integrated measure), with sampling uncertainty across the 84 dates, presented as a table or compact figure. Use it to substantiate or soften "substantially improves," and to support or withdraw the specific ranking claims about the edge-CRPS and spectral magnitude CRPS experiments.
-
Report the weighting factors numerically and requalify the weighting claim. Give the actual values of the scale weights ζ_i, the per-variable factors applied to geopotential and mean sea-level pressure, and the spectral band weights, for every experiment in Table 2. The conclusion that weighting "can matter as much as, or more than, the specific mechanism" must either be supported by at least one alternative weighting run or comparable evidence, or be explicitly downgraded to a hypothesis consistent with the observed weighted/unweighted contrast, with the confound stated.
-
State the composite nature of the graph energy objective where the result is claimed. In the abstract, Section 4.1 and Section 5, make explicit that the graph energy experiment optimizes fGES_graph + 0.1·fES and that no experiment isolates the graph term; therefore any advantage cannot be attributed to localization rather than to the anchor or their interaction. Remove or rephrase wording that implies the localization mechanism has been validated.
-
Make Figures 3–14 self-explanatory and legible. Each caption must state the plotted quantity and its units, the normalization, the lead times shown, the reference field and — for the ratio panels — the direction of the ratio. Describe the salient features in the running text. Consolidate where possible: twelve near-identically captioned figures with distinguishable-only-by-legend curves are hard to read; consider fewer panels with clearer line styling, or move per-variable duplicates to an appendix.
-
Close the reproducibility gaps identified by the methods audit. Specifically: (a) state the random seed(s) and whether each reported result is a single run or an average over runs; (b) give Anemoi (version or commit), PyTorch and Triton versions, GPU type and count, and an approximate compute budget; (c) specify the ERA5 product variant with a DOI or CDS reference, and the exact train / validation / test date ranges, including whether a validation set was monitored during training; (d) state whether early stopping was used or whether all scheduled iterations were completed; (e) describe how the 8 ensemble members are generated at inference (noise injection mechanism and whether members are independent forward passes); (f) give the exact definition of the graph-based Gaussian smoothing operators (kernel form and width parameter), since "approximately 100/200/400/800 km" is not reproducible as stated; (g) state the batch size and, if relevant, data shuffling.
-
Add code and data availability statements. Provide a repository or archive DOI for the scoring-rule implementations and training/evaluation configurations used here (or an explicit statement of the conditions of availability), together with the Anemoi commit for the spherical-harmonic transform. State where, if anywhere, the 2022 forecast output underlying Figures 2–14 can be obtained.
-
Note the unablated hyperparameters as such. α = 0.95, k = 16, the 0.1 anchor weight, the T191 truncation and the band boundaries are all fixed without justification or sensitivity testing. Say so in one place, so readers do not assume they were selected.
-
Position the graph energy score against patch-based localization. Since Pacchiardi et al. is cited as the motivation, state plainly that no head-to-head comparison against the patched energy score was performed, and confine the advantage claimed for the graph formulation to generality of applicability (irregular grids, sparse networks) rather than performance — noting that the irregular-grid capability is argued but not demonstrated here.
Minor Suggestions
- The ablation the panel wanted most — fGES_graph without the global anchor — would be the single most informative additional run and would convert the paper's mechanistic argument from plausible to tested. Similarly, one repeat of the CRPS and graph energy experiments under a different seed would settle how much of the tropical separation is run-to-run variability. Neither is required for this revision, but both would strengthen a follow-up substantially.
- A trivial reference baseline (climatology or persistence) would let readers confirm that all three models are in a skillful regime, which the current relative plots do not establish.
- Sensitivity to ensemble size (here fixed at 8) is worth a sentence, given that the fair/almost-fair corrections exist precisely because of finite-M bias.
- The compute cost per training step of the graph and edge scores relative to CRPS is worth reporting; if it differs materially, a nominally identical iteration schedule is not an identical training budget, and readers will want to know.
- Please give arXiv identifiers or DOIs for the 2025–2026 preprint references so readers can resolve them.
- Consider adding a brief statement of funding and competing interests; the ethics reviewer found nothing substantive at issue, but the statement is conventionally expected.
- Several equations in the current PDF-to-text rendering are badly mangled (for example the CRPS and afCRPS displays, and the definition of ‖d‖_{G,n}). This is almost certainly a conversion artifact rather than an error in the source, but since the definitions are the part of the paper most likely to be reused, it is worth checking the compiled version that readers will see.