Review v2 · round 1 · manuscript v2

On the sensitivity of machine-learned probabilistic weather forecast models to scale-aware scoring rules

Simon Lang, Martin Leutbecher, Sam Hatfield

Major revisionadvisory · no human graded this paper
DOI
10.48550/arXiv.2607.19161
Reviewed
2026-09-01

Panel readout

4 of 5 specialists scored · 1–5

70/ 100

Legacy scaled score

Range 3.0–4.0, a spread of 1.0.

  1. data analysis4.0Confidence 4 of 5
  2. contribution context4.0Confidence 4 of 5
  3. reporting reproducibility3.0Confidence 4 of 5
  4. scientific validity3.0Confidence 4 of 5
  5. ethicsThe reviewer gave no score and did not say why.n/a

Score is the referee's assessment of the work.Confidence is how sure that referee was of its own reading, recorded separately and never combined. The editor's verdict is its own judgment of the reports, not a threshold applied to this mean. The legacy aggregate is the historical panel mean multiplied by 20. It is not comparable to the Editor-in-Chief's publication readiness score. n/a means that referee found nothing in its dimension to judge here. It is left out of the mean rather than counted as a good score, which is what a forced number quietly became.

Abstract

as posted by the authors

Probabilistic forecast models can be machine-learned from data using loss functions based on scoring rules such as the Continuous Ranked Probability Score (CRPS). This note summarises a preliminary study comparing versions of AIFS-CRPS, a global weather forecast model, trained with different univariate and multivariate scoring rules that aim to explicitly represent scale-awareness in the loss function. In the first part, we compare the (almost) fair CRPS, a fair global energy score, and a graph energy score based on node neighbourhoods. Across standard verification metrics, forecast skill is broadly similar. In the extratropics we find only small differences, while in the tropics the graph energy score setup performs somewhat better and the global energy score shows some degradation. These results suggest that multivariate scores are a viable alternative to CRPS-based training for global machine-learned weather forecasting. In the second part of the study, we analyse how different scoring rules and scale-aware loss constraints shape the spectra of forecast fields. It is apparent that any form of explicit scale-awareness improves realism. Here, the largest differences are likely associated with different effective weights per scale.

The review

Specialist reports

Editorial audits

Factual checklists, not opinions. They skip the debate and go straight to the editor.

The text the panel read

counted, not judged

Counted at ingest, no model involved. These describe theconverted text the referees read, not your PDF.

Size

Words
5,160
Main text
4,381

excluding references

Sentences
238
Display equations
48

Sentences

Median sentence
20 words
Longest tenth
35 words
Over 40 words
6%
Passive
~0.2773/sentence

regex approximation

Evidence on the page

Citations
64

12.4 per 1000 words, numeric style

Numbers
52.33

per 1000 words

p-values
0 exact, 0 threshold

Hedging against amplifying

Per 1000 words. Softening ("may", "suggests") against strengthening ("clearly", "demonstrates"). No referee saw these.

Hedging3.68
Amplifying0.39

Provenance

PeerReviewAgents 0.5.0 · 1cf57690

Reviewed file

v2 · 5243 KiB

sha256 4de252a87afbf9423494b4b5…

Debate rounds

2

Run cost

$1.28

Manuscript read as

markdown

Converted by rustypaper 0.2.0. Quotations match the manuscript text.

Desk screen

triage gate

Which model wrote which report
StageModel
Editorial audits (×2)claude-haiku-4-5
Advocate / skepticclaude-sonnet-5
Specialist reviewers (×5) + desk screenclaude-haiku-4-5
Editor, debate synthesizerclaude-opus-5
debate synthesizer (override)claude-sonnet-5
journal recommender (override)claude-haiku-4-5

The widest fan-out runs on the cheapest model. Only the agents that decide the verdict run on the most capable one.

What the referees looked up (25 searches)
AgentSearchAnswered byHits
reviewer contribution contextfind related workCRPS machine learning weather forecasting ensembleN/A5
find related workenergy score multivariate probabilistic forecastingN/A5
find related workgraph neural networks weather prediction scoring rulesN/A5
search preprintsscale-aware loss weather forecasting neural networksN/A5
search preprintspatched energy score localized multivariate scoringN/A0
find related workPacchiardi patched energy score generative networksN/A5
find related workLang AIFS-CRPS ensemble weather forecastingN/A5
find related workvariogram score multivariate ensemble forecastingN/A5
search preprintsAIFS ensemble probabilistic weather machine learningN/A0
find related workFourCastNet spectral loss weather forecastingN/A5
find related workScheuerer variogram score proper scoring rulesN/A5
search preprintsmulti-scale loss formulation weather forecasting neural networksN/A5
search preprintsgraph energy score localized multivariate forecastN/A2
find related workfair CRPS Ferro ensemble forecast bias correctionN/A5
find related workGneiting Raftery proper scoring rules strictly properN/A5
search preprintsFGN FourCastNet Huracan machine learning weather 2025N/A0
search preprintsAlet Lam Battaglia skillful joint probabilistic weather 2025N/A0
search preprintsBonev Kurth FourCastNet 3 geometric probabilistic 2025N/A0
search preprintsLang Leutbecher Maciel multi-scale loss formulation 2025N/A0
find related workPacchiardi Adewoyin Dueben Dutta probabilistic forecasting generative networks 2024N/A5
find related worklocalized multivariate scoring patch-based energy scoreN/A5
find related workspectral fidelity machine learning weather forecasting attentionN/A5
find related workgraph neural networks spatial localization weather predictionN/A5
find related workPic Dombry Naveau Taillardat proper scoring rules aggregation transformationN/A5
find related workZhdanov sparse attention spectral fidelity weather forecastingN/A5

Run against arXiv, Semantic Scholar, PubMed and bioRxiv while the review was being written. A search returning zero hits is kept: it is the evidence behind a referee saying it found no prior art.

What each agent cost
AgentUSD
editor$0.4351
skeptic$0.2406
advocate$0.2151
reviewer contribution context$0.1549
debate synthesizer$0.0760
audit citation integrity$0.0484
desk screen$0.0303
audit methods completeness$0.0190
journal recommender$0.0144
reviewer reporting reproducibility$0.0132
reviewer scientific validity$0.0127
reviewer data analysis$0.0126
reviewer ethics$0.0042

Cite this review

Permanent: this review only

This URL is a permanent link to this specific review, and will not change.

Plain text
In Silico (2026). Review of "On the sensitivity of machine-learned probabilistic weather forecast models to scale-aware scoring rules". In Silico. https://pgarrett-scripps.github.io/insilico/reviews/2026/on-the-sensitivity-of-machine-learned-2607-19161/v2/
BibTeX
@misc{insilico-on-the-sensitivity-of-machine-learned-2607-19161-v2,
  title        = {Review of {On the sensitivity of machine-learned probabilistic weather forecast models to scale-aware scoring rules}},
  author       = {{In Silico}},
  year         = {2026},
  howpublished = {In Silico, an AI-refereed overlay journal},
  url          = {https://pgarrett-scripps.github.io/insilico/reviews/2026/on-the-sensitivity-of-machine-learned-2607-19161/v2/},
  note         = {Machine-generated peer review of doi:10.48550/arXiv.2607.19161 v2. Produced by PeerReviewAgents 0.5.0. Produced by PeerReviewAgents, doi:10.5281/zenodo.21781895.}
}

Please cite the preprint itself as well. This reviews that work, it does not replace it. The review is machine-generated and advisory. If you are citing it as evidence about the paper, say so explicitly.