LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration

Decision letter

Major revisionpanel verdict · 2026-09-01

VERDICT: major

Summary of Evaluation

LATTICE addresses a genuine gap: spatial multi-omics cohorts are increasingly assembled from heterogeneous assays, but downstream analysis usually reverts to single-modality pipelines. The paper's proposal — a single spatial graph whose node attributes concatenate five harmonized assay blocks, trained with masked reconstruction, cross-modal alignment and a Laplacian spatial term — is a sensible and clearly described design. The implementation reporting is unusually thorough for a preprint of this kind (Appendix H gives architecture, loss weights, optimizer, masking ratio, temperature, seeds, and compute footprint), and the authors deserve explicit credit for reporting the uncomfortable result rather than hiding it: adding chromatin and histone channels reduces agreement with the RNA-derived reference while increasing spatial coherence. That candour is the strongest feature of the manuscript and it is what makes the paper worth reviewing seriously.

The verdict is major revision because one central interpretive claim is not supported by the evidence presented, and the fix requires new computation whose outcome could change the conclusion.

The load-bearing problem. All five reports independently identified the same issue, and the debate did not resolve it: the abstract and §4.2 attribute the M4→M5 ARI/NMI decline plus contiguity/MUS increase to the embedding "capturing chromatin and regulatory structure beyond transcriptomic similarity." As written, this is an inference from metric movement, not a demonstration. Two properties of the evaluation make the inference circular rather than merely unproven. First, MUS (Eq. 11) is composed of spatial contiguity, silhouette, same-cluster neighbour fraction and embedding–spatial kNN Jaccard — quantities that the spatial regularization term (Eq. 9) directly optimizes. Second, spatial contiguity appears to be computed on the same k=6 kNN graph used for message passing and for the spatial loss. So a rise in contiguity and MUS as noisier features are added is consistent with the spatial term dominating the objective, and is equally consistent with the authors' preferred biological reading. Table 3 shows the spatial term is causally responsible for contiguity (0.850 → 0.783 when removed) but does not report MUS separation between "signal" and "smoothing" explanations. The decisive controls are cheap and named in the reports: MUS and contiguity for M5 with λ₃ = 0; M5 trained with the spatial ATAC and CUT&Tag blocks permuted across spots; and contiguity evaluated on a graph not used for training. If the permuted-chromatin run reproduces the contiguity/MUS gain, the paper's headline interpretation is wrong and must be withdrawn. That is why this is a major revision rather than a wording fix.

The baseline comparison. Table 2 compares methods on non-matched inputs. LATTICE M1 underperforms GraphST and STAGATE, which the authors attribute to architectural design for multimodal input; this is honest but untested. No baseline is run on the M2–M5 tensors LATTICE actually consumes, so the reader cannot separate "LATTICE's architecture helps" from "more data helps." SIMO and MaxFuse do produce integrated embeddings and the manuscript's characterization of them as producing "maps or fused views rather than a unified encoder" understates them. Running at least one graph baseline (GraphST or STAGATE) on the identical M2 feature matrix is the minimum needed to support the positioning claim.

Statistics. The abstract's "substantially improved" for M1→M2 rests on n=11 means and SDs with no paired test, and the manuscript does not state unambiguously whether the ± values are across the 11 samples or across seeds. The panel converged that the M1→M2 effect is real and mechanistically plausible; the fix is reporting, not new data. Related: the claim of "reproducible embeddings across analysis seeds" is asserted but never quantified.

Reproducibility and traceability. The data restriction is legitimate and clearly stated, and I do not penalise it. However, several load-bearing methodological elements are currently unreachable. ReCAST performs the QC that reduced 14 samples to 11, defines the five-way gene intersection, and produces the M4–M5 blocks, yet its algorithms, thresholds and code status are not specified. The "CGMC" prediction that generates the spatial ATAC and CUT&Tag blocks is never defined or attributed to a platform or method. SARSIM is cited as "bioRxiv, 2026" with no DOI, and the 10x URLs carry a future access date; the citation audit flags these as unverifiable, and SARSIM in particular is load-bearing for M2–M3. Reference genome build, library versions and per-modality QC thresholds are absent. None of these individually invalidates the work, but collectively they mean the pipeline that generates the paper's inputs cannot be inspected even in principle, which weakens the "practical framework" framing.

Claim calibration. The manuscript is generally careful, but three statements outrun the evidence: the chromatin-structure interpretation discussed above; "practical and empirically grounded framework" on the basis of a single internal cohort with internal upstream tooling; and "LATTICE is modular with respect to multimodal feature construction," which is asserted but never demonstrated on tensors from an alternative pipeline.

I want to be explicit about what I am not asking for. External validation on a public five-modality cohort, ChIP-seq confirmation of regulatory domains, and clinical-outcome prediction would all strengthen the paper, but they are not conditions for publication here; the authors already flag external benchmarking as future work and the claims can be scaled to the cohort in hand. What I do require is that the interpretation of the M4–M5 behaviour be tested against the obvious alternative explanation using the data already available, and that the input-matched baseline gap be closed.


Required Revisions

  1. Disambiguate the chromatin-signal interpretation with the negative controls. Report, for the full M5 input: (a) MUS and its four components with λ₃ = 0; (b) spatial contiguity, silhouette and MUS for a model trained with the spatial ATAC and CUT&Tag blocks permuted across spots (label shuffling within each block, spatial coordinates unchanged), averaged over at least three permutations; (c) spatial contiguity evaluated on a spatial graph not used for training or for the spatial loss (e.g. k=10 or a radius graph). If the permuted-chromatin control reproduces the M4–M5 contiguity/MUS gain, remove the "captures chromatin and regulatory structure beyond transcriptomic similarity" interpretation from the abstract, §4.2 and the conclusion, and replace it with the smoothing explanation.

  2. Run at least one spatial-graph baseline on matched multimodal input. Apply GraphST and/or STAGATE to the identical M2 feature matrix (and M5 if the method accepts it), report ARI, NMI, contiguity, silhouette and MUS under the same graph construction, K-matching and Leiden sweep, and add these rows to Table 2. State explicitly in the Table 2 caption which rows are input-matched and which are not. Revise the §2 characterization of SIMO and MaxFuse so it does not understate that both produce integrated embeddings; if they cannot be run on the M2–M5 tensors, say why in one sentence.

  3. Provide paired statistical tests for the modality-ladder transitions. State unambiguously whether the ± values in Tables 2–3 are SDs across the 11 samples or across seeds. For each adjacent transition (M1→M2, M2→M3, M3→M4, M4→M5) and for the LATTICE-vs-baseline comparisons you rely on, report a paired test (Wilcoxon signed-rank is appropriate at n=11) with effect size and a stated multiple-comparison correction. Retain or temper "substantially improved" according to the result.

  4. Quantify the seed-reproducibility claim. For the 11 analysis seeds already run, report the spread of ARI, NMI, silhouette and MUS, plus at least one embedding-stability measure (e.g. mean nearest-neighbour overlap or Procrustes distance between seed pairs). If the claim cannot be quantified, remove "reproducible embeddings across analysis seeds" from the abstract.

  5. Report per-sample results for the modality ladder. Provide a supplementary table of ARI, NMI, contiguity, silhouette and MUS per sample for M1–M5, and state whether the M2–M5 gains are consistent or concentrated in samples with higher multiome depth or larger gene intersection (Table 1 shows a 20-fold range in multiome cells and a 2.5-fold range in intersected genes).

  6. Make ReCAST and the spatial-epigenomic blocks methodologically inspectable. Define "CGMC" and name the spatial ATAC and spatial CUT&Tag platforms/protocols. State the QC criteria that excluded 3 of 14 samples and the per-modality QC thresholds actually applied. Describe the five-way gene intersection algorithmically (or provide overlap_genes.txt). State whether ReCAST code will be released; if it will not, provide enough procedural detail in the appendix that the harmonization could be reimplemented.

  7. Fix the citation and version record. Correct the dates on refs [4], [7], [8], [9] and the URL access dates on [14], [15], [17]; provide a DOI or accessible link for SARSIM [4], which is load-bearing for M2–M3. Add exact versions for PyTorch, PyTorch Geometric, Scanpy and Space Ranger, and the reference genome build used upstream. Confirm in the text where the supplementary code, environment.yml/requirements.txt and run snapshots are deposited and under what access terms (repository URL or archive DOI).

  8. Add the missing compliance statements. Include a funding statement and a competing-interests declaration (an explicit "none" is acceptable). State whether informed consent was obtained, and give the inclusion/exclusion criteria for the cohort; protocol numbers and committee names may remain withheld for double-blind review with a commitment to restore them.

  9. Define the metrics used as evidence. Give the explicit formula for spatial contiguity (SpotCut_v) in the main text or Appendix A.2. State that MUS is min–max normalized within the comparison pool and is therefore not comparable across papers or benchmark sets, and either justify the equal weighting or report a sensitivity check under one alternative weighting.

  10. Scale three claims to the evidence. (a) Temper "practical and empirically grounded framework" to reflect a single internal cohort with internal upstream tooling. (b) Either demonstrate modularity on tensors from a non-ReCAST/non-SARSIM source or restate it as compatibility-by-design rather than a demonstrated property. (c) Move the data-availability restriction from Appendix G.1 into a short limitations statement in the main text.

  11. Specify the cross-modal alignment configuration. State in the main text which modality pairs enter Eq. 8 (Appendix H implies only indices 0 and 1), justify the choice, and either report a variant that aligns additional pairs or acknowledge explicitly that three of five blocks are not directly aligned.


Minor Suggestions

  • The theoretical appendix (Lemmas I.1–I.3, Theorem I.4) restates standard spectral-graph and NCE results; Theorem I.4 as sketched is true by construction of the additive objective. Consider either shortening it to a short remark, or strengthening it to say something non-trivial about how the three terms trade off (e.g. conditions under which the spatial term overwhelms the reconstruction term).
  • Justify or ablate ρ = 0.15. Masked-autoencoder practice elsewhere uses much higher ratios; a two-point check (ρ = 0.15 vs 0.5) would settle whether the choice matters here.
  • Justify k = 6 briefly. On a hexagonal Visium lattice this is essentially the immediate neighbourhood, which is defensible but worth one sentence.
  • Report the Leiden resolution sweep range and tie-breaking rule; a data-dependent sweep to a SARSIM-derived K is a degree of freedom readers will want bounded.
  • M2 and M3 are near-identical across all metrics, suggesting scMultiome ATAC gene scores add little. This is worth a sentence rather than being passed over.
  • The marker-gene enrichment score is defined in Appendix A.2 but never reported. Either report it or remove the definition.
  • Figures 4–6 show one patient only; a brief cohort-level summary or a note that these are illustrative would prevent over-reading.
  • Learning curves or per-sample early-stopping epochs would substantiate the "stable optimization behavior" claim cheaply.
  • Consider whether the M2–M5 embeddings agree better with chromatin-derived labels than with Space Ranger labels; this is a direct positive test of your interpretation and could be reported alongside revision item 1.
  • De-identified clinical covariates (age, sex, stage, treatment class), if the data agreement permits, would help readers judge whether the patient-specific pre/post behaviour in §4.3 is plausible.

← All documents in this review