Preview of the new IC2 website. It is not public yet and is hidden from search engines.

Publications

Benchmark validity in graph neural network scoring of metabolic reaction activity on Recon3D: detecting label leakage, memorized noise and input-invariant models

T Phongwattana, JH Chan

Medical Imaging

Abstract

Context-specific genome-scale metabolic modeling begins with scoring which of the ~10,600 human reactions are active in a patient’s tumor. Methods in this literature are routinely benchmarked against activity labels obtained by thresholding the same transcriptomic matrix that is supplied to the model as input. We report a self-audit of our own graph attention scorer, MetaGNN, evaluated on TCGA colorectal ( n = 624), breast ( n = 1,095) and lung adenocarcinoma ( n = 517) cohorts, in which two independent failure modes produced a near-ceiling benchmark score and a positive architectural result, neither of which survived inspection. First, under expression-thresholded supervision the framework reaches AUROC 0.9864 ± 0.0008 on TCGA-BRCA. That figure partitions into 5,925 reactions whose labels are a deterministic threshold of the model’s own input, where ranking by the cohort-mean input alone gives AUROC 1.000; and 4,675 reactions whose stored labels we reproduce bit for bit from a seeded pseudo-random number generator, where the model nonetheless reaches 0.9291 ± 0.0030 by memorizing a patient-invariant label vector that patient-level splitting leaves fully visible during training. Second, on the cohort supervised independently of the input, the archived models never received patient data at all. Their released feature tensors are uniformly zero, and independently trained models show no agreement on which patient deviates where (| r | ≤ 0.004 on per-patient output residuals, against r = +0.32 between output and input residuals on expression-bearing reactions for a model with verified features). A dispersion ratio comparing between-patient output spread against Monte Carlo Dropout sampling spread sits at 1.02–1.03 for all three configurations, against a no-signal null of 1.02 and 2.44 for the verified model. We therefore withdraw a +0.105 AUROC gain attributed to relational edges in an earlier draft of this work. Retraining on rebuilt, verified features gives AUROC 0.5800 ± 0.0017, below both the raw-expression baseline of 0.6342 ± 0.0058 that we establish for this cohort and an information-free indicator baseline of 0.6085. Zero-shot transfer of the BRCA model is at or below chance on METABRIC microarray (0.4926 ± 0.0113, n = 200) and on same-platform CPTAC-BRCA RNA-seq (0.4986, n = 106). We release the code, the curated colorectal cohort, a script that replays the label vector from its generating seed, and the screening checks we now run before reporting any score. Source code: https://github.com/thiptanawat/MetaGNN-Framework (MIT). Highlights A near-ceiling benchmark score with no biological content . Under cohort-shared, expression-thresholded supervision the framework reaches AUROC 0.9864 ± 0.0008 on TCGA-BRCA. It partitions into 1.0000 on the 5,925 reactions whose label is a deterministic threshold of the model’s own input, and 0.9291 ± 0.0030 on the 4,675 reactions whose labels are pseudo-random. The randomized partition is the airtight part of the audit . We reproduce 4,675 of the 10,600 stored BRCA labels bit for bit from a seeded pseudo-random number generator, and we release the script that does it. A model scoring 0.93 against labels that encode nothing is memorizing a vector that patient-level splitting never hides. Any benchmark whose labels are shared across patients admits this failure mode, whatever the labels contain. A model that never read its input, and three checks that reveal it . The released colorectal feature tensors are uniformly zero. Independently trained models agree on nothing patient-specific (| r | ≤ 0.004 between their per-patient output residuals), and a dispersion ratio of between-patient output spread to MC-Dropout sampling spread sits at 1.02–1.03 for all three archived configurations, against a no-signal null of 1.02 and 2.44 for a model with verified features. We withdraw our own positive result . The +0.105 AUROC gain we had attributed to reaction–reaction relational edges came from a run that changed the edge topology, the input feature set and the parameter count together. That feature set was never archived, and it shows no patient signal. It is not evidence about network structure. Nothing beat the input, and the input barely beat nothing . On the cohort with input-independent labels, raw expression alone attains AUROC 0.6342 ± 0.0058, of which an information-free indicator of which reactions carry any expression already supplies 0.6085. The one configuration we retrained on verified features reaches 0.5800 ± 0.0017, below both. What we now run before reporting a score . Compare against the raw input feature and against an information-free partition indicator; report the median inter-patient output correlation; check that the output depends on the input at all. Each costs minutes. We release all of them with the curated 624-patient cohort and its rebuilt gene-identifier mapping (94.9 % of GPR-bearing reactions resolved).

Authors: Thiptanawat Phongwattana, Jonathan H. Chan

Published in: bioRxiv (Cold Spring Harbor Laboratory) (2026)

DOI · Full text · Google Scholar