A standardized glucose challenge can be highly repeatable without being useful for predicting responses to other meals. That distinction matters for CGM-based personalization, because a test meal is valuable only if it measures something that travels beyond the test itself.
This research note reanalyses open CGM challenge data from CGMacros and the Stanford/Hall study, then uses Food & You as an external comparison. The aim is not to identify a “best breakfast.” It is to ask a narrower question: what makes one standardized meal informative about future responses to different meals?
Three questions that should not be mixed
We separate three quantities:
- Same-meal repeatability: does the same person respond similarly when the same challenge is repeated?
- Cross-meal transportability: does a person’s deviation on the challenge track their deviation on other meals?
- Out-of-sample predictive gain: does observing the challenge actually improve prediction of later, different meals?
These are not interchangeable. A repeatable test can measure something specific to that test. Conversely, a moderately noisy test can still be useful if the part it captures is aligned with a broader person-level response.
Datasets
CGMacros contains 45 participants spanning no diabetes, prediabetes and type 2 diabetes, with repeated standardized breakfasts, CGM and clinical measurements. The public dataset is available through PhysioNet.
Stanford/Hall includes a 30-person standardized-meal subset in which participants consumed cornflakes with milk (CF), a peanut-butter sandwich (PB) and a protein bar (Bar), generally twice each. The original study and standardized-meal glucose traces are available from PLOS Biology.
Food & You provides a larger independent context: 50,463 free-living meals from 992 adults plus 4,524 standardized meals. Its standardized challenges were a glucose drink, white bread, and white bread with butter. We use published and repository-level summary results from the recent Food & You analysis; we do not treat it as a completed one-shot prospective validation.
A stricter harmonized test
The original CGMacros analysis suggested a striking one-meal calibration effect. To test whether the contrast with Hall was partly methodological, we re-ran both datasets under a common pipeline:
- 5-minute temporal grid;
- baseline from the -10 and -5 minute glucose values;
- signed mean glucose change over 0–120 minutes;
- premeal glucose level and premeal slope entered separately as state covariates;
- leave-one-person-out evaluation;
- the first challenge predicts only later different recipes;
- no future data from the test participant used to fit calibration weights.
For Hall, the primary analysis required all needed 5-minute points to be directly observed. A sensitivity analysis allowed interpolation only across gaps of 10 minutes or less.
The CGMacros–Hall contradiction became smaller
Under the older CGMacros setup, the first standardized breakfast reduced future prediction error by roughly one third. Under the stricter cross-recipe, state-adjusted analysis, the effect remained but became smaller.
| Challenge | Same-meal repeatability | Cross-meal transport r | MAE gain | Standardized gain |
|---|---|---|---|---|
| CGMacros reference | 0.667 | 0.660 | 15.7% | +0.118 SD |
| Hall cornflakes | 0.859 | 0.590 | 21.1% | +0.139 SD |
| Hall PB sandwich | 0.634 | 0.020 | -9.3% | -0.073 SD |
| Hall protein bar | 0.252 | 0.538 | 0.2% | +0.001 SD |
The Hall cornflakes challenge no longer looks like a failed replication of CGMacros. On the strict sample, its standardized predictive gain was at least comparable to the CGMacros reference challenge, although Hall’s sample is small and its bootstrap interval is wide.
This means part of the earlier story, “CGMacros works while Hall does not,” was created by differences in evaluation design. The more durable result is challenge specificity: different standardized meals can have very different value as calibration tests.
Repeatability is not transportability
The clearest example is the Hall peanut-butter sandwich.
Its state-adjusted same-meal repeatability was about 0.63. If repeatability were the main requirement for personalization, this should have been a reasonably good challenge. Instead, its correlation with later responses to different meals was approximately 0.02, and using it for personalization made out-of-sample MAE worse.
The protein bar shows a different failure mode. It retained some cross-meal association, but its same-meal reliability was weak and its practical predictive gain was approximately zero.
Cornflakes combined both ingredients more successfully: substantial repeatability and substantial cross-meal transportability, followed by positive predictive gain.
So repeatability answers, “can I measure this response again?” Transportability answers, “does this measurement tell me anything about other meals?” They are different properties.
A large glucose excursion is not enough either
Food & You provides a useful independent stress test of another simple explanation: perhaps a challenge is informative merely because it produces a large glucose excursion.
The mean positive iAUC differed strongly across its standardized meals:
- glucose drink: 3,470 mg·min/dL;
- white bread: 2,613 mg·min/dL;
- white bread with butter: 1,959 mg·min/dL.
Yet the correlation between each standardized-meal iAUC and the study’s dominant free-living person-level response axis (PC1) was almost unchanged: 0.53, 0.51 and 0.51, respectively.
The same pattern appears for peak glucose. Mean peak glucose fell from 149.0 to 133.0 to 120.2 mg/dL, while the corresponding correlations with PC1 were 0.65, 0.66 and 0.68.
This does not prove that response amplitude is irrelevant. A larger excursion can improve measurement signal-to-noise. But it argues against the simpler claim that the largest spike is automatically the most informative calibration challenge.
A candidate mechanism
The results fit a simple decomposition:
observed challenge response = general responder component + challenge-specific component + premeal/day state + measurement noise
A useful standardized challenge would then be one that captures enough of the stable general person-level component, while keeping challenge-specific and transient noise manageable.
Our current working approximation is:
challenge utility ≈ cross-meal alignment × measurement reliability
This is a hypothesis, not a validated formula. In the discovery datasets, transportability and predictive gain are partly constructed from related future outcomes, so their association cannot itself establish the mechanism.
What this note does not establish
- It does not identify a universally optimal standardized meal.
- It does not show that one breakfast is sufficient for everyone.
- It does not prove that glucose amplitude is unimportant.
- It does not establish the “alignment × reliability” mechanism independently.
- It does not provide a completed prospective first-challenge test in Food & You. The available summary correlations use repeated standardized challenges, and the published temporal split is not guaranteed to begin after the first challenge for every participant.
CGMacros and Hall should now be treated as discovery datasets: the current hypothesis was developed while looking at them. Independent confirmation requires a new dataset in which the analysis rules are fixed in advance.
What would falsify the current hypothesis?
The mechanism would be weakened if an independent dataset showed that challenges with stronger training-set cross-meal alignment and adequate reliability repeatedly failed to improve held-out predictions, or if amplitude or same-meal ICC consistently predicted out-of-sample gain while alignment contributed little.
The broader idea of one-meal personalization would become weak if adequately powered independent datasets repeatedly showed approximately zero or negative prospective cross-meal gain after chronology, same-recipe exclusion and leakage control were enforced.
Conclusion
A standardized metabolic challenge should not be judged only by how strongly glucose rises or by how reproducible the same meal is.
The central distinction is simpler:
A repeatable response is not necessarily a transportable response.
Some standardized meals appear to measure a person-level signal that generalizes to other meals. Others mainly reproduce meal-specific physiology or transient state. The next step is not to search for a more dramatic glucose spike, but to test, prospectively and out of sample, which challenges capture the transferable part of individual glucose response.
This study is part of the Null Institute’s Verification Pilots, a series of small, reproducible studies that test whether claims and analytical results can be independently checked or reproduced.
Sources
- CGMacros: a scientific dataset for personalized nutrition and diet monitoring. PhysioNet v1.0.0. Dataset.
- Hall H, Perelman D, Breschi A, et al. Glucotypes reveal new patterns of glucose dysregulation. PLOS Biology. 2018;16:e2005143. Article and supplementary standardized-meal data.
- Toumi M, Salathé M. Short-term postprandial glucose monitoring reveals stable traits from noisy free-living meals. medRxiv. 2026. Preprint.
- Food & You analysis repository. Code and derived results.
Research status: exploratory reanalysis with a preregistered-style falsification specification for the next independent dataset. Results reported as our analysis should be considered provisional until independently reproduced.