The Label Can’t Tell You
One measurement can’t settle whether a food is healthy, and one expert can’t either. At both levels the fix turned out to be the same: don’t choose.
Nutri-Score cannot be calculated from a Nutrition Facts panel, and neither can the other systems used to rate foods. Closing that gap means joining eight USDA and FDA sources that do not share a key — and once every system has what it needs, they turn out to disagree about half the shelf.
The short version. Claims link to where they are worked out in full.
Nutri-Score cannot be calculated from a Nutrition Facts panel. Neither can the other systems used to rate foods. The label does not carry what they need.
- The data exists, in pieces. Closing the gap means joining eight USDA and FDA sources that do not share a key — walk that path and the fragments collapse into one record per food.
- Then the systems still disagree. Once every rating system finally has what it needs, they split on about half the shelf.
- Same problem, two levels. No one measurement settles whether a food is healthy, and no one expert does either. At both levels the fix turned out to be the same: don’t choose. Integrate the measurements, run all the experts, and report the disagreement.
- Which makes “Uncertain” a real verdict. Report the middle with its spread and a large share of products honestly resolve to uncertain — not because the analysis failed, but because the experts do not agree closely enough to place them.
The point. A single number on a package implies a settled question. For about half of what is on the shelf, it is not settled — and saying so is more useful than picking a side.
Ask whether a packaged food is healthy and you have asked a question no single measurement answers. The answer depends on nutrients, on food groups, on bioactives, on how heavily the food is processed, on what contaminants it carries, and on how much of it people actually eat.
It is not a question any single expert answers either. More than 200 food-rating systems have been published, and they disagree.
Same problem, two levels. It turned out to have the same answer at both.
No one measurement
Nutri-Score is the simplest of the five systems we work with, and the one printed on packaging across Europe. It needs energy, sugars, saturated fat, sodium, fibre and protein — all of which sit on a Nutrition Facts panel — plus the fruit, vegetable, nut and legume fraction of the product.
That last term is mandatory. It is not a nutrient, it is a food-group quantity, and it never appears on packaging. Across 80,000 label cells I checked, it appeared zero times.
So Nutri-Score cannot be computed from a label. Not computed poorly — not computed. And Nutri-Score is the least demanding of the five: it asks for seven inputs, NRF 4:3:3 for ten, NRF 9f.3 for twelve, Food Compass for fifty-four, FPro for fifty-eight. The panel does not grow to meet them.
Pool everything the five systems ask for and you get 236 distinct attributes of a food. A package is a small document. Here is how small.
The other attributes are not printed on any package, and no amount of squinting at one will produce them. They are in federal databases — which is where the second problem starts.
USDA has the answers, in eight places
The data exists. It just sits in eight separate USDA and FDA sources — survey foods, reference commodities, food-group equivalents, flavonoids, contaminants, iodine — on three keys that do not join. Food-group values exist twice, once per vocabulary. Flavonoids exist twice the same way. A question as ordinary as how much whole grain, and how much quercetin, is in this food crosses at least two vocabularies, and joining on food name is not a join.
The join is possible because USDA already documents how its survey foods are built out of reference commodities. That existing link works in both directions — the deeper commodity nutrient panel rolls up to the survey food, and the food-group and flavonoid axes push down onto the ingredients.
Walk it and eight fragments collapse into one record per food.
That record is also honest about its own holes. Where USDA has genuinely never measured something — vitamin D isoforms, menaquinone-4, fluoride, individual tocotrienols — the gap gets declared rather than filled, because integration can propagate a measurement but it cannot create one.
The methodology and benchmark summary draws the line this whole essay rests on. FRESH Composition — the integration and the matching, versioned, checksummed and deliberately judgment-free — sits underneath FRESH Food, the interpretive layer where the five expert systems, the Meta-Score and its stability, and the four verdicts live. Six figures on how the sources are joined, how the matches are benchmarked, and what the pipeline does not claim.
No one expert
With that record in hand, holes and all, the systems can finally run. One of them could not have run at all before it: NRF 9f.3 extends the classic nutrient-density score with a flavonoid term, and flavonoids were one of the two axes sitting in two vocabularies. Score whole-wheat bread — a food most people would call straightforwardly healthy — against roughly 6,450 US foods, and ask all five where it belongs.
Same loaf, near the bottom of the food supply and near the top of it, depending on who you ask.
That spread is not noise. FPro marks the loaf down for being an industrially produced bread; NRF 4:3:3 rewards its nutrient density. Both are behaving correctly, because they are measuring different things. Collapse them into one confident number and you have destroyed the only information that mattered. Report the middle with its spread and the honest answer is Uncertain.
That argument is the subject of Same Bread, Five Verdicts, part one of this series, and The Disagreement Has a Shape extends it across the whole carbohydrate supply.
Half the shelf
Extending it to the shelf is what the integration finally allows. Run the whole stack on 666 matched retail products and count how often the five land close enough together to say anything at all.
The FRESH-Food product surface opens all 666 of them: label image, ingredient statement, the USDA food each was matched to, its intake tier, and where all five systems placed it. The 51% is easier to believe once you have watched two respected systems split over a cereal you have eaten.
That verdict gap counts as a finding only because of the integration underneath it. Fill a missing composition value quietly and the disagreement stops being a property of the systems and becomes an artifact of the filling.
Two levels, one shape. Don’t pick a measurement — integrate them. Don’t pick an expert — run them all and report how far apart they land. What falls out is uncomfortable: for half of a supermarket shelf, the most honest label we could print is nobody knows yet.
Resources
- The FRESH-Food product surface — the 666 matched products, one card each, with the label, the matched USDA food, and the five-system spread.
- The methodology and benchmark summary — how those matched products got their numbers: the integration architecture, the matching and its benchmarks, and what the pipeline does not claim.
- Same Bread, Five Verdicts — part one: why five systems split on one loaf.
- The Disagreement Has a Shape — part two: the same split across the carbohydrate supply.
The published framework
Erndt-Marino J, O’Hearn M, Menichetti G. An integrative analytical framework to identify healthy, impactful, and equitable foods: a case study on 100% orange juice. International Journal of Food Sciences and Nutrition. 2023;74(6):668–684. doi:10.1080/09637486.2023.2241672
Erndt-Marino J, Ghirardelli A. Educational gaps and communication priorities for the healthfulness of carbohydrate foods. Journal of the American Nutrition Association. 2026. doi:10.1080/27697061.2026.2687436
Comments