Same Bread, Five Verdicts

Part 1 of a series on the meta-NPS. Five expert systems score one loaf of whole-wheat bread five different ways. Here is how we turn that disagreement into a measurement, and why it matters for policy.

meta-NPS
NPS
IAFNS
FRESH
interactive
series

Five expert-built food-rating systems score the same whole-wheat bread anywhere from the 12th to the 99th percentile of the US food supply. This first part of the meta-NPS series walks through what a rating system is, why the five disagree, and how a meta-rating turns that disagreement into a number you can act on.

Authors

Josh Erndt-Marino, PhD

FRESH’s AI assistant

Published

June 23, 2026

The short version. Part 1 of the Meta-NPS series. Claims link to where they are worked out.

Five expert-built food-rating systems score the same loaf of whole-wheat bread anywhere from the 12th to the 99th percentile of the US food supply.

  • They are not measuring the same thing. Each encodes a different definition of healthy — nutrients to encourage and limit, food groups, degree of processing, bioactives. The disagreement is not error; it is a disagreement about the construct.
  • The fight is mostly about processing. The systems that account for how heavily a food is processed separate sharply from the ones blind to it. That single axis explains most of the spread on this bread.
  • Every score here is a percentile — a food’s rank against roughly 6,450 US foods — so five systems can be compared on one axis at all.
  • The move is not to pick a winner. It is to treat the disagreement as data: run all five, report the middle with its spread, and let the width of the spread carry the uncertainty instead of hiding it.

Why it matters. A label, a policy, or an app that quotes one system is quoting one opinion with the confidence of a measurement.

Part 2 asks how often the five split across the whole carbohydrate supply, and where.

The prevailing story about nutrition goes something like this. We basically know what is healthy. The hard parts are downstream: getting people to act on it (the behavioral environment), and building places where the healthy choice is available and affordable (the physical, or food, environment). On this account the science of what is largely settled, and the real work is delivery.

It is a reasonable story, and it is mostly right at the level of slogans. Eat more vegetables. Cut back on added sugar. The trouble starts the moment you have to be specific, because labels, “healthy” claims, school and hospital procurement, reformulation targets, and diet apps all run on the specific version, not the slogan. They need a verdict on this food.

Here is one loaf of whole-wheat bread that shows how quickly the specific version falls apart.

What a rating system actually is

Most people have never heard of a nutrient profiling system, or NPS, even though their food choices are increasingly shaped by them. The short version: take something like a Nutrition Facts panel, run it through a set of expert-chosen rules about what to reward and what to penalize, and collapse the result into a single score. That is the entire move.

Nutrition Facts Calories Total fat Sodium Fiber Sugars Protein Expert-chosen rules + reward fiber, protein, nutrients, food groups − penalize sugar, sodium, saturated fat, processing …each with a chosen weight one score 0 – 100

An NPS is a formalized, transparent, reproducible expression of one expert group’s definition of “healthy.” Change the rules, change the weights, or change which features you look at in the first place, and you change the score. There is no underlying true value the rules are estimating badly. Each set of rules is a definition.

Five definitions, one bread

Researchers have published more than 200 such systems, plus a long commercial tail. For this series we use five widely cited ones. They were built by serious people, and none is wrong on its own terms. They simply encode different definitions.

Nutrient density + food groups

NRF 4:3:3

Frontiers in Nutrition · 2020 New Nutrient Rich Food Nutrient Density Models That Include Nutrients and MyPlate Food Groups Drewnowski & Fulgoni

Adds up beneficial nutrients and MyPlate food groups, subtracts the nutrients to limit. Rewards nutrient-dense whole foods. Says nothing about processing.

Read the paper →
Nutrient density + flavonoids

NRF 9f.3

Journal of Nutrition · 2009 Development and Validation of the Nutrient-Rich Foods Index: A Tool to Measure Nutritional Quality of Foods Fulgoni, Keast & Drewnowski

Nine terms to encourage minus three to limit, per 100 kcal. Extends the classic nutrient-density index of the linked paper with a total-flavonoid term. Also blind to processing.

Read the paper →
Front-of-pack grade

Nutri-Score

Int. J. Vitamin & Nutrition Research · 2022 The Nutri-Score nutrition label: a public health tool based on rigorous scientific evidence Julia & Hercberg · adapted from the British FSA/Ofcom model

Balances energy, sugar, saturated fat and salt against fiber, protein and fruit/veg into an A–E color grade. In use across Europe.

Read the paper →
Composite · weighs processing

Food Compass

Nature Food · 2021 Food Compass is a nutrient profiling system using expanded characteristics for assessing healthfulness of foods Mozaffarian, El-Abbadi, O’Hearn et al.

Scores nine domains at once, including additives and degree of processing, on a 1–100 scale. Broad in scope, and openly contested.

Read the paper →
Processing only

FPro

Nature Communications · 2023 Machine learning prediction of the degree of food processing Menichetti, Ravandi, Mozaffarian & Barabási

A machine-learning estimate of how processed a food is, from whole to ultra-processed. Ignores nutrients entirely.

Read the paper →

These five are a sample. With 200-plus systems in the literature, any comparison sees a fraction of the field. Each system is also a living, partial definition: it captures some of what people mean by “healthy” and leaves the rest out. Keep that in mind when the numbers below start to look large. They are a floor.

Same bread, five verdicts

So: one loaf, five definitions. Watch what happens when each one scores it. The walk-through builds itself; let it run, or step through with the arrows.

Every score here is a percentile: a food’s rank against roughly 6,450 US foods. That percentile is just the common ruler that lets five very different systems be compared on one axis. The number to watch is the score.

The middle of the five scores is the meta-score, a single stand-in for the loaf. How tightly the five agree is its stability. The FRI (Food Recommendation Index) folds the two together: a food with a respectable middle can still be marked down when the experts split on it, and whole-wheat bread is exactly that case. Its scores run from 12 to 99, so the meta-rating does not call it “so-so.” It calls it Uncertain, one of four verdicts the method can return: Certainly Healthy, Certainly Neutral, Certainly Unhealthy, or Uncertain. “Uncertain” is not a hedge. It is a measured statement that the experts genuinely do not agree.

The fight was about processing

The disagreement on bread is not noise. It has a source, and the meta-rating can point straight at it.

Two of the five systems weigh one extra thing the others ignore: how processed the food is. Set those two aside, and the remaining three nearly agree (69, 93, 99). The meta-score climbs from 74 to 93, the spread collapses, and the verdict flips from Uncertain to Certainly Healthy. Same loaf, opposite reading. The entire fight was about processing.

That does not settle whether processing should count. It does something more useful: it locates the disagreement precisely, and tells NPS designers exactly what they would need to agree on (or transparently disagree about) to stop producing this particular split. Treating the disagreement as information, rather than choosing a winner among the systems, is the whole idea of a meta-rating.

Five systems, five food supplies

Zoom out from one loaf to the whole supply, and the disagreement gets harder to wave away. Each system draws its own picture of how healthy American food is.

The food supply does not have one healthiness shape. It has five, depending on which expert you trust. Run the meta-rating across all of them and the contested zone is not a rounding error: about one in five carbohydrate-rich foods, and roughly half of all grains, fall into “Uncertain,” where the experts cannot agree.

Those numbers are a floor, not a ceiling, and conservatively so on four counts:

  1. We sampled only 5 of 200-plus published systems. More systems, more disagreement.
  2. We used one way of drawing the cutoffs (spread and rank-range). A different defensible rule flags more foods.
  3. We held each system to one faithful implementation, when a single system can contradict itself on something as basic as how to normalize its inputs. The disagreement also lives inside one expert.
  4. Every system here is a universal rater: one score per food, the same for everyone. None account for the eater (physiology, goals, context). Healthfulness that depends on the person is disagreement these numbers never even attempt to measure.

Even under all four conservative choices, the disagreement is this large.

Where the leverage actually is

Step back to the prevailing story we opened with. We mostly know what is healthy; the work is behavior and access. Bread suggests a third environment the story leaves out: the informational one. Do we even agree on, and clearly communicate, what “healthy” is? On the high-stakes, food-by-food questions that policy actually runs on, we frequently do not, and we have no shared way to show the public how settled or contested any given verdict is.

That third environment is the neglected, lower-hanging, higher-leverage lever. Harmonizing definitions and communicating uncertainty is faster and cheaper than rebuilding food access or shifting population behavior, and it feeds both of the others. Clearer definitions make for better labels, procurement, and reformulation (the physical environment). Honest, uncertainty-aware information starves the misinformation vacuum and gives behavior something solid to stand on.

Two honest limits keep this from becoming a slogan of its own. Information enables; it does not cure. Healthy foods are not the same as healthy diets, and no rating system fixes the gap between the two. And harmonizing the universal disagreement is only the first move. The second is personalizing beyond it, since “healthy for whom” is a question none of these systems ask. Harmonization is the floor; personalization is the ceiling. Both are informational-environment work, and both are buildable now.

Whole-wheat bread’s five scores are real, drawn from the published systems run across the US food supply. This is the first of a three-part series. Part 2 takes the meta-rating across the whole supply to ask where the disagreements cluster. Part 3 brings in consumers, and what happens when expert verdicts and public belief pull apart. The webinars and papers behind all of it are on the resources page.

Get new essays by email

One a week. Sign up on freshfoodrecs.com — free.

Subscribe

Thoughts on this piece? Send feedback — it goes straight to us.