We Measure Sickness, We Mean Health

The health gap is the most consequential thing nutrition, medicine, and wellness have never agreed how to measure. Naming it is overdue.

health-gap
framework
thought-leadership
FRESH

There are nearly 75,000 codes in ICD-10-CM for what’s wrong with you. There are roughly zero agreed metrics for what’s right with you. That asymmetry quietly shapes the health outcomes we all claim to be chasing — and almost nothing has been built to close it. First in a series.

Authors

Josh Erndt-Marino, PhD

FRESH’s AI assistant

Published

May 18, 2026

The short version. Claims here link to where they are worked out in full.

There are 74,719 codes in ICD-10-CM for what is wrong with you. There is no agreed metric for whether you are healthy. That asymmetry is the health gap.

  • A metric is an agreement, not a discovery. It is how a society writes down what it has decided to care about, precisely enough that strangers can act on it together. Every one is an imperfect representation, and that is the cost of being comparable at all. What goes unstandardized does not stay neutral, it goes to whoever is most confident. So the missing health metric is really a missing agreement.
  • It is layered. The disease side has built a great deal and deployed little — cardiovascular risk assessment alone has more than a thousand published algorithms, and most clinicians can name one or two. The health side has built almost nothing to deploy.
  • Inverting a risk score does not give you health. A 95% probability of not having an event is not vitality, reserve, or resilience. The vocabulary collapses to absence-of-disease the moment you press on it.
  • Disease measures thresholds; health is a gradient. Two 30-year-olds at 118/76 and 105/65 are both “not hypertensive.” They are not equally healthy, and the disease vocabulary cannot tell them apart. Most people live in that range, and it has no agreed instruments at all.
  • The near-misses are laundered. Biological age clocks aim at the positive pole but are largely validated against disease-mortality outcomes. Food scoring has 200+ systems that disagree by ninety percentile points on the same loaf of bread.

What any defensible answer would need: explicit operationalization — imperfect is fine, vague is not; uncertainty carried in public; independence from disease prediction as its sole justification; legibility to the people it is for, in a country where a third of adults sit at or below the lowest measured level of numeracy; an institutional patron willing to fund the question on its own merits; and no pretending we are further along than we are.

Where this leaves FRESH. It does not close the gap. It tries to work honestly inside it — treating disagreement between rating systems as data rather than noise, and showing the structure of what we do not know.

Information can be health. But only if we can agree what health is.

Part 2 follows the gap downstream: We Measure Harm, We Mean Benefit.

There are 74,719 codes in ICD-10-CM as of last October1 — the diagnostic vocabulary US clinicians use to find what is wrong with you. You can be diagnosed with septic shock (R65.21), with frostbite of an unspecified ear (T33.011A), with an ingrown nail (L60.0). The system is built to find sickness with extraordinary granularity. Every test, every lab value, every clinical guideline turns is this person sick? into a binary you can bill against.

There is no agreed-upon metric for whether you are healthy.

Not for healthspan. Not for “physiological vitality.” Not for whatever the CDC means when it defines health equity as “the state in which everyone has a fair and just opportunity to attain their highest level of health.” Highest level of what, exactly? Measured how? Compared to what reference? The definition presumes a measurement that does not exist.

I posted a version of this on LinkedIn in March:

“We have TONS of ways to tell if you’re diseased. You’ve probably already been told that. But health? That piece is still elusive. Why?”

I posted a version of the same question in March 2024, fourteen months earlier:

“How do we want to define and measure health? Why hasn’t this been undertaken yet?”

And again last March, framing through an essay by John Singer:

“What would it mean to build a platform for the production of health, not the management of disease?”

I keep asking it. The question keeps not being answered.

What a metric is for

Before the gap, the thing the gap is made of.

A metric is not a description of the world. It is a statement of what a society has agreed to care about, written down precisely enough that strangers can act on it together. Those 74,719 codes are not a discovery about bodies. They are a decision, accumulated over decades, about what is worth naming, tracking, paying for, and arguing over. Somebody sat in a committee and concluded that frostbite of an unspecified ear earns its own entry. That was a judgment about what matters, and it is now infrastructure that shapes what several hundred thousand clinicians are able to see.

Which means the missing health metric is not really a missing number. It is a missing agreement. We have never collectively said what health we want, in terms specific enough that anyone could be held to them, and the empty space where the metric should be is the evidence that we never had the conversation.

Having that conversation would not produce a perfect instrument, and it does not need to. Every metric is an imperfect representation, and it has to be. Measurement compresses something complicated into something comparable, and the compression always loses. But comparability is the entire point of the exercise. A standard is what lets a clinician in Ohio, a researcher in Lisbon, and a regulator in Brussels find out that they disagree, and about what. Without one, a disagreement cannot even be located, let alone settled. Standardization is not a claim to have captured the truth. It is the machinery that makes collective correction possible, which is the only mechanism any of these fields has ever had for getting less wrong over time.

What goes unstandardized does not stay neutral in the meantime. It goes chaotic, and we already know what that chaos looks like, because the health side of the ledger has been living in it the whole time. Two hundred ways to score a food that disagree by ninety percentile points on the same loaf of bread. Wellness marketing where “healthy” means whatever the seller needs it to mean this quarter. A family at one dinner table holding four contradictory ideas about what a good diet is, each borrowed from a different source, none of them checkable against anything.

The strongest objection to this entire essay is that health is too rich, too personal, and too multidimensional to reduce to a number. The objection is correct. Health is all of those things, and any metric will flatten some of it. But that is an argument about which metric, not about whether. A rich thing measured roughly can still be compared, audited, and revised by people who were not in the room when it was built. A rich thing left unmeasured just gets handed to whoever is most confident. Imperfect cannot be the enemy of good here, because the alternative on offer is not something more faithful than a number. It is the noise we already have.

The asymmetry

We have built an extensive measurement literature for what is wrong. We have built almost nothing for what is right. And before I lean too hard on what we’ve built for the disease side, an honest qualifier: even the extensive disease literature sits largely unused. Most clinicians could name a published risk algorithm or two if pressed; most researchers outside specialty cardiology have never engaged with the methodology. The accumulation of published tools and their absence from actual clinical workflows is itself a long-running diagnostic, flagged for years by critics like the Sensible Medicine community. I have made some version of this observation in my own posts: we’ve had risk tools for >25 years; they’re accumulating on the shelves in the thousands.2

So the asymmetry is layered. The disease side has built a lot and deployed a little. The health side has built almost nothing to deploy.

The asymmetry shows up everywhere I look. Cardiovascular risk assessment has more than a thousand published algorithms — 1,382 of them sit in the Tufts PACE registry, and 58% have never been externally validated even once3 — Framingham, QRISK, SCORE2, the Pooled Cohort Equations, PREVENT. Most clinicians can name one or two. The rest live on academic shelves, most of them never checked against a second dataset by anybody. When PREVENT replaced the Pooled Cohort Equations in 2023 and halved baseline risk estimates overnight, the change moved through specialty cardiology with little fanfare and through general primary care with even less. You can mathematically invert any of these algorithms — your 5% 10-year risk is a 95% probability of not having the event — and call that inverted score a probabilistic view of your cardiovascular health. The inversion is technically valid. It also exposes the deeper problem: a probability of not having a specific event over a specific window is not cardiovascular vitality, or functional reserve, or resilience to perturbation. The vocabulary collapses to absence-of-disease the moment you press on it. The thousand algorithms also do not agree with each other. We have a thousand parallel disease-absence measures, mostly unused, none agreed upon, none of which is actually a measure of cardiovascular health in any positive sense.

You might reasonably ask: isn’t absence of disease a prerequisite for vitality, reserve, or resilience in the first place? Yes. Absence of disease is necessary. The disease-measurement infrastructure does real work — it tells you when you have crossed a threshold that means something has gone wrong, and crossing the threshold rules out most of the positive states by definition.

But necessary is not sufficient. Disease infrastructure measures threshold crossings. It does not measure the dynamic range above the threshold. The space where most “healthy” people actually live — between “no diagnosable condition” and “optimally thriving” — has no agreed measurement infrastructure at all. A 30-year-old whose blood pressure runs 118/76 and a 30-year-old whose blood pressure runs 105/65 are both “not hypertensive.” They are not equally healthy. The disease vocabulary cannot tell them apart. The same is true for HRV, VO2 max, lipid profiles, fasting glucose, sleep architecture, inflammation markers — every variable where disease has a threshold and health has a gradient. We have agreed on the threshold. We have not agreed on what to do with the gradient.

This is the shape of the gap. Disease measurement covers half the dynamic range — the half below the threshold. The other half — the positive-health side of the threshold, where most of the population sits — is where we have no agreed vocabulary, no agreed instruments, and no institutional patron for building either.

Biological age clocks come closer to the positive side than risk scores do. The underlying research question is legitimate — how do you measure how healthy someone is on the inside? — and the construct at least attempts to point at the positive gradient. But press on the validation and most of these clocks are trained against disease-mortality outcomes. Positive-health framing laundered through disease-prediction inputs. And the commercial proliferation has run miles ahead of the analytical foundation: a paper claiming a “strong association” from a hazard ratio of 1.055 gets accessed ten thousand times in a month. The translation, in my reading, is borderline awful.

The food-rating systems I have been working on for years have the same problem from a different angle. There are 200+ ways to score the healthfulness of a food in the literature. They disagree on the same loaf of whole wheat bread by ninety percentile points. They are measuring something, but not the same thing, and certainly not health-in-the-abstract.

The Dietary Guidelines repeatedly invoke “health equity” and “nutrient density” without offering formal definitions of either. The CDC’s health-equity statement presumes that attainable-health is a knowable quantity. The “production of health” language requires a productive measure to operationalize.

We are stacking second-order claims on top of a missing first-order measurement. The whole vocabulary of healthcare, nutrition, and wellness is built on a number nobody has agreed how to compute.

Why the gap is still open

Three guesses. I do not know which is most right.

One. We built the infrastructure for what could be billed. Disease has a payer. Disease has a billing code. Disease has a clinical trial endpoint. Health has none of those things. The market never had to organize around it. There has never been an institutional patron for is this person flourishing? the way there is one for does this person have cancer?

Two. Health may be genuinely harder to measure than disease. Disease is the perturbation; health is the equilibrium. We are better at noticing what is broken than describing what is whole. This might be a fundamental cognitive asymmetry, or it might be an artifact of the measurement tradition we inherited. I am not sure.

Three. The people best-positioned to build a health metric have weak incentives to do so. Academic publication rewards novel mechanisms of disease. Drug pipelines reward treatments of disease. Clinical trials reward endpoints related to disease. Even prevention research — the area closest to a positive health agenda — gets evaluated by how many cases of disease it prevents. A health metric would have to compete with the metrics that already work for everyone whose career and funding depend on those metrics.

None of these guesses gets less true if we keep not naming the gap.

Naming it

I have been calling this the health gap. It is a clunky name — easily confused with health-equity gaps, healthspan gaps, the access-to-care gap, every other “health ___ gap” in the literature. Better names are invited.

What I want the name to do is make the asymmetry visible. We have a measurement infrastructure for one half of what we say we want. We have almost nothing for the other half. The space between them is where most consumer confusion, most policy thrashing, most well-meaning-but-ineffective wellness investment, and most of nutrition’s perennial fights actually live.

You cannot solve a problem you cannot name.

What it would take to start closing it

I do not have a research program for this in my back pocket. But I can name conditions that any defensible answer would have to satisfy.

Explicit operationalization. Not “health” as the absence of disease, not health as a vague proxy. A specific construct, with specific subcomponents, that can be measured with specified instruments and aggregated by stated rules. Imperfect is fine. Vague is not.

Honest uncertainty. Health is a “the one and the many” problem before it is an empirical one. Any rigorous health measure must carry its uncertainty publicly — not hide it under summary statistics that pretend to more precision than the underlying construct allows.

Independence from disease prevention as the sole justification. A health metric that only matters because it predicts disease incidence is a disease metric in disguise. That is where most “wellness” measurements collapse on inspection. The metric should mean something on its own terms.

Legibility to the people it is for. A metric is a coordination device, and a coordination device nobody can read coordinates nothing. This is the condition I see skipped most often, and the numbers on it are not comfortable. In the 2023 international adult skills survey, 34% of US adults scored at or below Level 1 in numeracy, against an OECD average of 25%, and the share at or below Level 1 in literacy rose from 19% to 28% in six years.4 The last time anyone measured US health literacy at national scale, in 2003, the task of determining a healthy weight range from a graph relating height and weight to BMI was scored as Intermediate difficulty, which is to say beyond roughly half the adult population.5 That is a federal agency putting a difficulty score on an artifact the nutrition field hands out for free.

The reflex is to call this a deficit in the audience. It is mostly a deficit in the packaging, and there is a clean demonstration of the difference. Given a standard screening problem in conditional-probability form, 160 gynecologists picked the right answer 21% of the time, slightly worse than guessing. Given the same facts restated as natural frequencies, 87% of them got it right.6 Same people, same room, same afternoon. Nor is the professional half of this reassuring on its own terms: in a national survey of US primary care physicians, 47% said that finding more cancers in screened populations proves that screening saves lives, which it does not.7 Meanwhile patient education materials in high-impact journals run at a mean reading level of grade 11 to 14, and 2% of them meet the American Medical Association’s own recommendation.8 Whatever health metric eventually gets built will be shipped into that. If it is designed the way we design everything else, it will be legible to the people who built it and to almost nobody else.

Institutional patronage. NIH will not fund the basic question without disease endpoints. NSF does not see it as basic science. Foundations chase legible outcomes. Industry wants something it can sell. Somebody has to commission this on its own merits.

Not pretending we are further along than we are. Almost every public conversation about health prediction — biological age clocks, healthspan trackers, longevity supplements, FRESH-style food scoring — is operating in advance of the foundational measurement. The honest move is to say so out loud.

FRESH is one attempt

I have been building FRESH for three years on the food side of this. It does not solve the health gap. It tries to operate honestly inside it — by treating disagreement among food-rating systems as data rather than noise, by carrying uncertainty in the output rather than hiding it, by being explicit that we do not yet know what “healthy” means at the food level and showing readers the structure of that not-knowing.

If a health metric ever gets defined, FRESH-style methodology becomes more useful, not less. We will still need ways of asking is this food consistent with what we have measured as healthy? — and we will still need to do it with explicit uncertainty.

But FRESH is one attempt. There need to be more. The work is too important to be done in one project, or one paper, or one company.

The closer

I will keep posting some version of this question for as long as it stays unanswered. Each time I do, more people who have also been circling it reach back. The gap is real. The people who could close it are scattered across nutrition, medicine, public health, wellness tech, and philosophy of science. They mostly do not talk to each other.

If you have been working on this — building a metric, designing a study, writing a critique, thinking through what health even means as a measurable thing — I want to hear from you.

Information can be health. But only if we can agree what health is.

Agreements start with value alignment. Our kids are not taught what a value system is — philosophically or practically. We were not taught either. I have bet before that >95% of PhDs in my adjacent fields never took a philosophy course.9 The methodological gap I have been describing sits on top of an educational one: we have not built the scaffolding that would make the conversation we need to have possible.

Why are we still not asking those questions seriously?


This is the first in a planned series naming the health gap and what it would take to close it. The next post examines the most quantitatively striking case of the gap in action: how one cardiovascular risk equation, replacing another in 2023, may have silently halved the value ceiling for every CVD intervention overnight.

Related on LinkedIn:

Footnotes

  1. Correction, 2026-08-01. This essay published with 74,260, which is the FY2025 count. “As of last October” means FY2026, effective 1 October 2025, which is 74,719. Both figures are billable-code counts taken from the CMS icd10cm_order files for the respective years. The description line said “roughly 70,000” — a rounding that had drifted well below the actual number — and now says nearly 75,000.↩︎

  2. From January 2025: https://www.linkedin.com/feed/update/urn:li:activity:7281331846217437184/. The full line: “We’ve had risk tools for >25 years. They’re accumulating on the shelves in the thousands. Why? My bet (other than incentives) is education.”↩︎

  3. Wessler, Nelson, Park, et al., “External Validations of Cardiovascular Clinical Prediction Models: A Large-Scale Review of the Literature,” Circulation: Cardiovascular Quality and Outcomes 2021;14(8):e007858, DOI 10.1161/CIRCOUTCOMES.121.007858 — 2,030 external validations of 1,382 CPMs in the Tufts Predictive Analytics and Comparative Effectiveness registry, with 807 of them (58%) never externally validated. (Citation upgraded 2026-08-01; this claim originally rested on my own June 2024 LinkedIn post, which is not a source.) Note the denominator: 1,382 is cardiovascular CPMs of all kinds. The narrower set of general-population risk models is smaller — Damen et al., BMJ 2016;353:i2416 found 363 — so “more than a thousand” is right for the field and would be wrong for Framingham-style tools specifically.↩︎

  4. OECD, Survey of Adult Skills 2023, United States country note, and NCES, Highlights of the 2023 U.S. PIAAC Results (NCES 2024-202). US adults averaged 249 in numeracy, below the OECD average, with 34% at or below Level 1 against an OECD average of 25%; the share at or below Level 1 rose from 29% to 34% in numeracy and from 19% to 28% in literacy between 2017 and 2023. Two caveats worth carrying: the overall response rate was 28%, and NCES itself warns that comparisons across cycles need care because the framework was revised, administration moved to tablet-only, and basic-skills items counted toward overall scores only in 2023, which mechanically moves some adults down a level. OECD reports ages 16–65 and NCES sampled 16–74, so the denominators are not interchangeable.↩︎

  5. Kutner, Greenberg, Jin & Paulsen, The Health Literacy of America’s Adults: Results From the 2003 National Assessment of Adult Literacy, NCES 2006-483, free PDF. The BMI-graph task is scored at 290 on the 0–500 health-literacy scale, inside the Intermediate band (226–309); 12% of adults reached Proficient and 53% Intermediate, so “beyond roughly half” is my reading of a task sitting in the upper part of the Intermediate range, not a figure the report states. Two things I am not doing with this source: NAAL’s own category names are Below Basic, Basic, Intermediate and Proficient, and the report never calls Basic “inadequate,” so the widely quoted “36% have limited health literacy” is a derived sum with an editorial gloss on top. It is also 23 years old and has never been repeated, which is its own comment on how much we wanted to know.↩︎

  6. Gigerenzer, Gaissmaier, Kurz-Milcke, Schwartz & Woloshin, “Helping doctors and patients make sense of health statistics,” Psychological Science in the Public Interest 2007;8(2):53–96, free PDF. The 21% and 87% figures are from a continuing-education session with 160 German gynecologists, not a survey with a sampling frame, and the question was multiple-choice, which is what makes 21% “slightly less than chance.” I am using it as a demonstration that format carries the failure, which is what it demonstrates, rather than as a population estimate.↩︎

  7. Wegwarth, Schwartz, Woloshin, Gaissmaier & Gigerenzer, “Do physicians understand cancer screening statistics? A national survey of primary care physicians in the United States,” Annals of Internal Medicine 2012;156(5):340–349, PMID 22393129. 412 physicians across two waves from a research panel; 69% recommended a test on statistically irrelevant evidence versus 23% on relevant evidence. The scenarios were hypothetical rather than observed practice, and the data are from 2010–11.↩︎

  8. Rooney, Santiago, Perni, et al., “Readability of patient education materials from high-impact medical journals: a 20-year analysis,” Journal of Patient Experience 2021;8, PMC8205335. 2,585 materials, mean grade level 11.2 to 13.8 across seven metrics; 2.1% met the AMA’s sixth-grade recommendation and 8.2% the NIH’s eighth-grade one. Readability formulas count syllables and sentence length rather than conceptual difficulty, and journal materials are a narrow slice, so this is evidence of a mismatch rather than a measure of comprehension.↩︎

  9. The “95% of PhDs never took a philosophy course” bet is from a January 2025 post: https://www.linkedin.com/feed/update/urn:li:activity:7288609024235720704/. “Probably <50% took a stats course. Statistics and philosophy are tightly related, but training in either is very uncommon.”↩︎

Get new essays by email

One a week. Sign up on freshfoodrecs.com — free.

Subscribe

Thoughts on this piece? Send feedback — it goes straight to us.