Skip to content
The Supar Health library·Pillar 03 of 05·Science

How biological age tests actually work - and why most are limited.

The category of "biological age testing" has expanded dramatically over the past decade. Each new product comes with marketing copy that suggests a number representing how old your body really is. The mechanics behind these tests, and the actual scientific limitations, are rarely communicated honestly. This article is our attempt to do that.

Reviewed17 May 2026
Reading time~13 min
Maintained bySupar Health scientific team
SeriesThe clinical library, 03 of 05

There are now methylation-based clocks, glycan-based clocks, blood biomarker composites, multi-omic models, AI-derived scores, and dozens of consumer products built on top of these.

We sell a biological age test. We also believe that the people taking biological age tests deserve a clearer picture of what they are getting than the marketing of this category typically provides. What follows is a clinician-grade explanation of how biological age tests are built, what they measure, where they work well, and where they break down. We have tried to present this honestly enough that it would still hold up if you read it before deciding which test to use - including if you decided not to use ours.

01What "biological age" actually is

Let us start with the most important conceptual point: biological age is not a clinical endpoint.

There is no biological measurement that directly reads out the age of a human body. You cannot put a number on how "old" a heart is or a brain is or a cell is in any direct sense. Cells have many features that change with age - DNA methylation patterns, telomere length, mitochondrial function, glycation products, protein damage, cellular senescence markers - but none of these is the age of the cell. They are features that correlate with age.

What every biological age test actually does is the following:

  1. Measure one or more biological features that change in patterned ways with age
  2. Use a statistical model, trained on a dataset of people of known chronological ages or known clinical outcomes, to translate the measurements into an "age-equivalent number"
  3. Present that number to you as your biological age

This is a useful exercise. It gives you a comparison point against population norms. It can highlight whether you appear to be aging faster or slower than your chronological peers. It can support a coherent intervention strategy and a longitudinal tracking practice.

What it cannot do is directly tell you how old your cells are. That number does not exist as a measurable physical quantity.

This is not a criticism of any specific test. It is a clarification about what the entire category is.

02The three things that matter in a biological age test

Once you understand what biological age is, you can evaluate any test by three properties.

Property 1 - What is the underlying measurement?

Every biological age test rests on something physical that is actually being measured in a laboratory. The choice of underlying measurement is the single most important design decision in the test.

Some examples:

  • Horvath and Hannum clocks: DNA methylation patterns at specific CpG sites in your genome
  • GrimAge: DNA methylation patterns, plus methylation-inferred levels of several plasma proteins and smoking history
  • DunedinPACE: DNA methylation patterns, but the model is trained on longitudinal aging rate rather than snapshot age
  • GlycanAge: patterns of sugar molecules attached to immunoglobulin G in blood
  • Multi-biomarker panels: dozens of blood biomarkers combined statistically
  • suPAR: a single circulating protein measured directly by clinical immunoassay

These are categorically different measurements. Their analytical characteristics - how reproducible they are, how stable they are over time, how sensitive they are to assay conditions - differ substantially.

A test built on a noisy underlying measurement will inherit that noise no matter how clever the statistical model on top. A test built on a stable, well-characterized underlying measurement will be inherently easier to interpret at the individual level.

Property 2 - What endpoint was the model trained on?

The number a test gives you depends entirely on what its statistical model was trained to predict.

  • If the model was trained on chronological age (as in Horvath and Hannum), the output approximates chronological age. Deviations from chronological age are interesting but indirect indicators of health.
  • If the model was trained on biomarker composites (as in PhenoAge), the output approximates a biomarker-defined age that correlates more directly with clinical risk.
  • If the model was trained on long-term outcomes (as in GrimAge), the output is essentially a risk score expressed in age units.
  • If the model was trained on longitudinal aging rate (as in DunedinPACE), the output is a rate, not an age.

These give meaningfully different numbers. The "biological age" output of two different tests on the same person is not directly comparable, because the two models are predicting different things.

This is one of the most consistently obscured points in consumer marketing. A user receives "your biological age is 42" from two different tests and assumes the two numbers are equivalent. They are not.

Property 3 - How responsive and reliable is it at the individual level?

A test that works at the population level - meaning, on average across thousands of people, the test stratifies risk meaningfully - is not necessarily a test that works at the individual level for tracking.

Population-level performance is governed by predictive validity (how well the test predicts an outcome on average). Individual-level performance is governed by test-retest reliability (how reproducible the result is in the same person under the same conditions) and responsiveness (how much the result moves when the underlying biology actually changes).

A test can have strong predictive validity at the population level - meaning it really does sort high-risk from low-risk people accurately - while still being too noisy at the individual level to track an individual's change over time. This is, in fact, the most common failure mode in the biological age category.

03Where most biological age tests are limited

With those three properties in mind, here is the honest picture of where most products in this category fall short.

Limitation 1 - The underlying measurement is often noisy

Methylation array technology - the foundation of the methylation clock category - has known measurement noise at the individual level. The aging field has been refreshingly honest about this. Higgins-Chen et al. in 2022 (Nature Aging) explicitly addressed the reliability problem and developed corrections to address it.

For someone using a methylation clock to compare their result to a population average, the noise is manageable. For someone using the same test to track whether their own number is moving over months and years of intervention, the noise is a substantial problem.

Limitation 2 - The model is often trained on the wrong endpoint

If your goal is to track and improve your healthspan, what you want is a test that predicts healthspan-relevant outcomes - morbidity, frailty, long-term health decline. A test trained on chronological age (Horvath, Hannum) does not optimally serve that goal. A test trained on a biomarker composite (PhenoAge) is closer. A test trained on long-term outcomes directly (GrimAge) is closer still.

Many consumer products bury the question of what their model was trained on. This is worth asking about any biological age test you consider.

Limitation 3 - Responsiveness to lifestyle change is often weak

Most consumers who pay for a biological age test do so because they want to know whether their efforts at health optimization are working. They expect that if they improve their diet, sleep, exercise, and stress over a year, their biological age will move favorably.

The evidence for methylation clock responsiveness to lifestyle intervention is mixed. The CALERIE caloric restriction trial showed that DunedinPACE responded to a sustained two-year caloric restriction protocol (Waziry et al., Nature Aging, 2023) - a meaningful result. But it is a single study of a heroic intervention, and the magnitude of the change was modest. Most lifestyle intervention studies have shown limited movement in first-generation methylation clocks.

If a test cannot register that you have improved, it cannot be used to track that you have improved. That is a fundamental limitation for the use case most consumers have in mind.

Limitation 4 - Regulatory status is wellness, not clinical

Most biological age tests on the market are sold as wellness products. They are not regulated as clinical instruments. This does not mean they are not built on real science. It means the assay performance, quality control, and validation requirements are governed by the manufacturer rather than by regulatory authorities. For research purposes, this is fine. For longitudinal individual tracking with clinical implications, it is a meaningful structural difference.

Limitation 5 - Interpretation often happens without a clinician

A biological age result, taken alone and outside of clinical context, can be misleading at best and actively harmful at worst. A high result can prompt unwarranted anxiety. A low result can prompt complacency that defers attention to genuine clinical issues. The optimal use of any biological age output is within a clinical conversation where the result is one input among many.

Many consumer products are designed for direct-to-consumer use without clinical interpretation. This is a structural design choice with real implications for how useful the output actually is.

04What a better biological age test would look like

If you redesigned the category from scratch, with all three of the properties above in mind, here is what you would want:

  • An underlying measurement that is stable, well-characterized, and analytically robust at the individual level. Single-analyte immunoassays generally outperform array-based methylation tests on this dimension.
  • A model trained on the clinical endpoints that actually matter for healthspan - morbidity outcomes, long-term health decline - rather than chronological age.
  • Responsiveness to lifestyle change, validated in published intervention studies across the major modifiable drivers of aging (smoking, weight, sleep, exercise, diet).
  • Regulatory status as a clinical assay, with documented analytical performance specifications and quality controls under genuine regulatory oversight.
  • Integration into a clinical workflow, with interpretation provided by a qualified clinician within a doctor-patient relationship.
  • Fast turnaround, supporting the cadence of serial testing that any useful biological age application requires.

This describes suPAR.

We did not design suPAR to be a biological age test. The biomarker was developed and validated in clinical medicine over twenty years for a different set of purposes - risk stratification, emergency department triage, prognostic assessment in cardiovascular and kidney disease. What we did was recognize that the properties that make suPAR an exceptional clinical biomarker also make it an exceptional biological age test, and build a platform that makes it accessible in that role.

The point of this article is not to oversell our product. The point is to give the people taking biological age tests a clear-eyed framework for evaluating any of them. By that framework, suPAR happens to perform well. By that framework, other tests perform less well.

05What you should ask before paying for any biological age test

If you are evaluating a biological age test - ours or anyone else's - the questions that matter are:

  • What does the test actually measure, physically, in the laboratory?
  • What endpoint was the underlying model trained on?
  • What is the published test-retest reliability at the individual level?
  • What is the documented response to specific lifestyle interventions in published trials?
  • What is the regulatory status of the assay?
  • Is clinical interpretation included, and by whom?
  • What is the turnaround time, and what is the recommended cadence for serial testing?
  • What is the total cost over a meaningful longitudinal period (one year, two years, five years)?

If a vendor cannot answer those questions clearly, that is information.

06Bottom line

Biological age testing is a legitimate and useful category. It is also, in its current consumer form, often less rigorously framed than it should be. A "biological age" is not a measurement of how old your cells are; it is a statistical translation of a real measurement into an age-equivalent number. The quality of that translation depends on the quality of the underlying measurement, the relevance of the training endpoint, and the responsiveness of the system to actual change.

We built Supar Health because we believed the category could be more honest, more clinically grounded, and more useful. We believe suPAR is the best single biological age test currently available against the criteria that matter most. We have written this article to give you the framework to evaluate that claim - and any competing claim - on its merits.

Built for tracking

A clinically validated biological age test, built for tracking.