Training Science
This page is how Allure computes what it shows you: the models we use, the published research they stand on, and the standard every claim has to meet before you see it.
Every threshold, zone and grade in Allure is computed from what you actually did, measured against a model of you. We do not grade you against tables of other people. If a number cannot be traced back to your own recorded efforts, we do not show it.
Where we do lean on external research, it describes physics and human physiology in general: the energy cost of running uphill, the way heat taxes pace, how altitude lowers the aerobic ceiling. Those are properties of the sport, not norms of a population, and each one is cited below.
Every insight carries the evidence it was computed from, and every detector has to stay quiet for an athlete who almost triggers it. When your data cannot support a claim at the standard we hold, Allure says less, or nothing. A confident wrong answer is worse than none.
What one session costs you, in one currency across sports.
The slow build, the recent strain, and the difference.
Your engine, mapped from your own best efforts.
The same framework, on pace instead of watts.
Bands anchored to your thresholds, not a table’s.
Weighted power, variability, decoupling, intervals.
What your engine does deep into a session, not off the line.
Climbing rate, always in the context of its gradient.
Hilly runs and flat runs, made comparable.
Projected from your bests, within honest limits.
Ground contact and oscillation, reported without judgement.
Naming the tax the environment charged you.
HRV, resting heart rate and sleep against your own normal.
Why Allure sometimes says nothing at all.
The engine
Effort is Allure’s single currency for what a session cost you. Rides, runs and swims all pay into it, which is what lets one fitness curve describe a triathlete’s whole week.
When a session carries power, Effort is computed from your weighted intensity relative to your own sustainable threshold, scaled by duration. Riding at your limit for an hour is a fundamentally different event from riding easy for three, and the computation reflects that separation rather than adding up minutes.
When there is no power, Effort comes from heart rate, in the lineage of E.W. Banister’s training-impulse work: time spent, weighted by how far your heart rate sat above easy and below your threshold heart rate. And when a run carries power from a pedal-style pod or a watch, we still score it from heart rate, deliberately. Running watts and cycling watts are different quantities, and grading one against the other’s threshold would quietly inflate or deflate every number downstream.
Built on: Banister’s impulse–response and training-impulse (TRIMP) framework (Banister & Calvert, 1975; Banister, 1991).
Your Effort history feeds two rolling quantities with deliberately different memories. Fitness is the slow one: it accumulates over weeks and decays over weeks, which is why a big day today does not make you fit tomorrow. Fatigue is the fast one: it spikes with recent strain and clears in days. Form is the difference between them, and it is the number that behaves the way tapering feels: it dips while you build, and surfaces when you absorb.
This is the impulse–response model of training adaptation, proposed by Banister and colleagues in 1975 and validated across five decades of endurance literature. The memory lengths default to values the literature supports, and both are yours to tune in Settings, because a fifty-year-old ultrarunner and a twenty-year-old crit racer do not absorb work on the same clock.
Built on: Banister & Calvert’s systems model of training (1975) and the body of impulse–response modelling that followed.
Your profile
Every ride you record contributes to a curve of your best average power at every duration, from a five-second sprint to an hour at the limit. Allure fits that curve with the Critical Power model: a formal description of the boundary between work you can sustain and work you cannot, plus a fixed reserve, called W′, that you spend above that boundary and recover below it.
The model was introduced by Monod and Scherrer in 1965 and refined into the three-parameter form we use by Morton in 1996. It is one of the most validated constructs in exercise physiology, and everything in your profile flows from fitting it to your own bests: your estimated FTP, your rider phenotype, and your named limiter, which is simply the region where your own curve falls furthest below what your own model says you should produce. No percentile tables, no comparison to riders your age. The gap analysis is you against you.
The fit also reports its own confidence. A profile built from three easy weeks says so, and the numbers it feeds are held more loosely until you have given the model something hard to work with.
Built on: Monod & Scherrer (1965); Morton’s three-parameter Critical Power model (1996); Jones et al.’s reviews of the critical power framework (2010, 2017).
Running gets the same framework on pace. Your best times over the distances you have actually raced and trained form a speed–duration curve, and the fit produces your Critical Speed: the pace boundary between sustainable and unsustainable, with D′ as the distance-above-threshold reserve. Jones and Vanhatalo’s 2017 review describes this as one of the most robust predictors of endurance performance in the literature, and it is the anchor for your running threshold pace and pace zones.
Your threshold pace is set from this fit only when the fit is trustworthy, and never overwrites a value you have set yourself. Swimming works the same way over swim distances.
Built on: Jones & Vanhatalo, “The ‘Critical Power’ Framework” (Sports Medicine, 2017).
Zones in Allure are bands of your own thresholds: power zones from your FTP, heart-rate zones from your lactate threshold heart rate, and running pace zones from your Critical Speed. The band boundaries express definitional physiology, the difference between recovery, aerobic work, tempo, threshold and work above the ceiling. The anchors those bands hang from are measured from you, and when your threshold moves, every zone and every stored session’s classification moves with it.
This matters more than it sounds. A zone table copied from a book is only correct for the athlete the book imagined. Anchoring to your own measured boundary is what makes “Zone 2” mean the same thing for you in March that it meant in November, even though you are a different athlete by then.
Inside a session
Average power flatters a surgy ride. Allure’s weighted power emphasises surges the way your physiology experiences them: hard minutes cost disproportionately more than the average suggests, an insight formalised in the normalized-power tradition of power analysis. Intensity is that weighted output relative to your threshold, and Variability is how far the ride’s surges pulled it away from a steady effort, which is how a criterium and a steady climb of identical average power tell two different stories.
Over a long effort, a strong aerobic system holds its output at a steady internal cost. Decoupling measures the drift between output and heart rate from the first half of a session to the second. On rides with power we compare power to heart rate, because on a bicycle speed belongs to the course, and a ride that descends its second half would otherwise read as superhuman endurance. On runs we compare pace to heart rate, where pace is the athlete’s own output. The underlying phenomenon is cardiovascular drift, described in the physiology literature for decades and popularised as a field test by coach Joe Friel.
Finding hard efforts is easy. Refusing to call terrain a workout is the hard part, and it is where most detectors fail: a hilly ride has repeated efforts separated by descents, and a naive detector reports it as intervals with a straight face. Allure’s detector demands what actual structure has and terrain does not: regularity on both sides. The work repeats, and the recoveries repeat too, and the work has to be genuinely hard for you, measured against your own threshold. A session that does not meet the bar produces no interval report at all, which is the correct answer for a session that was not intervals.
Built on: the normalized-power tradition of cycling power analysis; Coyle & González-Alonso on cardiovascular drift (2001); decoupling as popularised by Joe Friel.
Fresh power flatters everyone. The number that separates road racers is what the engine produces after hours of work, and the sports-science literature has recently formalised this as durability: the decline in your best short efforts once real work has accumulated. Allure follows the fixed-work method from that literature: your best one-minute and five-minute outputs while fresh, against the same bests after a substantial, fixed quantity of work, with the fade between them as the reading.
Runs get the same idea in running’s own units. Since a run accumulates metabolic work rather than kilojoules on a crank, the split point is expressed in grade-adjusted distance, and the comparison uses grade-adjusted pace, so a course that climbs late does not masquerade as the athlete fading.
Built on: Maunder et al., “The Importance of Durability in the Physiological Profiling of Endurance Athletes” (Sports Medicine, 2021); Spragg and colleagues’ fatigued-state profiling work in professional cyclists (2022–2023).
Allure detects the climbs inside your rides from the altitude trace itself, and computes VAM, vertical metres per hour, per detected climb. VAM is only meaningful next to the gradient it was ridden on, because steep roads produce high VAM at modest power while shallow ones cap it regardless of fitness. So Allure never shows a VAM without its gradient, and never derives one from a whole ride’s totals, where descents and flats would turn it into noise. The dashboard’s climbing chart plots every detected climb against its gradient so improvement is visible as your recent climbs sitting above your older ones on equal terms.
Running
Running uphill costs more energy per metre than running on the flat, and downhill less, in a relationship measured directly in the laboratory by Minetti and colleagues in 2002, spanning gradients far steeper than most trails. Allure uses that published cost curve to express your pace as its flat-ground equivalent, sample by sample. It is what makes a mountain run and a track session comparable, and it is the backbone of the running durability measure above.
Built on: Minetti et al., “Energy cost of walking and running at extreme uphill and downhill slopes” (Journal of Applied Physiology, 2002).
Your projected race times come from Riegel’s endurance model, published in 1981 and still the standard for relating performances across distance. Allure anchors it to your own best reliable effort, and only projects within a bounded window around that anchor, because the model degrades far from the effort it was fitted to. Where you hold an actual best at a distance, the real performance is shown instead of a projection. A prediction is never allowed to overwrite a fact.
Built on: Riegel, “Athletic Records and Human Endurance” (American Scientist, 1981).
Ground contact time, vertical oscillation, vertical ratio and left–right balance are measured by your watch or pod, and Allure reports them exactly as recorded, alongside cadence and stride length derived from your own speed and step rate. We deliberately do not grade these against population norms. The running-economy literature connects these quantities to efficiency, but the honest use of them is as trends against your own history: your contact time on fresh legs against your contact time deep into a long run, and how both move over a season.
Context
Some numbers need an asterisk the athlete should not have to supply themselves. When a run happens in real heat, Allure notes the expected pace cost, drawn from marathon field studies of weather and performance, so a slow-looking session in 30°C reads as what it was. When a session averages genuine altitude, Allure notes roughly how far the aerobic ceiling drops up there, drawn from the altitude physiology literature. In both cases we name the tax without touching your data: the recorded pace stays the recorded pace, and the context sits beside it.
Built on: Ely et al., “Impact of Weather on Marathon-Running Performance” (Medicine & Science in Sports & Exercise, 2007); Wehrlin & Hallén on altitude and maximal oxygen uptake (European Journal of Applied Physiology, 2006).
Morning heart-rate variability, resting heart rate and sleep are only meaningful against your own normal. Allure builds that normal as a rolling personal baseline with a robust measure of your natural spread, so a single odd morning cannot distort what “usual” means for you. Each day’s reading is then expressed as a deviation from your baseline, which is how the same HRV number can be a green flag for one athlete and a warning for another. The approach follows the sports-science consensus on HRV monitoring: trends against an individual’s rolling baseline, never raw values against a chart.
Those deviations are then tied to your training, per session type, so the app can show what a hard day actually does to your recovery markers and how long they take to return. That relationship is computed from your history alone.
Built on: Plews et al. on heart-rate variability monitoring in elite endurance athletes (Sports Medicine, 2013).
Everything above produces numbers. The insight engine is what turns them into sentences, and it operates under three rules. First, every claim carries its evidence: the numbers it was computed from ride along with it, inspectable. Second, every detector is tested in both directions before it ships: it must fire for an athlete who has genuinely earned the observation, and it must stay silent for an athlete who almost has. The negative cases are what stop a detector from becoming a horoscope. Third, confidence is graded. A conclusion from a rich power history is stated plainly; a conclusion from thin data is hedged or withheld.
This is also why Allure will sometimes tell you less than other platforms would. A session with no clean structure gets no invented intervals. A profile with three data points gets no phenotype. We think the numbers you act on should be the ones that would survive you checking them.
Free while in private beta
Connect Strava or Garmin and every model on this page starts working on your data.
Join the beta