Ladakh Build — where the numbers come from

1. What this is

Ladakh Build is the training dashboard I built for my own 19-week preparation for the Ladakh Half Marathon on 13 September 2026, run at 3,524 m. An AI coach reads each week's logged training, scores it, and writes the next week's plan. The site publishes all of it.

This document classifies every number the system uses. Evidence means it traces to a study and the study says what the system uses it for. Convention means a threshold the field uses widely, sitting on a relationship that is continuous underneath and a cut point nobody validated. Judgement means I set the value on coaching reasoning, and no literature fixes it.

Verifying the seven citations against the actual papers moved three of them. The Efficiency Factor, which the scoring leans on hard, has no peer-reviewed source. The HRV paper I was citing does not support the direction the system uses HRV in. The mathematical-coupling argument I had credited to Impellizzeri is Lolli's, and Lolli's paper names the exact defect this codebase found in itself.

2. The numbers, classified

Phase weight vectors. Judgement. The composite is a weighted sum of four pillars: consistency, fitness, recovery, health. In base_build and build the weights are 30/30/25/15. In peak, 20/25/35/20. In taper and race, 20/20/40/20. No study says recovery is worth 40% of readiness in a taper. I decided a taper is mostly about absorbing training already done, so recovery should dominate.

Pillar sub-weights. Judgement. Inside consistency the four terms are weighted 40/30/20/10; inside fitness, 35/25/25/15. These rank the terms the way a coach would rank them, and no study fixes the ratios. The weights are held in two places, computed in code and described to the model in prose, so the two representations have to be kept in step by hand.

The Efficiency Factor. Convention. EF is total metres divided by total heartbeats, where heartbeats are average heart rate times duration, summed across a week's qualifying runs. The metric comes from Joe Friel and lives in his books and TrainingPeaks documentation. It appears in no peer-reviewed journal, and nothing validates it against a lab measure of running economy, which is why it is classified as convention here.

Friel's condition is that EF only means anything between closely matched workouts, so the system filters twice before computing it. Outdoor runs only, because pace at a given heart rate differs systematically on a belt and belt distance is unreliable. Runs at a perceived effort of 6 or below on a 10-point scale, because efficiency rises with intensity and a quality session appearing in some weeks and not others would make a change in session mix look like a change in fitness. The threshold is 6 rather than 5 because at 5, four of nine weeks had fewer than two qualifying runs before the surface filter even applied. When no qualifying run exists the system scores the aerobic term at its midpoint instead of manufacturing a number, and every week is labelled on the site as measured, thin or neutral so a reader can see which of the three happened. A day logged with distance but no duration is excluded from numerator and denominator both, never counted as zero.

The 300 coefficient. Judgement. The aerobic term scores 70 plus 300 times the fractional change in EF against a trailing three-week baseline, clamped 0 to 100. The 300 is a presentation scale. It sits alongside 300, 200 and 150 on the other sub-scores in the same file so that a realistic week-to-week change moves the number by a readable amount. I set it to that scale, and the code comment records it as a presentation choice rather than a physiological constant.

What recovery is made of. Judgement. Recovery is never less than a quarter of the composite and rises to 40 percent in a taper, which makes it the heaviest pillar at the point in the block where the scores matter most. It is built from four terms: resting heart rate at 35 percent, sleep duration at 30 percent, sleep consistency at 15 percent, and perceived effort at 20 percent, with the HRV modifier applied to the total afterwards. The perceived-effort term scores 100 when the week's average sits inside a target band, 4 to 6 in base_build and 5 to 7 in every later phase, and drops 15 points for each point of average effort outside it. Those bands are mine. They encode the ordinary coaching view that a base week should feel easier than a build week, and no measurement fixes where the edges fall.

HRV. Convention. A week-on-week drop of ten percent or more costs five recovery points; a comparable rise adds three. I had cited Plews and Laursen for this. Their conclusion is that in elite endurance athletes the relationship runs both ways: rises and falls have both been observed alongside negative adaptation, and athletes have improved while HRV fell. They state that practical monitoring uses are not established. The system needed a direction, so I took the common one and weighted it as a secondary modifier rather than a primary input.

Sleep. Convention. Sleep scores 100 at 7.5 hours and 30 below 6.5. Milewski is a real study of 112 adolescent athletes, mean age 15, at one school, with injuries taken from a retrospective records review. It supports the association between short sleep and injury in teenagers. It does not establish a threshold for a 24-year-old runner. The cut points are mine.

The long-run progression cap. Convention, with a trial against it. The system caps week-over-week long-run growth at fifteen percent, a descendant of running's 10 percent rule. Buist randomised 532 novice runners between a standard 8-week program and a gradual 13-week program built on that rule. Injury rates came out the same. The cap stays because progressing to a schedule beats progressing arbitrarily, and because this athlete has a shin that rewards caution. Nobody should read it as protective in any demonstrated sense.

The ACWR band. Evidence, with a large asterisk. The acute-to-chronic workload ratio compares this week's distance against the average of the recent weeks, and the system treats 0.8 to 1.3 as the safe band, from Gabbett. Load here means kilometres and nothing else. Intensity, terrain and perceived effort do not enter it, so a week of easy running and a week with a time trial in it are the same load if the distances match. Gabbett is a review of team-sport studies covering rugby league, cricket and Australian football, with no distance runners in the underlying data, and the specific figure that produces the band is the one Impellizzeri's group formally asked the journal to retract. A later paper from that group found that substituting invented chronic-load values predicted injury about as well as real ones. The band is in the system because it is the field's shared reference point.

The ACWR window. Lolli. As originally written, the chronic average included the acute week in its own denominator. Lolli's 2019 paper names that as mathematical coupling: it damps the ratio toward 1.0 and manufactures correlation. The chronic window is now the three weeks preceding the acute week, acute excluded, and the same window is used in both of the two places the ratio is computed. Weeks 9 and 13 are where the fix bites hardest, because those follow real load jumps and the old ratio had been hiding them.

HR caps. Judgement, held in the library. Heart-rate ceilings are attached to aerobic families only: 160 for easy Z2 and the long run, 150 for a recovery run, 165 for a fuelled long run and the dress rehearsal, 145 for the altitude jog in Leh. The fourteen quality workouts carry no heart-rate field at all and are prescribed by pace and effort, because capping heart rate on an interval session would be capping the stimulus the session exists to produce. Nothing here is a claim about what the athlete's heart rate does on a given run. It constrains what a plan is allowed to ask for on easy days, which is where the constraint matters, since the failure mode on an easy day is running it too hard. Because the ceilings sit in the segment definitions, a plan cannot exceed one without inventing a segment the library does not contain, so the constraint holds by construction rather than by an output check. The planner may raise an aerobic ceiling by 5 bpm, but only when temperature is above 32°C and humidity is above 85 percent at the same time. Kolkata clears the humidity condition for most of July and August. The temperature condition is what withholds the allowance, because a 5:15 AM start sits at the day's minimum and this athlete's training range runs 26 to 32°C. The conditions are also read as a single live sample at the moment the plan is generated rather than for the hour the session is actually run.

The health pillar. Judgement. The pillar exists to carry one input no wearable can produce: a pain flag the athlete logs directly. For a runner whose single hard stop is a shin that must be respected at the first sign of tightness, that signal needed a pillar of its own rather than a seat inside recovery, where it had previously been worth about 5 percent of the composite and could be outvoted by the metrics around it. It is now roughly 15 percent, and the reweighting is a judgement call.

The pillar is 70 percent severity and 30 percent persistence. Severity maps the week's worst pain flag: none 100, tightness 70, a modified session 40, a stopped session 10. Those four numbers have no source beyond my own sense of how much each one should cost. Persistence is the share of logged days carrying any flag, kept continuous so that one tightness day in seven and five tightness days in seven score differently.

What "fitness" actually measures. Judgement. Two of the fitness pillar's four terms read week-over-week change rather than absolute level, so the pillar rewards progression and not condition. A disrupted week therefore penalises the following week twice, through two separate terms. Week 13 is the clearest case: it was executed almost perfectly and scored 96 on consistency, but scored 31 on fitness, because its 13.74 km long run measured against week 12's collapsed 5.84 km clamped the progression term to zero, and its volume against a depressed three-week average clamped the load term too. The arithmetic is correct and surfacing a sharp load jump is the formula doing its job. The pillar's name is what misleads, and a reader should take the fitness score as a progression score.

Timeline anchoring. Bookkeeping. In the scoring path, the countdown to race day and the count of days on iron treatment both derive from the scored week's own Monday, never from the clock at computation time. Without it, a recompute run today would stamp every historical week with today's race proximity and fabricate a trend across nine weeks of history.

3. Limits on what the system decides

Each entry below states one thing the system does not decide well, the reason, and what stops it reaching the athlete.

It reduces the wrong session. When recovery signals fire, the generator cuts the long run and keeps the easy filler around it. A coach does the reverse: shorter easy days, a lighter gym session, a softened quality session, long run intact. Week 15 showed it clearly, where the generator's draft and my revision agreed almost exactly on total weekly load and disagreed completely on how that load was distributed. No check on total volume would ever catch that.

The reason is that nobody wrote down what to protect. The prompt said to reduce load and never named the session that must survive a reduction, so the model reduced the largest reducible number in front of it. The project calls this protective overcorrection. What holds it in check is an assertion in the evaluation harness that fails any plan whose longest run falls below 90 percent of that week's long-run target, and which grades reductions rather than excusing them, because reductions are the whole subject.

Most of the rules are advice, not enforcement. Two rules are enforced in code and will block a plan outright: whether a workout is allowed in the current phase, and whether the workout exists in the library. The heart-rate ceilings hold because they are built into the segment definitions. Everything else is softer than the prompt makes it sound. The mileage-spike check runs and records but does not block. The Monday alternation has its input computed in code and its output checked nowhere. The recovery-cost budget and the contraindication gates live only as sentences in the prompt, and for the gates the triggering signal is never computed at all, so the model is expected to read it out of the raw logs. What holds this in check is that no plan is published without review, and that the harness now tests the rules the code does not.

The weekly mileage target ignores what was actually run. The targets come from a fixed 19-week schedule set at the start of the block, and the generator reads that schedule alone. Week 14 is the clearest case. Its target was 37 km, set with no reference to week 13's logged 29.63 km, and the plan that came back measured 61.1 km against that prior week, an increase of 106 percent where the rule allows 15. The check records the breach and does not block on it, so what holds it in check is the review every draft passes through before publication. That review is a designed step in the pipeline, not a rescue.

A headline distance on the site may not match the plan behind it. The large distance shown at the top of a week is not read from a stored field. It comes from a routine that scans the text of each day's session, picks out every number that looks like a distance, and keeps the largest. The routine runs at two different points, once before a day's full workout text has been filled in and once after, so the number the mileage check measured is not always the number the page displays. Week 14 is the visible case: the plan targeted 37 km, the check measured 61.1 km, and the page displays 61 km. It stays uncorrected as the worked example, and the stored plan is the authoritative record.

4. Method notes

The model never changed. Both of the system's server-side functions, the one that writes plans and the one that writes reflections, are pinned to claude-sonnet-5, and that model served every generation and every reflection in the corpus. The two workout libraries have been frozen since ID-based selection began. Any difference between two plans is the planner's decision against stable definitions. This is discipline rather than limitation: it is what makes week-to-week comparison mean anything.

The harness grades prescriptions, not runs. The evaluation harness checks the coach's stored plans against the rules the system claims to follow, never the athlete's logged training. Eleven assertions cover it; each is defined below. The live eval panel on the site reads results from a file at runtime and updates on every re-run. This document carries the reasoning and no counts.

The eleven rules, defined

Recovery-cost budget. Every leg exercise carries a recovery cost. The ones picked for leg day have to add up to less than the ceiling for that phase.

Quality-volume floor. Tuesday's hard session must contain at least 2 km of genuinely fast running. A token interval set on the one hard day of the week is not a hard week.

Contraindication gates. When last week's data raised a flag, the plan cannot prescribe a workout the library marks off-limits under that flag.

Phase eligibility. Every workout has to be one the library permits during the phase the build is currently in.

Heart-rate ceiling. No run may carry a heart-rate target above the limit for the phase — 165 through the build, 145 on race day.

Weekly increase limit. When planned distance goes up, it cannot go up more than 15% over what was actually run the week before. Flat and reduced weeks have no rule and aren't graded.

Monday alternation. Monday alternates between an upper-body gym session and a cycling session. The same one twice running is a failure.

Week shape. The seven days have to match the fixed weekly structure, and each gym day has to carry the muscle groups that day is for.

Public read. A logged-out visitor has to actually be able to read that week's published reflection. Checked live against the database, not assumed.

Weekly distance floor. Planned running has to reach at least 85% of the week's target. If a genuine reduce signal was active the week isn't graded — a justified cutback is not a failure.

Long-run floor. The longest run has to reach at least 90% of its target. Unlike the weekly total this is graded even when reduce signals were active, deliberately: the question is whether the long run survived a week when load was cut.

Writing the harness found a defect in its own specification. A1 was defined as a rolling three-to-five-day recovery budget graded against per-phase caps. Those caps were never implemented as a rolling window; they govern a single adaptive Leg Day session, and under the rolling reading no conforming plan could pass, since one Leg Day alone can exceed the cap. The assertion was rescoped to the session the caps actually govern rather than recalibrating thresholds that exist nowhere in the system.

The recompute delta. When the scoring defects were fixed, all nine stored weeks were rescored under the corrected formulas with the originals preserved beside them. This is stored data; re-running the harness does not change it.

Week Consistency Fitness Recovery Health Overall
5 75 → 84 84 → 93 43 → 46 69 → 70 65 → 75
6 100 → 100 68 → 60 52 → 52 77 → 100 75 → 76
7 93 → 93 58 → 58 61 → 57 77 → 100 74 → 74
8 54 → 54 29 → 27 62 → 63 75 → 100 58 → 55
9 100 → 100 44 → 39 64 → 62 71 → 100 70 → 72
10 87 → 87 43 → 47 75 → 71 71 → 100 68 → 73
11 92 → 92 74 → 70 52 → 52 68 → 100 73 → 76
12 60 → 60 35 → 33 69 → 67 68 → 100 56 → 60
13 96 → 96 46 → 31 47 → 46 68 → 100 64 → 65

The health column's jumps to 100 are correct. The corrected pillar drops two terms that had been depressing it, and 100 now means no pain flag logged that week. Week 5, the only week with pain flags, stays down. Week 13 is the worked example on fitness: 13.0 of its 15-point drop comes from the ACWR self-exclusion fix, where the self-inclusive ratio of 1.62 was already past the top of the band and scoring 52 of 100, and the corrected 2.03 scores zero, and 1.36 points come from the aerobic term, where measured EF fell 1.81% against baseline. The correction makes the score worse. The old formula damped load spikes toward safe; the new one reports them.

The reasoning fields are prospective. Every plan carries a written explanation of why each session was chosen, produced before the week and then frozen, never edited afterwards. Week 13's plan says it is capping at 12 km rather than 19 km, and the athlete logged 13.74. It stays as written, and the correction belongs to the following reflection. A reader's default assumption is that these were written with hindsight, and the mistakes left standing in them are the evidence that they were not.

A few late-block weeks are the exception and should be read differently. Generation fell behind the calendar near the end of the block, and those plans were written after their weeks had run. Their reasoning is reconstruction rather than prescription. The affected weeks are named on the site's evaluation panel, which reads them at runtime rather than from anything typed here.

5. Close

Nineteen weeks of plans and reflections were generated, scored and published, on one pinned model and against two frozen workout libraries. That is what makes the record readable. Any week can be compared against any other, and where two weeks differ, the difference belongs to the planner's decision rather than to a version change underneath it.

The document you have just read is the other half of that. Every number the system uses is named, sourced where a source exists, and marked as mine where it does not. A reader who disagrees with the 300 coefficient or the sleep cut points or the weighting of the health pillar can find the number, see who chose it, and argue with the choice. That was the point of writing it down.

The clearest thing the build taught me is where the next version starts. Most of what the system decided badly, it decided badly because a rule lived in my head and not in the specification: which session a reduction has to protect, which floor a plan may not cross, which signals describe the week rather than the workout. Those are written down now, and the harness tests them. The next one begins from the specification instead of arriving at it nineteen weeks later.