Duration evaluation: synthetic corpus¶
Synthetic data — not a measurement of anything real
Every number on this page comes from a manufactured run history, not from people cooking. The corpus is rhylthyme-cli-runner/tests/fixtures/history/thanksgiving-synth-runs.json: 20 records generated by rhylthyme_cli_runner.history.synthesize_runs with a known generating law, so the correct answer is known in advance and the evaluation can be checked against it. It says whether runs evaluate measures what it claims to measure. It says nothing about whether prediction helps a real cook. That question is answered by the catalog report, which is gated on the 30-day observation window.
What the corpus contains¶
20 runs of the Thanksgiving example (thanksgiving-one-oven), one per day, all completed on a wall clock at speed 1, in one environment (home-kitchen), with two declared variance factors answered on every run: turkeyKg (5, 6, 7 or 8) and oven (gas or convection).
Two steps are measurements (the rest are fixed, whose observed duration only confirms their timer):
| Step | Duration in the program | Generating law in the corpus |
|---|---|---|
turkey-roast | indefinite, defaultSeconds: 9900 | 600 + 90 · turkeyKg seconds, × lognormal noise, CV 0.05 |
potatoes-boil | variable, defaultSeconds: 1200 | 1200 seconds, × lognormal noise, CV 0.10 |
The two are deliberately opposite cases:
- the roast's real duration is a function of a declared factor and is nowhere near the author's number — the author wrote 2h45m for something that takes about 20 minutes;
- the boil's real duration scatters around the author's own number, which is therefore already the truth.
The command¶
rhylthyme runs evaluate rhylthyme-server/static/examples/thanksgiving_one_oven.json \
--runs-dir <a directory holding the 20 records> --format md
Result¶
Runs: 20 usable of 20; trained on 16, held out 4 (latest 20%) Scored 8 of 8 held-out measurements (8m) Overall MAE: predicted 0:01:28 vs planned 1:13:31 (+98.0%) Claim (predicted beats planned on every predictable step type): FAILS for stove
Per step¶
| Step | Type | Task | Verdict | n | Planned | Median actual | MAE predicted | MAE planned | Improvement | Basis |
|---|---|---|---|---|---|---|---|---|---|---|
| *turkey-roast | indefinite | oven | predictable | 4 | 2:45:00 | 0:19:06 | 0:01:21 | 2:25:46 | +99.1% | 4m |
| *potatoes-boil | variable | stove | predictable | 4 | 0:20:00 | 0:20:52 | 0:01:35 | 0:01:15 | -25.8% | 4m |
By step type (task)¶
| Group | Steps | n | MAE predicted | MAE planned | Improvement | Predictable steps | Predictable n | Predictable improvement |
|---|---|---|---|---|---|---|---|---|
| oven | 1 | 4 | 0:01:21 | 2:25:46 | +99.1% | 1 | 4 | +99.1% |
| stove | 1 | 4 | 0:01:35 | 0:01:15 | -25.8% | 1 | 4 | -25.8% |
By duration kind¶
| Group | Steps | n | MAE predicted | MAE planned | Improvement | Predictable steps | Predictable n | Predictable improvement |
|---|---|---|---|---|---|---|---|---|
| indefinite | 1 | 4 | 0:01:21 | 2:25:46 | +99.1% | 1 | 4 | +99.1% |
| variable | 1 | 4 | 0:01:35 | 0:01:15 | -25.8% | 1 | 4 | -25.8% |
Improvement is 1 - MAE(predicted)/MAE(planned): positive means the forecast beat the author's number. A step marked * was called predictable by runs report on the training runs; the acceptance claim is read off those rows only. n counts held-out measurements that had BOTH a prediction and a planned value, so the two MAE columns are over the same sample; Basis says which lookup branch answered (i=identical, m=model, n=none). Fixed steps never appear: their observed duration only confirms their timer, so there is nothing to forecast.
Reading it¶
The roast: +99.1 %. The plan is wrong by 2h25m on average; the forecast is wrong by 81 seconds. This is the case the feature exists for — a duration that depends on something the author declared, where one fixed defaultSeconds cannot be right for every run. runs report calls the step predictable (CV 0.10) and the held-out evaluation confirms that the verdict means what it says.
The boil: −25.8 %. The forecast is worse than the plan, by 19 seconds of mean absolute error. This is not a bug, and it is the most useful thing on this page: the corpus was built with the boil's durations centred exactly on the author's 1200 s, so the author's number is the true mean and a median of 16 observations can only add estimation noise to it. Any estimator would lose this comparison. Real authored durations are not the true mean of anything, which is exactly why the honest measurement has to be made on real runs.
The acceptance claim therefore fails on this corpus, for stove. The claim — predicted beats planned in mean absolute error for at least the step types the inferentiality report marked predictable (PRD §8) — is a statement about real history, and a synthetic corpus in which one planned value is already perfect is not a fair test of it. What this page does establish:
runs evaluatesplits chronologically (16 train / 4 held out, the latest four run ids listed above), predicts each held-out run in its own context, and scores both estimators on the same observations;- where a declared factor really drives a duration, the lookup recovers the law and the improvement is nearly total;
- where it does not, the lookup degrades to roughly the author's number rather than to nonsense — and the report says so instead of hiding it.
Caveat on basis¶
Every held-out measurement here was answered by the model branch (8m above), and in this corpus the model is the median fallback: with 16 training measurements no factor clears the Pearson threshold for the boil, and for the roast the fit is on turkeyKg. A corpus with more runs per (turkeyKg, oven) combination would be answered by the identical branch instead. Both are reported per row, so a reader can tell which lookup was being measured.
Reproducing¶
rhylthyme-cli-runner/tests/test_evaluate.py::TestCommittedCorpus asserts these numbers, so they cannot drift silently. The fixture's generator block records the exact synthesize_runs arguments.