Skip to content

Duration evaluation: synthetic corpus

Synthetic data — not a measurement of anything real

Every number on this page comes from a manufactured run history, not from people cooking. The corpus is rhylthyme-cli-runner/tests/fixtures/history/thanksgiving-synth-runs.json: 20 records generated by rhylthyme_cli_runner.history.synthesize_runs with a known generating law, so the correct answer is known in advance and the evaluation can be checked against it. It says whether runs evaluate measures what it claims to measure. It says nothing about whether prediction helps a real cook. That question is answered by the catalog report, which is gated on the 30-day observation window.

What the corpus contains

20 runs of the Thanksgiving example (thanksgiving-one-oven), one per day, all completed on a wall clock at speed 1, in one environment (home-kitchen), with two declared variance factors answered on every run: turkeyKg (5, 6, 7 or 8) and oven (gas or convection).

Two steps are measurements (the rest are fixed, whose observed duration only confirms their timer):

Step Duration in the program Generating law in the corpus
turkey-roast indefinite, defaultSeconds: 9900 600 + 90 · turkeyKg seconds, × lognormal noise, CV 0.05
potatoes-boil variable, defaultSeconds: 1200 1200 seconds, × lognormal noise, CV 0.10

The two are deliberately opposite cases:

  • the roast's real duration is a function of a declared factor and is nowhere near the author's number — the author wrote 2h45m for something that takes about 20 minutes;
  • the boil's real duration scatters around the author's own number, which is therefore already the truth.

The command

rhylthyme runs evaluate rhylthyme-server/static/examples/thanksgiving_one_oven.json \
    --runs-dir <a directory holding the 20 records> --format md

Result

Runs: 20 usable of 20; trained on 16, held out 4 (latest 20%) Scored 8 of 8 held-out measurements (8m) Overall MAE: predicted 0:01:28 vs planned 1:13:31 (+98.0%) Claim (predicted beats planned on every predictable step type): FAILS for stove

Per step

Step Type Task Verdict n Planned Median actual MAE predicted MAE planned Improvement Basis
*turkey-roast indefinite oven predictable 4 2:45:00 0:19:06 0:01:21 2:25:46 +99.1% 4m
*potatoes-boil variable stove predictable 4 0:20:00 0:20:52 0:01:35 0:01:15 -25.8% 4m

By step type (task)

Group Steps n MAE predicted MAE planned Improvement Predictable steps Predictable n Predictable improvement
oven 1 4 0:01:21 2:25:46 +99.1% 1 4 +99.1%
stove 1 4 0:01:35 0:01:15 -25.8% 1 4 -25.8%

By duration kind

Group Steps n MAE predicted MAE planned Improvement Predictable steps Predictable n Predictable improvement
indefinite 1 4 0:01:21 2:25:46 +99.1% 1 4 +99.1%
variable 1 4 0:01:35 0:01:15 -25.8% 1 4 -25.8%

Improvement is 1 - MAE(predicted)/MAE(planned): positive means the forecast beat the author's number. A step marked * was called predictable by runs report on the training runs; the acceptance claim is read off those rows only. n counts held-out measurements that had BOTH a prediction and a planned value, so the two MAE columns are over the same sample; Basis says which lookup branch answered (i=identical, m=model, n=none). Fixed steps never appear: their observed duration only confirms their timer, so there is nothing to forecast.

Reading it

The roast: +99.1 %. The plan is wrong by 2h25m on average; the forecast is wrong by 81 seconds. This is the case the feature exists for — a duration that depends on something the author declared, where one fixed defaultSeconds cannot be right for every run. runs report calls the step predictable (CV 0.10) and the held-out evaluation confirms that the verdict means what it says.

The boil: −25.8 %. The forecast is worse than the plan, by 19 seconds of mean absolute error. This is not a bug, and it is the most useful thing on this page: the corpus was built with the boil's durations centred exactly on the author's 1200 s, so the author's number is the true mean and a median of 16 observations can only add estimation noise to it. Any estimator would lose this comparison. Real authored durations are not the true mean of anything, which is exactly why the honest measurement has to be made on real runs.

The acceptance claim therefore fails on this corpus, for stove. The claim — predicted beats planned in mean absolute error for at least the step types the inferentiality report marked predictable (PRD §8) — is a statement about real history, and a synthetic corpus in which one planned value is already perfect is not a fair test of it. What this page does establish:

  1. runs evaluate splits chronologically (16 train / 4 held out, the latest four run ids listed above), predicts each held-out run in its own context, and scores both estimators on the same observations;
  2. where a declared factor really drives a duration, the lookup recovers the law and the improvement is nearly total;
  3. where it does not, the lookup degrades to roughly the author's number rather than to nonsense — and the report says so instead of hiding it.

Caveat on basis

Every held-out measurement here was answered by the model branch (8m above), and in this corpus the model is the median fallback: with 16 training measurements no factor clears the Pearson threshold for the boil, and for the roast the fit is on turkeyKg. A corpus with more runs per (turkeyKg, oven) combination would be answered by the identical branch instead. Both are reported per row, so a reader can tell which lookup was being measured.

Reproducing

rhylthyme-cli-runner/tests/test_evaluate.py::TestCommittedCorpus asserts these numbers, so they cannot drift silently. The fixture's generator block records the exact synthesize_runs arguments.