Backtest and error ranges
The backtest replays the forecast for each past walk week, in order, using only what was known then, and compares it with the harvest that followed.
/app/forecast. The backtest table sits below the chart.Reading the backtest
GET …/backtest and the app's table report, for each horizon:
| Column | Meaning |
|---|---|
| Weeks checked | Past forecasts that have a harvest to compare with |
| Mean error | Mean absolute error in kilograms (mean_absolute_error_kg) |
| Mean % error | Mean absolute percent error of the issued forecast (mean_absolute_percent_error) |
| Repeating last week | The same error for plain persistence (baseline_mean_absolute_percent_error) |
| Inside range | How many harvests fell between low_kg and high_kg, such as "5 of 8" |
If the forecast's error is not below "Repeating last week" on your cycle, the ripening model isn't adding anything there yet. The per-cycle choice will already lean on the history for those horizons.
For the same figures pooled over a greenhouse, with a trust level per horizon, see How far to trust the forecast.
Reading the error range
- The range appears once a horizon has two checked weeks. Before that, the forecast is a single number without a range.
- It is the forecast plus and minus this cycle's mean relative error at that horizon. It is not a statistical confidence interval.
- With few checked weeks, the range can be too narrow. Check "Inside range" before relying on it.
- The weeks further ahead usually have wider ranges, because their forecasts were further off.
Stored forecasts
POST …/forecasts issues the current forecast and stores it. GET …/forecasts lists stored forecasts next to the harvest logged since, so you can check what was actually said at the time, not only a replay.
The backtest and accuracy replay each past week with your walk corrections as they are today, as if they had been known then. Correcting an old walk moves the error of the weeks it fed. Stored forecasts don't change: they keep the number they were issued with, before any later correction.
Results on public data
No public dataset has filmed walks. The six-compartment trial has harvests but no fruit counts, so its numbers measure the harvest-history forecast, the baseline the ripening model has to beat. The smaller trial also has weekly fruit counts, without ripeness, so the walk models were run there with every fruit treated as green.
| Dataset | Model | This week | Next week | Week after |
|---|---|---|---|---|
| Six-compartment cherry tomato trial 86 / 80 / 74 checked weeks | Repeat last harvest | 48.9% | 32.0% | 46.6% |
| As issued (blend selectable) | 47.9% | 32.2% | 43.9% | |
| Smaller hand-recorded trial 135 plants in 3 series, 33 / 30 / 27 checked weeks | Repeat last harvest | 65.4% | 100.3% | 102.6% |
| As issued, earlier walk model | 78.5% | 91.2% | 70.7% | |
| As issued, current walk model | 64.5% | 69.7% | 61.3% |
On the six-compartment trial, compartments ranged from 37.6% to 67.8% error for the week itself. It is picked every 4 to 5 days, so a week holds one pick or two and the weekly total zig-zags. The smaller trial had small plots and irregular picking, so its harvest swung hard from week to week. There the current walk model lowered the error for next week and the week after in 3 of 3 series. Most of that gain comes from learning each cycle's pace: the published curve alone, at a pace of 1, did worse than the earlier model there.
Read these with their limits. The smaller trial is one greenhouse and one season. Its counts don't split ripeness, so the pace was learned from fruit picked the week after a count, not from turning fruit on the next walk as it is on your rows. Its weekly error stays high either way. On the six-compartment trial both walk models give the same numbers, because it has no fruit counts.