DeepLeaf Yield DeepLeaf Yield

Backtest and error ranges

The backtest replays the forecast for each past walk week, in order, using only what was known then, and compares it with the harvest that followed.

The Forecast page in the app for a demo cycle: the forecast table with week, fruit ripening, temperature, forecast, based on, range, and harvest logged, then a chart of each week's forecast against the harvest. The backtest table, How past forecasts compared with the harvest, comes further down the page.
The Forecast page at /app/forecast. The backtest table sits below the chart.

Reading the backtest

GET …/backtest and the app's table report, for each horizon:

ColumnMeaning
Weeks checkedPast forecasts that have a harvest to compare with
Mean errorMean absolute error in kilograms (mean_absolute_error_kg)
Mean % errorMean absolute percent error of the issued forecast (mean_absolute_percent_error)
Repeating last weekThe same error for plain persistence (baseline_mean_absolute_percent_error)
Inside rangeHow many harvests fell between low_kg and high_kg, such as "5 of 8"

If the forecast's error is not below "Repeating last week" on your cycle, the ripening model isn't adding anything there yet. The per-cycle choice will already lean on the history for those horizons.

For the same figures pooled over a greenhouse, with a trust level per horizon, see How far to trust the forecast.

Reading the error range

  • The range appears once a horizon has two checked weeks. Before that, the forecast is a single number without a range.
  • It is the forecast plus and minus this cycle's mean relative error at that horizon. It is not a statistical confidence interval.
  • With few checked weeks, the range can be too narrow. Check "Inside range" before relying on it.
  • The weeks further ahead usually have wider ranges, because their forecasts were further off.

Stored forecasts

POST …/forecasts issues the current forecast and stores it. GET …/forecasts lists stored forecasts next to the harvest logged since, so you can check what was actually said at the time, not only a replay.

The backtest and accuracy replay each past week with your walk corrections as they are today, as if they had been known then. Correcting an old walk moves the error of the weeks it fed. Stored forecasts don't change: they keep the number they were issued with, before any later correction.

Results on public data

No public dataset has filmed walks. The six-compartment trial has harvests but no fruit counts, so its numbers measure the harvest-history forecast, the baseline the ripening model has to beat. The smaller trial also has weekly fruit counts, without ripeness, so the walk models were run there with every fruit treated as green.

Mean absolute percent error, pooled over series
DatasetModelThis weekNext weekWeek after
Six-compartment cherry tomato trial
86 / 80 / 74 checked weeks
Repeat last harvest48.9%32.0%46.6%
As issued (blend selectable)47.9%32.2%43.9%
Smaller hand-recorded trial
135 plants in 3 series, 33 / 30 / 27 checked weeks
Repeat last harvest65.4%100.3%102.6%
As issued, earlier walk model78.5%91.2%70.7%
As issued, current walk model64.5%69.7%61.3%

On the six-compartment trial, compartments ranged from 37.6% to 67.8% error for the week itself. It is picked every 4 to 5 days, so a week holds one pick or two and the weekly total zig-zags. The smaller trial had small plots and irregular picking, so its harvest swung hard from week to week. There the current walk model lowered the error for next week and the week after in 3 of 3 series. Most of that gain comes from learning each cycle's pace: the published curve alone, at a pace of 1, did worse than the earlier model there.

Read these with their limits. The smaller trial is one greenhouse and one season. Its counts don't split ripeness, so the pace was learned from fruit picked the week after a count, not from turning fruit on the next walk as it is on your rows. Its weekly error stays high either way. On the six-compartment trial both walk models give the same numbers, because it has no fruit counts.