History fallback and model choice
When a week has no finished walk, or the ripening model can't run, DeepLeaf Yield forecasts from your harvest history. On weeks with a walk, each week ahead keeps the method that has earned it on this cycle, as below.
The history models
harvest-persistence-v1- Each week repeats the most recent harvest logged in the three weeks before. It is offered on every week ahead, and is picked only after it beat the current pick by 20% over 8 checked weeks of this cycle, this week included.
harvest-blend-v1- The mean of that last harvest and the mean harvest of the last three weeks. It is offered for next week and the week after only, and has no fitted parameters.
harvest-pretrained-v1- The History model: a forecasting model trained in advance on many crops' weekly series, given only this cycle's harvest history. It is offered on every week next to the others, when the server has it, and replaces the current pick for a week ahead only on the terms below.
harvest-finetuned-v2- The Tuned history model: the same kind of model, further trained on public greenhouse tomato harvests. It is offered next to the others when the server has it, and is picked for a week ahead only after it beat the current pick by 20 % over 8 checked weeks of this cycle, this week included.
harvest-mean3-v1- The three-week mean: the mean of the harvests logged in the last three weeks (with fewer weeks logged, the mean of those). This is the default, offered on every week ahead, with no fitted parameters. When the History model runs, the average below takes its place as the default.
harvest-mean3-pretrained-v1- The three-week mean and History model: the plain average of the two, for each week ahead, with no fitted parameters. When the server has the History model, this is the default and every other method has to beat it by 20% over 8 checked weeks. On 305 public tomato greenhouses it was more accurate than the three-week mean for single weeks and for 4- and 8-week totals; on the 11 earlier public series, only the 4-week total was clearly better. Without the History model, nothing changes.
harvest-mean3-pretrained-sim-v1- The three-week mean and two history models: from 3 weeks ahead, the plain average of the three-week mean, the History model and a second history model trained first on simulated greenhouse seasons and then on public tomato harvests, with no parameters fitted per cycle. When the server has both models, this is the default from 3 weeks ahead and every other method has to beat it by 20% over 8 checked weeks; this week and next week keep the average above. On 305 public tomato greenhouses and on the 11 earlier public series, it was more accurate than that average from 3 weeks ahead and for 4- and 8-week totals, also on greenhouses and regions it had not seen. Without the second model, nothing changes.
From next week on, the three-week mean is the pick until another method beats it by at least 20% on mean error over 8 or more checked weeks of this cycle where both were offered. The pick then stays until something beats it the same way, so one lucky week cannot flip it back. This week keeps the earlier rule for the History model: the history value moves off the three-week mean once it was closer on two checked weeks; repeating last week and the Tuned history model need the 20% over 8 weeks here too. A short history behaves like the plain three-week mean.
Why the three-week mean is the default: on two public greenhouse tomato datasets, with repeating last week as the default the pick stayed on it for 409 of 484 scored weeks, and the three-week mean was more accurate on the same weeks: 50.9% against 48.4% for a single week, 69.4% against 66.6% for 4-week totals and 71.8% against 67.0% for 8-week totals (accuracy is 100 minus the mean error as a share of the average weekly harvest). The cost is on seasons where picking zig-zags week to week: on the synthetic season, next week's mean error is 31.3% against 26.2% for repeating last week.
Choosing per cycle
On a week with a walk, the ripening value is one more method on the same terms. This week, it is kept until the history value was closer on two checked weeks. From next week on, it has to beat the current pick by 20% over 8 checked weeks. Only weeks before the walk are compared. Each row reports:
chosen: the model that gave the value;model_kg: the ripening value, even when it wasn't chosen;model_errorandbaseline_error: the mean relative errors that decided the choice;baseline_kg: the persistence value, on every row, for comparison;selection: which rule made the pick and why (default,held, orbeat_pickwith the weeks and errors compared).
Why the stricter rule from next week on: with only two checked weeks, a method that was closer by chance on zig-zag picking weeks took over and then lost, so next week's forecast was worse over a season than repeating last week (28.3% against 26.2% mean error on the synthetic season, 28.5% against 23.9% on a 38-week test cycle). With the stricter rule, and repeating last week as the default then, next week matched repeating last week on both, and at every distance it stayed within 4 points of it on synthetic, public and test seasons.
A forecast stored with POST …/forecasts records the chosen model's name.
Why the blend, and not more
Several simple history forecasts were scored strictly out of sample on two public datasets. Only the blend was at or below persistence for next week and the week after on both datasets. The gain is small, about 1.5 points of error on the larger trial, and it is no better for the walk week, so it is not offered for that week. Ratios from temperature and radiation were not consistent across datasets, so they are not used.