DeepLeaf Yield DeepLeaf Yield

Accuracy and limits

What was measured, on which kind of public data, and what has not been validated yet. These are other growers' greenhouses, cameras, and labelling rules, not a validation on your rows.

The capture check alert, which warns when a walk is filmed too far, too close, or blurry, the conditions under which the measured accuracy no longer holds.
The capture check flags walks filmed outside the conditions these numbers were measured in.

Fruit detection and ripeness

All numbers are on public greenhouse tomato datasets with every fruit labelled by hand. A detection matches a labelled fruit when their boxes overlap at IoU 0.3 or more, one to one.

DataResult
Aisle images
449 frames, 1280×720, from a robot driving between tomato rows
Fruit large enough for ripeness (≥ 29 px): 79% of labelled fruit found, 64% of detections matched a label. Per image the count was off by about 4 fruit, 44% of the labelled count. In total, 22% more than the labels.
All sizes: 89% found, but four times as many boxes as labels, almost all small fruit in the rows behind that the labellers left out. Only fruit large enough for ripeness reach the forecast.
Ripeness: 96.6% agreement (4,686 of 4,850). Labelled turning fruit were most often confused with red (71 of 427).
Commercial greenhouse images
95 held-out images, phone and drone
Close-range phone frames (60, 4320×7680): 68% found, but about 1.5 times the labelled count (53% high). Most extra boxes are small pieces of one large labelled fruit. On frames the capture check calls too close, boxes far smaller than the frame's large fruit are dropped. The same happens on frames under that limit when many small boxes sit on larger fruit. Together these took the count from 93% high to 53% high for 3 points of fruit found.
Drone frames (35, 1280×720): 72% found, count 35% high. Among fruit large enough for ripeness, 29% low.
Ripeness: 92.9% agreement (1,026 of 1,104). Labelled red fruit were rare, 30 in total, so red agreement rests on few fruit.

Harvest forecast

On a public six-compartment cherry tomato trial (about 15 harvest weeks each, picked every 4 to 5 days), the harvest-history forecast was off by 47.9% this week, 32.2% next week, and 43.9% the week after, as issued. On a smaller public trial with irregular picking, repeating last week was off by 65% to 103%. There the current walk model, run on fruit counts without ripeness, was off by 69.7% next week and 61.3% the week after, against 91.2% and 70.7% for the earlier one. See the full table in Backtest and error ranges.

What is not validated yet

  • The ripening forecast on a real season with walks. No public dataset has weekly walks and harvests of the same row. The current model was checked on counts without ripeness from one greenhouse and one season (the smaller trial), and on one ripening-rate comparison, where it still ran slow on a dwarf cultivar that ripened in one wave.
  • Plant counts. Plants are proposals from stem geometry, checked on example photos and one test clip, not against a labelled plant count.
  • Sizes and mass. The sphere model and outlier rule are reasoned from example photos, not checked against weighed fruit.
  • The cover balance uses default coefficients per house type, not ones fitted to your greenhouse. Logging indoor temperature corrects it.
  • Close-range counting. Close phone frames still count about 53% more fruit than the labels, and frames just under the too-close limit about twice as many. Overcount when the largest fruit (90th percentile) are above about 60 px is a known limit (see Roadmap).

How to check it on your rows

  1. Film the same row every week the same way, following the filming guide.
  2. Log every harvest for that row.
  3. After a few weeks, read the backtest. It shows the error on your own cycle, next to simply repeating last week.