Field-collected sustainability data is a claim shaped by whoever collects it. A defensible verdict needs at least one check that cannot peek at the passing score.
William Ouellette, CTO, Epoch Blue · Last updated 28 August 2026.
A field agent stands on a cocoa farm in Côte d'Ivoire with a smartphone. The app shows the number of shade trees the farmer reported, and whether today's count will qualify the farm for its premium. The agent counts, and the count comes up short. What happens next?
Researchers at Cambridge and ETH Zurich asked this same question, and the results, published in Science this month, are unsurprising. The study ran a randomized experiment inside a cocoa trader's sustainability program covering 407 farms across 85 villages, 17 cooperatives and 28 field agents, followed from 2024 to 2026. About one in four entries had been manipulated by the agents collecting it. When a farm missed its target on the first pass, agents amended the count in 33.6% of cases, compared to 8.8% on farms that passed straight away. And when the researchers withheld the farmer's originally reported count, so the agent couldn't see what number the check needed to match, false reporting fell from 24.8% to 11.2%. That’s a 55% reduction.
We think this is one of the most useful sustainability results of the year, and the main takeaway is: ground-collected sustainability data fails structurally when its collector can see the passing score, so a defensible verdict needs at least one check they can't see. We are observing this in our own cocoa work from the buyer's chair, where teams assembling plot files and audit-ready data rooms are asking what weight field evidence can carry and what to do when the ground story and the satellite verdict disagree.
The flawed design is fixable
The temptation is to read this as a fraud scandal. The researchers read it as a design result, and we agree. Federico Cammelli, the lead author, put the cause plainly: "There are conflicts of interest throughout the entire value chain, and there is a lack of incentives to ensure that the data collected on cocoa farms is independently verified." The agents have every reason to stay on good terms with the farmers they check, and the program pays premiums on the numbers they input. The app accepted retroactive corrections, and the edits were almost exclusively in the fields that determined whether they passed the audit.
The 24.8% is a modeled estimate of manipulation, not a count of proven-fraudulent records, and this was just one trader's program, so a certifier would fairly say their scheme runs integrity checks that this program lacked. Third-party field checks do run in Ivorian cocoa, and their results are not public. Both points stand; the core message is clear: remove one piece of information, the passing score, and false reporting halves.
The same hands feed a lot of ledgers
Here is why this reaches well beyond one shade-tree program. Data collected in the field by interested parties is the raw material of sustainability claims in soft commodities. Certification premiums in cocoa run on it, in the same West African belt this study worked in, where Côte d'Ivoire and Ghana together grow nearly 60% of the world's crop. Buyer commitments have entered their verification phase and are discovering the difference between process and proof: the UK Soy Manifesto's own dashboard reports that 96% of signatories have policies and 84% report publicly, numbers about paperwork rather than verified volumes. Payment-for-outcome programs, including carbon projects, pay out on field-reported measurements by construction.
And then there's the EUDR, where the stakes are less forgiving. Operators filing due diligence statements from 30 December 2026 carry legal responsibility for their geolocation and legality data "regardless of the means or intermediaries they use to collect that information," as stated in the Commission's FAQ. A cooperative's manipulated entry becomes the operator's breach, with at least 3% of standard-risk operators checked each year and fines with a ceiling starting at 4% of EU turnover. The same FAQ names remotely sensed imagery as a legitimate way to verify what declared plots actually experienced. The paper's authors come to a similar conclusion under Article 9: companies are liable for accurate and truthful geolocation data, yet the services helping them comply track data quality rather than truthfulness. The regulation, in other words, already assumes the file and the land might tell different stories.
Which verification channels can see the passing score
Lay the verification channels side by side and the pattern is hard to unsee:

This is the practical difference between company-level ESG data and production geography. An entry by someone who can see your targets could lead to a bluff you can’t back up. A satellite time series over a plot won’t consider your certification threshold, premium structure, or your filing deadline, making it the most useful file you can have.
Before anyone quotes us saying satellites solve this: they don't solve it alone. A satellite can't count seedlings under a closed cocoa canopy, and single global maps carry their own failure mode. In one Mexican pilot, 75% of 600 smallholder coffee plots were wrongly flagged as non-compliant because the check relied on a one-layer map. We've argued before that no single map should decide a plot's fate, and we hold that position here. The paper is careful in the same way: it calls the remote-sensing cross-check the EUDR embeds "a promising avenue," and in the same breath notes that shade trees only become visible from orbit years after planting and that field collection stays irreplaceable for most behaviors. The argument is independence. A verdict only stands when checks with different incentives and different blind spots agree. Rachael Garrett, the paper's senior author, pointed to the operational consequence herself: "It's time to think more seriously about supply-shed monitoring rather than a constant focus on individual farms." A supply shed, the full growing landscape a mill or cooperative draws from, is the unit where observation is strongest, and gaming is weakest.
How we'd redesign the check
Four changes based on the evidence:
- Blind collection. Whoever records your data shouldn't see thresholds, targets, or the farmer's prior reports during collection. That single change halved false reporting in the study.
- Cross-check the deciding claims against observation. For anything a satellite can see (plot boundaries, deforestation status, land-use change), put an independent record next to the field entry.
- Spend field visits where the layers disagree. That is where our ground effort for cocoa went this month: georeferenced photos and farmer questionnaires for the specific plots where the verdict was contested, and full polygons for the plots that decide the most.
- Assess the supply shed before the farm. Farm-level precision invites farm-level gaming. Landscape-level observation sets the baseline individual claims must reside within.
This is the part Epoch works on: we derive plot and supply-shed assessments from observation of the land, built for the case where suppliers won't or can't hand over data, and we score agreement across six independent detection systems because we don't fully trust any single layer, ours included.
The belief worth updating is this: field data isn't ground truth; it's a claim with an incentive attached, and a verdict becomes defensible when a check that can't see the target agrees with it. A satellite doesn't know what your certificate needs to say. Right now, that's exactly what makes it useful.
If you'd like to learn more about how we verify plot and supply-shed claims against six independent observation systems without supplier cooperation, or put the product to work in your supply chain, reach out to us here.
Or, subscribe to our newsletter to get the latest on supply chain risk management, EUDR, deforestation, water stress, and the latest trends in geospatial AI.

