Good on-farm research trial design comes down to two conditions. Treatments must be replicated and randomized within the field, and the measured difference must be larger than the trial's own least significant difference (LSD). A single side-by-side strip satisfies neither. Most on-farm differences reported at 3–5% do not exceed the LSD of the trial that produced them, which makes the honest conclusion "no difference detected" rather than "the product worked."
Here is the part that should change how you plan a season. Adding more strips to one field will not rescue a small effect. A 0.25 t/ha (4 bu/ac) difference in maize sits below what any realistic number of strips in one field can resolve. Your field is one replicate. The network of fields is the experiment.
The number most on-farm trials never report
The least significant difference is the smallest gap between two treatment means that a trial can separate from background variation. Report treatment means without it and you have a demonstration, not a trial.
How often does this matter? Practical Farmers of Iowa reported in March 2026 that across 16 on-farm trials of biofertilizer and biostimulant products run by seven cooperators since 2019, 81% showed no benefit. The Manitoba Pulse & Soybean Growers On-Farm Network reached the same place from a different direction in December 2025: across 38 site-years in soybean, pea, and dry bean, biostimulant use produced no significant yield increase over untreated strips.
The pattern is not limited to products. In the Nebraska On-Farm Research Network's compilation of maize seeding rate work from 2010 to 2014, only six of 24 rainfed sites showed a significant yield increase from planting above 28,000 seeds per acre. Eighteen sites produced numbers that looked like responses and were not.
Why split-field comparisons produce false results
An unreplicated split-field comparison confounds the treatment with everything else that differs between the two halves of the field. Soil texture, drainage, compaction from last year's traffic, and weather damage all vary spatially, and none of them are randomized out when you treat one side and leave the other.
South Dakota State University Extension published a soybean row spacing trial that shows the mechanism clearly. Across six replications, 7.5-inch rows at 120,000 seeds per acre yielded 52 bu/ac against 44 bu/ac for 30-inch rows at 90,000 — an 8 bu/ac difference against an LSD of 2 bu/ac, so a real effect. A hailstorm then damaged the half of the field holding replications four, five, and six. Had the trial been a simple split, comparing the narrow rows on the hailed side against the wide rows on the undamaged side, the difference would have read 4 bu/ac instead of 8.
That example understated a real effect. The error runs the other way just as easily, and you cannot tell which happened without replication. Alesso and colleagues showed in a 2021 Agronomy Journal simulation that when spatial autocorrelation is left unmodeled, it both degrades treatment effect estimates and inflates Type I error — the rate at which a trial declares a difference that does not exist. Their designs with smaller experimental units and more replications performed best.
How many replications does a strip trial need?
Four to six randomized strip pairs is the practical minimum, and that minimum will only resolve differences of roughly 6–10%, not the 3–5% most product claims describe. The Nebraska On-Farm Research Network sets three replications as its floor and specifies more for noisier questions — its sulfur protocol calls for five. SARE's on-farm research guide recommends four to six blocks for most projects.
Here is the full calculation. Copy it for your own field.
Step 1. Estimate variation. Take last year's cleaned yield map for the field and calculate the standard deviation among combine passes of equal length. Suppose the field averages 13.8 t/ha (220 bu/ac) with a coefficient of variation of 5%. Your standard deviation is 0.69 t/ha (11 bu/ac).
Step 2. Apply the LSD formula. For two treatments in a randomized complete block design:
LSD = t × s × √(2 ÷ r)
where r is the number of strip pairs, s is the standard deviation, and t is the two-tailed critical value at your chosen significance level with (r − 1) error degrees of freedom. On-farm research commonly uses α = 0.10, which corresponds to 90% confidence.
Step 3. Read the result.
Strip pairs
Error df
t (α = 0.10)
LSD
4
3
2.353
1.15 t/ha (18 bu/ac)
6
5
2.015
0.80 t/ha (13 bu/ac)
8
7
1.895
0.65 t/ha (10 bu/ac)
12
11
1.796
0.51 t/ha (8 bu/ac)
Six strip pairs — a substantial trial by most standards — cannot separate anything smaller than 0.80 t/ha (13 bu/ac). To bring the LSD down to 0.25 t/ha (4 bu/ac) in this field, you would need roughly 43 strip pairs. On 24-row equipment that is more than 200 hectares of one field devoted to one question.
This is the arithmetic behind the thesis. Within-field replication has sharply diminishing returns because the LSD falls with the square root of r. Doubling your strips from 6 to 12 buys you a 37% reduction in the detectable difference, not half.
Replication alone is only half the protection. Piepho and colleagues set out the requirement in their 2011 review of on-farm experiment design: randomization, replication, and blocking work together. A systematic alternating layout — treated, untreated, treated, untreated across the whole field — can line up with a drainage or soil texture gradient and reintroduce the bias you were trying to remove. Randomize the order within each pair, and record which strip received which treatment before harvest rather than reconstructing it afterward.
Your field is one replicate, not the experiment
Small effects become detectable across sites, not within them. This is the single change that most improves on-farm research trial design, and it costs less than adding strips.
Kandel and colleagues quantified the tradeoff in a 2018 Plant Disease study comparing 230 on-farm trials against 49 small-plot trials of foliar fungicide in Iowa soybean. With 12 locations, three replications per location were enough to detect a 134 kg/ha response in the on-farm trials. Detecting 67 kg/ha — about 1 bu/ac — required 15 replications at each of those 12 locations. The site count does the heavy lifting; the within-field replication mainly protects each site's estimate.
That study also produced a result worth sitting with: the on-farm strips were less noisy than the research station plots. Laurent and colleagues confirmed it at scale in Agronomy for Sustainable Development in 2022, comparing 479 on-farm soybean trials against 83 small-plot trials and finding similar mean responses and similar within-trial variability, with higher between-trial variance in the small plots. Long field-length strips average over more of the field than a 10-metre research plot does.
The same principle holds where equipment and field sizes are entirely different. Lark, Manzeke-Kangara, Kihara, and Broadley modeled biofortification trial networks for Ethiopian cereal systems in npj Sustainable Agriculture in 2025 and found that replication at farm scale — treating each farm as a complete block — was what delivered statistical power of 0.8 or better. Spreading single unreplicated plots across many farms did not. CIMMYT's mother-baby trial design, formalized by Snapp in Malawi more than two decades ago, encodes the same logic: one replicated reference trial anchors many farmer-managed satellites.
A Uruguayan example shows the payoff. Researchers pooled 448 on-farm strip trials conducted from 2014 to 2024 across soybean, rice, maize, wheat, and barley to evaluate a humic biostimulant, and resolved mean yield responses between 7.6% and 15.7%. No single one of those 448 fields could have settled the question.
For a co-op agronomist, this reframes the season. Instead of one large showcase trial, run an identical protocol on eight cooperator fields with four strip pairs each: same product, same rate, same application timing and growth stage, same cleaning rules. Then analyze site as a random effect, so the treatment mean is estimated across environments rather than within one. Thirty-two strip pairs distributed this way carry far more weight than 32 strips in a single field, and they answer the question a farmer is actually asking — whether the practice works on farms like theirs, not whether it worked on one particular quarter section. The added cost is coordination, not land.
Strip width is set by your equipment, not your plan
Each treatment strip must be at least one full combine header width, and the strip width must also be a whole multiple of your applicator or planter width. Miss either constraint and you contaminate every harvested pass.
Nebraska's protocols state the harvest side plainly: rows planted in each treatment must equal or exceed the combine head width. Illinois Extension's farmdoc guidance addresses the application side — unless product can be placed precisely to strips exactly as wide as the combine harvests, strips need to be wider than the combine. The on-farm strips in the Kandel study ran 18 to 55 metres wide and 200 to 800 metres long.
The practical target is two header widths. Harvest the centre pass for your data and discard the outer pass, which absorbs application overlap and edge effects. If your applicator is 36 metres and your header is 12 metres, use 36-metre strips and keep the middle 12.
Which yield monitor data cleaning steps change the conclusion?
Start-and-end-of-pass delay correction and overlap removal matter most, because both types of error concentrate at strip boundaries and headlands — exactly where treatments meet in a strip trial. Grain flow lag, minimum and maximum velocity filters, and unrealistic yield limits matter less but are cheap to apply.
Sudduth and Drummond at USDA-ARS reported in Agronomy Journal in 2007 that their Yield Editor filters removed 13 to 27% of raw points, and cited earlier work removing anywhere from 10 to 50%. Uncorrected start and end delays leave erroneously low points in the dataset, which pull the treatment mean down and inflate the standard deviation. Both effects push directly into your LSD.
One discipline separates cleaning from data manipulation. Choose your cleaning parameters before you look at treatment means, apply the identical parameters to every strip, and record what you used and what percentage of points you removed. Tuning filters after seeing the result is how a non-significant trial becomes a significant one.
What a defensible trial report contains
Six items, none optional: the number of replications and error degrees of freedom; the coefficient of variation; the LSD with its significance level stated; the cleaning parameters and the share of points removed; the observed difference printed next to the LSD; and a sentence naming the smallest difference the trial could have detected.
That last item is the one extension services and co-op agronomists most often leave out, and it is the one that protects you. A trial that reports "we could not have detected less than 0.80 t/ha, and we observed 0.25 t/ha" is a useful contribution. A trial that reports "4 bu/ac advantage" without an LSD is not evidence, however carefully the strips were driven.
Valora Earth's Experiments feature calculates the LSD for each trial and displays it alongside the treatment means, so a difference that falls below its own detection threshold is visible before the result is shared with farmers.
Frequently asked questions
How many strips do I need for a valid on-farm trial? At minimum four to six randomized, replicated strip pairs, which will detect differences of roughly 6–10%. Smaller effects require pooling across sites. Nebraska's On-Farm Research Network sets three replications as its floor and specifies five for higher-variability questions such as sulfur response.
Can I use yield monitor data instead of a weigh wagon? Yes, provided the monitor is calibrated for the season and crop, and the data are cleaned with documented filters applied identically to all treatments. Uncleaned yield monitor data typically contain 10 to 50% erroneous points, concentrated at pass ends and overlaps where strip trial boundaries fall.
What does "no significant difference" actually mean? It means the trial could not separate the treatments from background variation — not that the treatments are identical. Report the LSD alongside the result so readers can see what the trial was capable of detecting. A small observed difference with a large LSD is an inconclusive trial, not a negative one.