Your complete guide to the Regeneron International Science & Engineering Fair

ISEF Experimental Design (2027): Controls, Replicates and the Decisions You Make Before Any Data

No amount of clever analysis rescues a badly designed experiment. Before you collect a single data point you have to fix five things: what exactly you are measuring, what you are holding constant, what your control is, what counts as one independent unit, and how many of those units you need. Get those wrong in August and no statistics in April will fix it.

Design is where projects are lost, and it is lost quietly

Analysis errors are visible. You can spot a wrong test, a missing error bar, an overclaimed p-value — and you can usually fix them, because the data still exists. Design errors are different. They are invisible while you work, they feel like productivity, and by the time a judge exposes one your season is over, because the fix requires collecting the data again.

The pattern is consistent enough to name. A student spends three months collecting an impressive quantity of data, produces beautiful charts, and then meets a judge who asks one question — “how do you know it was the treatment and not the position on the windowsill?” — and there is no answer, because the treated plants were all on the windowsill and the controls were all on the shelf. The data is fine. The design cannot separate two explanations, so the data cannot mean anything.

The good news is that this is a small, learnable checklist, and August is exactly the right time to work through it. If you are still assembling your season plan, our overview of how ISEF works sets out where design sits relative to the approval and fair calendar. Note that some project types must be reviewed and approved before experimentation begins — confirm what applies to your project on societyforscience.org and with your affiliated fair, because starting work early on a project that needed prior approval is not a mistake you can undo.

Step one: operational definitions, or how to stop measuring a word

“Does music affect plant growth?” is not yet an experiment, because none of the three key words has a measurable definition. An operational definition converts each variable into a procedure that another person could repeat identically.

Fuzzy version Operational definition What it fixes
“Music” A 30-minute recording at 65 dB measured 20 cm from the pot, played 09:00 daily Makes the treatment repeatable and the dose stated
“Plant growth” Stem height in mm from soil surface to apical meristem, measured every 48 h with the same ruler by the same person Removes measurement drift and observer variation
“Affects” A difference in mean height at day 21 large enough to matter biologically, tested against the variation between plants Forces you to define your outcome before you see the data
“Water quality” Dissolved oxygen in mg/L by the same probe, calibrated weekly, sampled 30 cm below surface at 08:00 Pins down time, depth and instrument, which all move the number
“Efficiency” Output energy divided by input energy over a 10-minute run at 25 °C ambient Stops the definition changing between conditions
Operational definitions turn a topic into a protocol. If a stranger could not repeat your procedure from the definition, it is not operational yet.

Write these down before you build anything. The act of writing “measured by the same person with the same ruler” is what makes you notice that last week your lab partner measured half the plants.

Step two: the four different things students call “the control”

This single word does four jobs in science, and conflating them is the most common conceptual error in school-level projects. Judges probe it deliberately, because knowing the difference is a reliable signal of whether a student designed the study or inherited it.

Four meanings of control in experimental design: controlled variable, control group, negative control, positive control, each with a definition and an example
Four distinct roles. The bottom pair — negative and positive controls — are what separate a working method from an unvalidated one.

The practical value of the bottom row is enormous and underrated. Without a positive control, a null result is uninterpretable: you cannot distinguish “the treatment does nothing” from “my assay never worked”. A judge who hears “we found no effect” will ask exactly this, and a student who can point at a working positive control has a genuine finding rather than a failed season.

Step three: what counts as one? The pseudoreplication trap

This is the design error that most often destroys otherwise excellent projects, and almost nobody is taught it before university. The question is deceptively simple: what is your independent experimental unit?

Comparison of pseudoreplication with thirty measurements from one tank giving n equals one, versus true replication with six independent tanks giving n equals six
Thirty measurements from one tank is one independent unit, not thirty. The diagnostic question is what a single failure would contaminate.

Apply the diagnostic to your own setup. Thirty leaves from one plant is one plant. Three hundred cells from one culture flask is one flask. Fifty survey responses from one class taught by one teacher is, for many questions, one class. Repeated measures are still worth taking — averaging them gives you a more precise value for that unit — but they belong inside the unit, not in the count of units.

If independent units are expensive, the honest response is not to inflate the count. It is to run fewer conditions with real replication and say so: “I could support six independent replicates, so I tested two concentrations rather than six” is a design decision a judge respects. Pretending thirty leaves are thirty samples is a design decision a judge catches.

Step four: randomisation, blinding and order effects

Once your units are defined, the remaining question is how systematic differences sneak in through the back door. Three cheap habits close most of the gaps, and none of them requires equipment:

  • Randomise assignment, and randomise position. Do not put treatment on the left and control on the right. Use a random number generator to assign both which unit gets which treatment and where each unit physically sits, then record the layout in your notebook. Position, light gradient, draught and bench vibration are all real effects that look exactly like your treatment.
  • Blind the measurement wherever you can. If you know which sample is the treated one while you are reading a colour, judging a behaviour, or deciding where the meniscus sits, your expectations will move the number. Relabel samples with codes and have someone else hold the key. This costs ten minutes and is one of the strongest credibility signals available to a school project.
  • Break up the order. If you measure all controls on Monday and all treatments on Friday, then day of week, instrument warm-up and your own growing skill are all confounded with treatment. Interleave conditions within each session.

Record all three choices in the lab notebook as you make them, with dates. A design decision documented before data collection is evidence; the same decision recalled in April is a claim.

Step five: the pilot run, and freezing the design

Give up two weekends in late summer to a pilot run at roughly one-tenth scale. The purpose is not to get results — you will throw the data away — it is to find out what breaks. Pilots reliably surface things no amount of planning does: the assay takes 40 minutes not 15, the sensor drifts after an hour, the seedlings die at that concentration, your measurement resolution is coarser than the effect you are chasing.

That last one is the most valuable output. If your instrument reads to the nearest millimetre and the difference you expect is a fraction of a millimetre, no sample size will save you, and you have found this out in September rather than March.

When the pilot is done, freeze the design and write it out in full: variables and their operational definitions, controls of all four kinds, unit of replication and how many, randomisation and blinding scheme, and the analysis you intend to run. Date it and do not change it once real data collection starts. Deciding how to analyse data after you have seen it is how honest students accidentally produce misleading results, and a dated pre-analysis note is the simplest defence there is.

Design decision Fix it before data The question a judge asks
Operational definitions Written procedure for every variable “How exactly did you measure that?”
Controlled variables Named list of what is held constant “What else changed between your groups?”
Control group Identical handling minus the treatment “What were you comparing against?”
Negative control Condition that should show nothing “How do you know that signal is real?”
Positive control Condition known to show something “How do you know your method works?”
Unit of replication Defined and counted honestly “What is your n, and why?”
Randomisation and blinding Scheme recorded with dates “Could you have influenced the reading?”
Analysis plan Written and dated before collection “Did you decide that before or after seeing the data?”
A pre-data design checklist. Every row is cheap in August and impossible to retrofit in April.

None of this requires a university laboratory or an expensive instrument. It requires deciding, in writing, before you start. That is the whole discipline, and it is the part of a project that transfers to everything you do afterwards. For readers still choosing where a project like this belongs, our guide to choosing your ISEF category covers the field-fit question separately, and what ISEF is covers the structure of the competition for students new to it.

Frequently asked questions

How many replicates does an ISEF project need?
There is no required number. What matters is that your replicates are genuinely independent and that you can justify the number you chose out loud.

What is the difference between a control group and a negative control?
A control group is your untreated comparison. A negative control checks your method itself, showing that a condition expected to give nothing really gives nothing.

Is a null result a problem for judging?
Not if your positive control worked. A validated method with no effect is a real finding; an unvalidated method with no effect is uninterpretable.

Can I change my design partway through the project?
You can, but record the change, the date and the reason, and treat the phases as separate datasets rather than silently pooling them.

This is an independent guide operated by Hanlin Education for China-based international-school students. We are not affiliated with, endorsed by, or sponsored by the Society for Science or Regeneron ISEF. Rules, forms, approval requirements and dates are revised between seasons — confirm current details on societyforscience.org and with your affiliated fair before starting experimentation. Factual errors reported to us are corrected within 7 working days.