A prototype that works is not an ISEF result. In our coaching experience, what separates engineering entries is evidence: a numeric specification, a baseline you measured yourself, repeated trials under fixed conditions, and honest failure analysis. The device switching on at the table proves only that it switches on. This guide turns a build into data a judge can score.
The demo trap
Engineering projects fail differently from experimental ones. The student has built something real, often over months, and it works. Then the first judge asks two questions — compared with what, and how do you know? — and the project has no answer, because nobody wrote down what “better” was supposed to mean before the soldering started.
The demo trap is seductive because a working device gets applause from teachers, parents and classmates. Judges are a different audience: they are looking for a claim, a measurement, and a reason to believe the measurement. A robot arm that lifts a bottle on stage but has never been tested 100 times, against anything, under controlled conditions, is a demonstration. A robot arm with a success rate, a comparison unit and a documented set of failure modes is a project.
Remember also that your first panel is not the ISEF final's panel. Your selection fair judges you months earlier, sometimes on its own criteria — see how ISEF works from affiliated fairs to the global finals — so the evidence has to exist by that earlier date, not by May. The official judging criteria and their weightings are published by the organiser; confirm the current version on societyforscience.org and read it before you design your tests, not after.

Write the specification and choose the baseline before you build
A specification is one paragraph with numbers in it. It states the metric you care about, the threshold that counts as success, the conditions the device has to work in, and any constraint you accept (cost, mass, power, size). Writing it in August is the single highest-return hour of an engineering season, because a spec makes your project falsifiable — and falsifiable projects are the ones that produce results worth defending.
| Vague goal | Specification with numbers | How you will measure it |
|---|---|---|
| Build a better robotic gripper | Lift 500 g objects in three shapes with at least 95% success over 100 attempts per shape, cycle time under 2 s | 100 logged attempts per shape, randomised order, success defined in writing before testing starts |
| Make a cheaper water filter | Remove at least 90% of turbidity from 1 L samples at room temperature, at a materials cost under a stated per-unit budget | Turbidity meter with stated resolution, three samples per run, five runs, untreated sample as control |
| Improve solar panel output | Raise daily energy yield by at least 8% against a fixed panel of the same area, over 14 consecutive days | Two identical panels side by side at one site, logged every 10 minutes, weather recorded |
| Make an AI app for diagnosis support | Cut false negatives from the untuned baseline to under a stated ceiling on a held-out set of labelled cases | Held-out set fixed and locked before any tuning; baseline is the untuned model on the same set |
Two rules keep a specification honest. First, the threshold is written before the first test, not adjusted afterwards to match what you got. Second, if you cannot say how you would measure the metric with equipment you can actually access, the specification is fiction and needs rewriting now rather than in April.
Choose the baseline before you optimise. “Better” is a comparative, so every engineering claim needs something to compare against. You have three legitimate options, and the weakest of them still beats none.
- A commercial unit. Strongest, because judges can situate your work in the real world. Buy the cheapest comparable product and measure it yourself on your own rig — do not quote the manufacturer's figure, which was measured under conditions you cannot verify.
- A published value. Acceptable when a device is out of reach, but you must cite the source, state the conditions it was measured under, and acknowledge that the conditions differ from yours.
- Your own version one. Perfectly valid and often the most instructive, because the comparison isolates the change you made. It only works if version one was measured properly before you improved it — which is why you measure early, even when the first build is embarrassing.
Build the test rig, not just the device
Students often spend most of their time on the device and very little on the measurement, then wonder why the numbers wobble. Invert some of that. A test rig is the boring apparatus that makes results comparable: the fixture that holds the sample the same way every time, the fixed distance, the controlled temperature, the same battery charge state, the instrument whose resolution you know.
- Fix what you are not testing. One variable moves; everything else is pinned down and written down.
- Know your instrument. If your meter reads to 0.1 units, you cannot claim a 0.05 improvement. Judges notice claims finer than the tool that produced them.
- Decide the number of trials in advance, and randomise their order so drift — a warming motor, a draining battery, a tiring human operator — does not masquerade as an effect.
- Blind anything subjective. If a human scores the output, the scorer should not know which version produced it.
- Define failure before testing. Write down what counts as a failed attempt, or you will unconsciously redefine it mid-experiment.
Iterate on the record: the design log
The engineering-design cycle is not decoration. Its value at the judging table is that it turns your project into a sequence of decisions you can explain, and explaining decisions is what the interview is for. Log each loop in five lines: what you predicted, what you changed, what you measured, what failed, and what you did next.

Judges reward failures that were understood. “Version two overheated after eleven minutes because the heat sink contacted the casing; version three moved the mount and ran for forty” is a better answer than “it works”. Keep the failed parts — a burnt component in a labelled bag is more persuasive than a sentence about it, and it also demonstrates that the record is real rather than reconstructed the week before the fair.
Claims that survive a judge
| Weak claim | Why it cannot be scored | Scoreable version |
|---|---|---|
| “Our device works well.” | No metric, no comparison, no conditions | “Across 100 attempts per shape it succeeded in 96, 94 and 89, against 71, 66 and 52 for the unit we compared against on the same rig.” |
| “It is 30% better.” | Thirty per cent of what, measured how, and how much did it vary? | “Mean daily yield rose 30%, from 1.0 to 1.3 units per day over 14 days, with a day-to-day range of 22% to 38%.” |
| “It is cheap.” | No bill of materials and nothing to compare the cost with | “Bill of materials of 12 listed parts, against the cheapest commercial unit meeting the same specification, priced on the same day.” |
| “The model is accurate.” | Dataset, split and baseline all unknown | “Recall of 88% on a held-out set locked before tuning, against 71% for the untuned baseline; errors concentrate in low-light images.” |
Five patterns we see most in China-based engineering builds
These are the recurring failure modes in our own coaching, de-identified and generalised; individual outcomes vary widely.
- Part substitution without re-testing. A component is unavailable, a near-equivalent is ordered, and nobody repeats the earlier measurements. The dataset now mixes two devices, and a judge who spots the discontinuity has lost confidence in all of it.
- Simulation only, or build only. Either extreme is fragile: a model with no physical test invites “did you ever build it?”, and a build with no model invites “why does it behave that way?”. Pair them, even at small scale.
- One successful run presented as a result. Usually the best run out of many. Report all attempts, including the ones that failed, and the honesty itself becomes a strength.
- “Improvement” with no baseline. The commonest single flaw, and the cheapest to fix — buy or build the comparison, and measure it on your own rig.
- Category mismatch. An engineering build is sometimes entered where the actual contribution is not engineering at all, or the reverse. Decide what your contribution really is, then use our guide to choosing your ISEF category and confirm your fair's own category list, since the final runs 22 categories but a selection fair may group subjects differently.
If you are early enough that none of this has been decided yet, that is the best position to be in: the specification, the baseline and the rig are all cheap in August and expensive in April. Our overview of what ISEF is and how the season runs sets out the wider timetable a build has to fit inside, including the abstract limit of 250 words that will eventually force you to say in one paragraph exactly what you measured and what it means.
Frequently asked questions
Does my engineering project need statistics?
It needs repeated trials and honest variation. Report how many attempts, how much they varied, and what counted as failure before testing began.
Can I use a manufacturer's specification as my baseline?
Prefer measuring the comparison unit yourself on your own rig. If you must cite a published figure, name the source and its test conditions.
Is a project still valid if the prototype never met the target?
Yes. A specification you missed, measured properly and explained, is stronger than an unmeasured device that appeared to work.
How many design iterations do judges expect?
There is no set number. What matters is that each iteration has a documented prediction, change, measurement and failure analysis.
This is an independent guide operated by Hanlin Education for China-based international-school students. We are not affiliated with, endorsed by, or sponsored by the Society for Science or Regeneron ISEF. Specifications, claims and figures in the tables are illustrative worked examples, not results from any particular project, and coaching observations are de-identified and vary by student. Rules, categories, judging criteria, required forms and deadlines change every cycle: confirm current details on societyforscience.org. Errors reported to our editorial desk are corrected within 7 working days.