A trained model with high accuracy is a result, not a project. What converts it into research is the comparison structure around it: a baseline that says what the number means, an ablation that says which part earned the performance, an error analysis that says where it fails, and a documented account of where your data came from. Everything else is a demo.
The 94% problem
Every fair season produces a wave of projects built the same way. A student finds a public dataset, fine-tunes a model, reports a headline accuracy figure, and builds a board around it. The work is real — the code runs, the number is honest — and yet at the judging table it dissolves in about ninety seconds, because a judge asks one of two questions and neither has an answer.
“What would a simple method get?” If 94% of the dataset is one class, then always guessing that class gets 94%, and your model has learned nothing. If a one-line logistic regression on the same features gets 92%, your architecture bought two points and a lot of complexity.
“Which part of your system produced the improvement?” If you changed the model, the preprocessing, the augmentation and the loss function all at once, you have one data point about a bundle, and no knowledge about any component.
Neither question is hostile. They are the same question a reviewer asks of any paper in the field: compared with what, and how do you know? A machine learning project answers them by being designed as an experiment from the beginning, with the model as the treatment rather than as the result.
Which of the 22 categories does an AI project belong in?
This trips up more students than it should, because “artificial intelligence” is a method, not a field, and the competition organises by field. The Society for Science lists 22 categories, and several can host a machine learning project depending on what the project is actually about. The names and codes below are as published on societyforscience.org; confirm the current list and the full subcategory descriptions there before you commit, since categories are revised between seasons.
| Category | Fits when your project is really about… | Typical giveaway |
|---|---|---|
| Software Design (SFTD) | The algorithm, architecture or software system itself — you are making a computational method better | Your contribution is a technique that would apply to other datasets |
| Computational Biology and Bioinformatics (CBIO) | A biological question answered computationally — sequences, structures, omics, biological signals | Your novelty claim is biological, and the model is the instrument |
| Robotics and Intelligent Machines (ROBO) | A physical machine that senses, decides and acts in the world | There is hardware, and it moves or manipulates something |
| Embedded Systems (EBED) | Computation running on constrained hardware — microcontrollers, edge devices, real-time systems | Power, latency or memory limits are part of the problem |
| Translational Medical Science (TMED) | Moving a clinical or diagnostic idea toward practical use | The endpoint is patient or clinical impact, not model performance |
| Behavioral and Social Sciences (BEHA) | Human behaviour, cognition or social systems, studied with computational tools | Your conclusion is about people, not about the model |
| Mathematics (MATH) | The underlying mathematics — proofs, bounds, theoretical properties of a method | You prove something rather than measure it |
| Technology Enhances the Arts (TECA) | Technology in service of artistic creation or experience | The output is an artistic capability, judged as such |
The rule of thumb: ask what your sentence of novelty is about. “I built a better segmentation method” is a software contribution. “I found that this protein family behaves unexpectedly” is a biology contribution that happened to use software. Judges in a category are specialists in that category, so placing a biology finding in a software category means being questioned by people who want to hear about the algorithm you barely changed. We work through the general logic in how to choose your ISEF category.
Baselines and ablations: the two moves that make it an experiment
A baseline is the answer to “compared with what”. An ablation is the answer to “which part mattered”. Together they cost you a weekend of compute and they are the single highest-return work available to a machine learning project.

Ablations follow the same logic as controlled variables in a wet-lab experiment: change one thing, hold the rest constant, and see what happens. If your system has four components, run it four more times with one component removed each time. Then report a table, not a sentence — and be honest when the answer is unflattering. “Removing my custom attention layer cost 0.4 points, while removing the augmentation cost 6 points” is a genuine finding about your own system, and it is far more impressive than a headline number, because it demonstrates that you interrogated your own work.
Error analysis is the third move and the most neglected. Pull out the cases your model gets wrong, look at fifty of them by hand, and group them. You will nearly always find structure: a class it confuses, a lighting condition it cannot handle, a recording device that shifts the distribution. That paragraph — “here is what my model systematically fails on, and here is why I think it does” — is often the most scientifically mature thing on the entire board.
Where did your data come from? The question before the code
Data provenance is where computational projects meet the rules, and it is the part students most often discover too late. The Society’s human participants rules address research using pre-existing data, and the treatment differs by data type. Two points published on societyforscience.org are worth knowing before you download anything:
- Data or record review studies that use pre-existing datasets which are publicly available and/or published, with no interaction with human participants and no collection of data from them, are treated as exempt studies.
- Projects using pre-existing de-identified or anonymous data have additional conditions — including written certification by a qualified professional that the data were appropriately de-identified — and the affiliated fair’s Scientific Review Committee reviews for compliance.
Because the boundaries depend on the specific dataset and the specific study, do not reason by analogy from someone else’s project. Confirm the current rules on societyforscience.org and ask your affiliated fair’s committee before you begin, since some categories of work require review and approval prior to experimentation.

Separately from the rules, keep a provenance record for scientific reasons: dataset name, version, licence, download date, and any filtering you applied. Half the reproducibility failures in student machine learning projects come from nobody being able to say which version of the data produced which figure.
Reproducibility for a project with no test tubes
Computational projects still need a research record. The equivalent of a lab notebook for code is a combination of version control and a dated log, and judges increasingly know to ask for it.
- Fix and record your random seeds, and report results across several seeds rather than the best one. A single lucky run is not a result, and reporting the range is a maturity signal.
- Split the data once, before you start tuning, and never touch the test set until the end. If you tune against your test set, your reported number is optimistic and you cannot say by how much. Use a separate validation split for all decisions.
- Version the code with dated commits. Your commit history is your notebook: it shows what you did, when, and in what order — which is exactly what a judge wants when asking how the project developed.
- Record the environment. Library versions, hardware, training time. Silent version changes are a common cause of numbers that will not reproduce two months later.
- Log the failures. The approaches that did not work, with dates, are evidence of a real research process and are the most convincing thing you can show a sceptical judge.
The five interview questions weak ML projects cannot survive
Rehearse a specific answer to each of these. If any answer is missing, that is your remaining work list, and August is a good time to have it.
- “What does a trivial baseline get on this data?” Have the number ready.
- “Which component actually produced the improvement?” Show the ablation table.
- “Where does it fail, and why?” Show the grouped error analysis.
- “How many times did you run this, and how much does the result move?” Report variation across seeds or splits, not one number.
- “Whose data is this, and what did you have to do to use it?” Name the source, the licence, and the review determination for your project.
None of this requires a GPU cluster or a university affiliation. A laptop, a public dataset, one clearly stated question and a disciplined comparison structure will beat a larger model with none of the above, because the competition rewards the reasoning around the result rather than the size of the result. If you are new to how the whole season fits together, start with what ISEF is and the route from an affiliated fair to the finals in May.
Frequently asked questions
Which ISEF category should a machine learning project enter?
It depends on your claim. Software Design for algorithmic contributions, Computational Biology and Bioinformatics for biological questions, Robotics or Embedded Systems when hardware is involved.
Do I need original data, or can I use a public dataset?
Public datasets are widely used. Record the source, version and licence, and confirm the review requirements for your project with your affiliated fair.
Is high accuracy enough to be competitive?
No. Without a baseline the number cannot be interpreted, and without an ablation you cannot say which part of your system earned it.
What replaces a lab notebook for a coding project?
Dated version-control commits, fixed random seeds, a frozen test split, recorded library versions, and a log of the approaches that failed.
This is an independent guide operated by Hanlin Education for China-based international-school students. We are not affiliated with, endorsed by, or sponsored by the Society for Science or Regeneron ISEF. Categories, rules, forms and review requirements are revised between seasons — confirm current details on societyforscience.org and with your affiliated fair before acting on anything here. Factual errors reported to us are corrected within 7 working days.