Sign in
Criterion F: Evaluation · 6 marks≈ 600 words · suggestedSecondary data

Evaluation

You arrive with

A conclusion, and the discussion of how your data source could mislead that you wrote in step 6. Most of the raw material for this section is already on the page. What changes is what you do with it.

You leave with
  • Specific weaknesses of your data, ranked by how much they mattered
  • Your own selection rule, evaluated
  • An improvement for each, naming something real
  • Questions your investigation opened but could not close

Six marks for saying what was wrong with data you did not collect.

That sounds like an odd thing to be marked on. But much of it was your decision. You chose the source, the indicator, the years, the places and the rule that kept some rows and dropped others. Every one of those is a methodological choice, and this is where you weigh them.

The bands turn on whether your weaknesses are generic or specific. The test is mechanical: could this sentence have been written by someone who never opened your file?

Generic · 1–2

The data may not be completely accurate.

Specific · 5–6

Eighteen sharks in my four areas are recorded only as UNKNOWN SHARK, and I counted them as other sharks because none is recorded as a white shark. If any of them was one, the white shares I report are slightly too low. That cannot reverse my result: even if all eighteen were white sharks caught in nets, the nets' share would still be below the drumlines' 3.8%.

Generic · 1–2

More data would have improved reliability.

Specific · 5–6

The file records what each gear caught, not how many nets and drumlines were in the water or for how long. So my result is a share of each gear's own catch, not a rate. It can say that white sharks were a larger part of the drumlines' catch; it cannot say that drumlines catch more white sharks. More years of the same file would not change that, because the missing thing is a column, not rows.

Generic · 1–2

The dataset may contain errors.

Specific · 5–6

My method has one decision in it that decides the direction of my result: which areas to include. Across the whole of Queensland the answer reverses (nets 1.8%, drumlines 0.9%), because drumlines are also set from Cairns to Bundaberg, where white sharks do not swim. I fixed the four areas that set both gears before I looked at the white sharks, and I report the statewide figures too, so a reader can see the choice I did not take.

Everything below is how we suggest you actually do it.

Three strands, one chain

2 min

The most common mistake in structure is to write three separate lists: weaknesses, then improvements, then questions at the end, with no links between them. The descriptors link them clearly. Improvements must address the limitations you identified. Questions must relate to your conclusion.

Write it as a chain, and the marks follow. Take your two or three most important weaknesses and give each one the full chain: what it was, what it did to your answer, what you would change, and what that change would improve.

Two things do not belong here: the strengths of your method, and an application of your findings. Neither earns a mark, and together they can use up two hundred words that the marked strands needed.

The weakness

Specific to your method. And what it did to your conclusion, not just that it existed.

The improvement

Addressing that named weakness. Realistic for a student with a school's equipment.

What remains

A question with a different focus from your original, and a bearing on your answer.

Write it out, at length, in sentences

Two formats lose marks. A table of limitations against improvements looks organised, but it reads as a list and scores lower than extended writing. Bullet points of three words each also lose marks: they are what a section looks like when the word count has run out.

The opposite failure is padding. Plan the space before you start: this is the section where students most often have no words left.

If your data came from people

Questionnaires have three built-in weaknesses that the IB expects you to know about. Naming them precisely is better than vaguely mentioning "possible bias". Participation was voluntary, so the people who answered are not a fair sample of the people you asked. Strong opinions turn up more often than mild ones. Respondents tend to avoid the two ends of a Likert scale (a 1 to 5 agree-disagree scale), which squeezes your data toward the middle. And a long or complicated survey makes people tired, so the later answers are worse than the early ones.

Each of these earns marks only when you connect it to your conclusion: which of your results would change, and in which direction, if the missing people had answered?

The weaknesses are in the data, not in your hands

2 min

This criterion evaluates a method, and the method that produced your numbers was carried out by somebody else. So the temptation is to write about your analysis instead, where there is very little to say: a spreadsheet does not misread a scale.

What you evaluate instead is how the data was built: who chose the indicator, in how much detail it was measured, which places and years were left out, and which rows you yourself decided to keep. Step 6 discussed those weaknesses. This step weighs them, ranks them, and says what you would do about them.

Where the weakness livesThe question it answers
The indicatorDoes it really stand for the thing your question is about?
The resolutionIs it detailed enough, in space and time, to show what you claimed?
The coverageWhich places and years are missing, and are they missing at random?
The collectionWho produced it, by what method, and with what reason to be careful or careless?
Your own selectionWhat did your inclusion rule remove, and could that have made the pattern?

The last row is the one students leave out and the one that is most obviously yours. You chose the rows. That choice is part of the method and it belongs in the evaluation of it.

If your line 2 is one individual, such as one tree recorded year after year, that is a weakness of the indicator, and a large one: what is true of one tree may not be true of the next. Name it, and say what it stops you claiming about the population.

And put a number on the uncertainty, taken from the documentation rather than from an instrument. How many significant figures? Is there a detection limit? Rounded to what? A series reported to the nearest whole unit cannot support a difference of half a unit, and saying that in a sentence is worth more than a paragraph about data quality in general.

Test the limitation instead of asserting it

3 min

This is the best thing you can do on this route. You still have all your data, and re-running an analysis takes minutes.

Take the limitation you think matters most and rerun under a different assumption. Change the inclusion rule. Drop the years you were least sure about. Use the median instead of the mean. Then say what happened. "I tested this and the conclusion held" is what "evaluates" means as opposed to "describes", and it is worth more than any amount of careful, uncertain wording.

Either result is useful. If the conclusion survives, you have evidence it is robust. If it does not, you have found something important. Reporting that honestly is worth far more than a finding that secretly depended on one choice.

A sensitivity test that failed, and what that was worthwhite shark study

The strongest objection to this study is that the two gears do not catch their sharks in the same months. White sharks pass in winter and spring: 124 of the 147 in my four areas were caught June to October. Nets catch most of their other sharks in summer, so only 30% of the nets' sharks were caught June to October, against 43% of the drumlines'. The obvious test is to rerun on June to October alone, and unlike most sensitivity tests, this one does not pass. Chi-squared falls from 11.6 to 0.8 (p about 0.38). In those months white sharks were 6.3% of the nets' sharks and 7.3% of the drumlines': close, and not significantly different.

That is not a disaster, and explaining why turns it into the most useful paragraph in the report. It turns a finding about gear into a finding about timing. The nets' big summer catch of other sharks dilutes their white-shark share, so most of the difference comes from when each gear catches its other sharks, not from which sharks it catches. So 3.8% against 2.2% is a statement about the whole year, and cannot be quoted as a statement about the gear.

The second test was the place. Rerun across the whole of Queensland and the answer reverses: nets 1.8%, drumlines 0.9%, chi-squared 19.7. Drumlines are also set from Cairns to Bundaberg, where the catch is tiger sharks, whalers and hammerheads and there are no white sharks, which drags the drumlines' share down. That is place, not gear, and it is why I fixed the four areas before I looked at the white sharks.

Those two tests also rank my weaknesses for me. Season comes first, because it explains most of the difference. Place comes second: it reverses the answer statewide, but my rule already keeps the north out. Having no data on how much gear was in the water comes third, because it limits what the result can claim rather than changing it. The eighteen unidentified sharks come last, because they cannot reverse it.

For your own investigation

A test that fails has told you what your result is really about. Report it, and let it change your conclusion, rather than burying it where a moderator will find it for you.

Improvements that name something

2 min

An improvement that is only stated earns the bottom band. Say what it would achieve and you are in the top one. That takes two extra clauses per improvement.

This route has one typical mistake, and examiners know it: "use more data". More countries, more years, a bigger dataset. It is not an improvement. It is a wish, and it addresses no weakness you actually named. A real improvement names a specific dataset, a finer resolution, or a second source that would settle something.

Keep them feasible for a student. A different open dataset, a finer time step, a second indicator to cross-check with, a shorter and better-matched window.

ImprovementWhat it would achieve
Compare the gears month by month instead of pooling the yearTakes season out of the comparison, the weakness that explained most of my result, without needing any new data
Pair the 15 beaches that set both gearsCompares the gears at the same beach, so place cannot stand in for gear even inside the four areas
Join the programme's equipment locations file, from the same dataset pageDivides each gear's catch by how much of that gear was in the water, turning a share into a rate. It lists where the gear is now, so it would have to be matched to the years it covers
Add the biological information file from the same page, which records lengthsTests whether the two gears catch white sharks of different sizes, which a count of animals cannot show

The last two are not "get more data": they are files the same publisher put on the same page as mine, each answering something my file cannot. Look at what sits next to your file before you look anywhere else. That is often the highest-value improvement on this route, and the one students never think of.

Questions you opened but could not close

2 min

These have to point somewhere new. "What if I had more data?" is an improvement written as a question, and the guidance says exactly that.

A real one has a different focus from your original question, arises from something you actually found, and affects how far your conclusion reaches. The best ones come out of the part of your result that did not fit.

Then two simple rules that give easy marks. Give them their own subheading, called Unresolved questions, and write them as questions, with question marks. A moderator cannot give marks for a strand they cannot find, and headings do not count against your word limit.

Two that go somewherewhite shark study

White sharks were alive when the gear was checked more often on drumlines (37 of 74) than in nets (28 of 73), a difference I never tested. Does a white shark caught on a drumline survive more often than one caught in a net? If it does, the larger share on drumlines would not mean more white sharks harmed, and that bears directly on which gear the programme should set.

I stopped at 2023 so that the gear stayed the same through my years. From 2024 the programme added new drumlines and new locations and began servicing them daily, and its shark catch rose from 948 in 2023 to 1,497 in 2024 and 3,430 in 2025. What did that expansion do to the white-shark catch? My result describes the programme before it, not the one in the water now.

For your own investigation

Look at what your own inclusion rule left out: the columns you never used and the years you cut off. Questions that come from there have a different focus from yours and bear on the strategy, and because the answer is often in the same file, somebody could actually go and answer them.

Using AI at this stepLevel 1 · AI Planning

It can be argued with once your own list exists. Ask it to attack a limitation you have already named and quantified, and see whether it holds.

It cannot generate the limitations. Ask a tool what is wrong with a dataset it has never opened, and you get the five things wrong with every dataset. That is exactly what generic means, and it is the bottom band. The marks here are for what you found by looking at your file.

What this level means

Ready for step 8?

Secondary data checklist0 of 18

The one students miss is evaluating your own inclusion rule.

Next: step 8, write it up

Every section now exists in some form. Step 8 is assembly, referencing and the word count. Aim for 2,900, not 3,000.

The white shark investigation used on the secondary-data route of this guide is the author’s own analysis of a published dataset: the Queensland Shark Control Program’s record of every animal caught on its nets and drumlines, published by the Queensland Government under CC BY 4.0 and downloaded on 26 September 2026. The choice of the four areas, the analysis and the conclusions are the author’s, not the Queensland Government’s.