Sign in
Criterion C: Method · 4 marks≈ 500 words · suggestedSecondary data

Select and document your data

You arrive with

A research question, a place, a tension, and probably the file you opened back in step 1. Until now, that file showed the investigation was possible. Here it becomes the dataset you can defend: counted, documented, and cut down by a rule. It may not be the same file. Swapping it now is much easier than swapping it in step 6.

You leave with
  • The file you will actually use, opened, counted and understood
  • A protocol precise enough to rebuild your exact file
  • An inclusion rule you fixed before you looked
  • A licence, a credit and a date accessed

Your method is a set of instructions for getting the same file back.

This criterion wants an account precise enough that a stranger could go to the same source, apply the same filters, and end up with your dataset row for row. That is a higher standard than it sounds, but an easy one to meet: everything you need to write down is on the screen when you press download.

Two questions before you go any further. Both have to be yes.

Could someone rebuild your file?

Not find similar data. Rebuild yours: same source, same version, same filters, same rows. If a reader would have to ask you which dataset you meant, the answer is no.

Is there enough data?

Enough to answer the question you asked, and enough for the test you intend to run. On this route you chose the amount, so you have to justify it.

The way this route fails, and it fails here

An investigation that only reviews what other people have written is explicitly not a repeatable method. It is not a weak method: it is not a method at all. It fails the first gate outright and lands Criterion C (Method) in the 1 to 2 band however good the argument is.

The line between the two is not how much you read. It is whether you took unprocessed numbers and did something to them yourself. Quoting a conclusion from a paper is a literature review. Downloading that paper's dataset and reworking it is an investigation.

Everything below is how we suggest you actually do it.

Four tests, before you spend too much time on it

2 min

Most secondary investigations are decided in the first twenty minutes, by whether the thing you found is usable. These four questions take a couple of minutes each. They protect you from the one failure you cannot recover from: finding out at the end that your dataset could never have answered your question.

Is it numbers?
Not a picture of numbers

A chart on a page, a figure in a paper, a PDF of a table. None of those are data yet. You need values you can put in a spreadsheet and change. If the only way to get them is to read them off an axis, keep looking.

Can you say what one row is?
Out loud, in a sentence

One station in one year. One country in one month. If you cannot finish that sentence, you do not yet know what you would be comparing, and neither will your reader.

Does it cover your years and your places?
Both, not one

A strategy introduced in 2015 needs data either side of 2015. A question about a river needs stations on that river. Check before you write the question, not after.

Can you have it in ten minutes?
Today, not in two weeks

Some of the best environmental archives need an account and a rights request that takes two weeks to approve. That is fine for a researcher and far too slow for you. If you cannot have it on your screen now, it is not your dataset.

People argue most with the fourth, and it is the one I would insist on most. An investigation you cannot start is worth less than a smaller one you can finish, and the good archive will still be there next year.

Datasets that have already been through these four
The catalogue: one card per dataset, searchable by the issue it can support. The counted ones have been downloaded and taken apart, so each says what one row is, how many rows are really in the file, and what caught out the person who opened it.
Browse the datasets

Open it before you trust it

4 min

A downloaded file arrives looking authoritative: a government logo, a licence, a last-updated date. None of that tells you what is inside.

So before anything else: how many rows are there really, what is one row, which years are missing, and does the current year look complete? Ten minutes with a spreadsheet, and it changes what you can honestly claim.

Then find your two lines in it: the column for line 1, what people do, and the column for line 2, what happens to nature. If one of them is in neither this file nor a second one you can join to it by place and year, your investigation is not in this data yet.

What to check, in this order
The row count, then the count of things you actually have. They are not always the same number.
One row, named. Write the sentence down; you will need it in your method anyway.
Who is in it. Sort the names and look for the one you came for. Nothing tells you when something is missing: a file can lack a whole continent without saying so.
The years present. Sort them and look for gaps. The file usually does not tell you a year is missing.
The current year. It is often there but incomplete, which makes it look like a sudden drop.
The extremes. A value that is suspiciously round, or repeated at the top of the range, is often a reporting ceiling (the highest value the system records) rather than a measurement.
The zeros. Sometimes they mean none. Sometimes they mean below the detection limit, which is a different claim.
One row was not one sharkwhite shark study

The shark file has 23,118 rows and looks like a list of animals caught. It is not. One row is one kind of animal caught on one kind of gear at one beach in one month, and a column called NumberCaught says how many. Of the 4,299 rows my question uses, 591 record more than one animal, so counting rows undercounts the catch.

The zeros are missing in the same way. A month in which no white shark was caught has no white shark row at all, not a row with 0 in it. Nothing tells you this. There is no note and no blank row: the missing row is the zero, and only adding up NumberCaught handles it correctly.

And before either, the download. The portal only hands the file to a browser: a script that asks for it gets back an empty file. "The download worked" and "I have the data" are not the same thing when a server can do that, and opening the file to count its rows is the only way to tell them apart.

For your own investigation

Before you analyse anything, write down what one row is, and ask what a missing row means. Neither comes with a warning. On this route, the commonest way to be wrong is to count something other than what you think you are counting.

The mistake this protects you from

An absence cannot be seen in a summary statistic. Every mean, every median and every test will run perfectly happily on a file that is missing a third of the world, and the output looks exactly like the output of a complete one.

It is one of the two errors on this route a reader cannot detect from your finished report. The other is duplicated rows. They leave your averages correct but multiply your sample size, so the statistics come back significant whatever is in the data. Both are found by counting before you calculate, and both are worth a sentence in the method saying you did.

Could a stranger rebuild your file?

5 min

This is what repeatable means on this route, and it is a higher bar than it sounds. The question is not whether someone could find similar data. It is whether someone could follow your written method and end up with the identical file, without asking you a single question.

Everything you need to record is on your screen at the moment you download. Write it down then. If you try to work it out three weeks later, the date accessed becomes a guess.

Record thisBecause without it
The publisher and the dataset's own namea reader searches the wrong catalogue
The version, release or last-updated datethe file changes under you and nobody can tell
The indicator or layer codeone publisher has forty things with similar names
Every filter, query or search term you appliedthe same page returns a different file
The aggregation you chose: hourly, daily, monthly, annuala reader cannot tell whose mean they are reading, the portal's or yours
Any quality-control option you tickedthe same query with the box ticked returns fewer rows, and better ones
The date you downloaded itthe numbers may since have been revised
The file format and how many rows arrivedyour reader cannot tell whether they got the same thing

Where the data comes through a query rather than a download button, it is easier: paste the query. A URL that returns the data is the most repeatable method anyone in this course will ever write, because a reader re-runs it rather than reconstructing it.

What will not do is the name of a website. "Data from the Queensland Government Open Data Portal" is the kind of line where most secondary methods stop, and it is insufficient: one dataset page on that portal offers a catch file, a file of equipment locations and a file of biological information from caught animals. The source has to be precise and reachable. Two extra lines are the difference between passing and failing the first gate.

The test, in one sentence

Hand your method to somebody in another class. If they can produce your exact file, you have passed the first gate. If they come back and ask you which of the four datasets you meant, you have not.

Screenshot the download itself

Where the data comes out of a portal rather than a link, photograph the route: the search page with your filters set, the station or indicator list with your selection showing, the download button you pressed. Number them and put them in the method.

It costs nothing against the word count and it proves the method is repeatable better than prose can, because a reader can follow the pictures without knowing the site. ⚠️ This is not the same as a screenshot of somebody else's chart, which is the worst habit in step 5: this is a picture of your procedure, not of their results.

Past tense throughout. This is a record of what you did, including the dataset you tried first and abandoned. That abandoned attempt is worth a sentence: it shows you chose your dataset deliberately, not by luck, and step 7 will want it.

Four filters to write down, one column to buildwhite shark study

The extraction protocol is a page, a file and four filters. The page is "Queensland's Shark Control Program Data and Information" on the Queensland Government Open Data Portal. The file is "Number caught by area, calendar year and species group" (scp_numbers-caught.xlsx), last modified 19 June 2026, downloaded in a browser because the portal answers a script with an empty file. The method records the date and what arrived, 23,118 rows on 26 September 2026, then the filters: Year 1996 to 2023, SpeciesGroup SHARK, Gear Net or Drum, and Area Gold Coast, Sunshine Coast North, Sunshine Coast South and Rainbow Beach. That leaves 4,299 rows. There is no address a reader can re-run, so this is where numbered screenshots of the filters earn their place.

The harder half was a column the file does not have. The dependent variable is the share of each gear's sharks that were white sharks, and no column holds it. So the method builds it in steps a reader can repeat. A helper column P sorts every row: =IF(J2="WHITE SHARK", "White shark", "Other shark"), where J is CommonName. Four SUMIFS formulas, on a tab holding only the filtered rows, then fill the 2 by 2 table: =SUMIFS(O:O, F:F, "Net", P:P, "White shark") adds up NumberCaught for white sharks in nets and gives 73. Each gear's percentage comes from that table. It is written into the method, not left for the results, because without it nobody could rebuild the number the whole study rests on.

For your own investigation

Where the variable you analyse is not a column in the file, the steps that build it are part of your method and belong there, formula and all. A number a reader cannot rebuild makes the whole analysis unrepeatable, however precisely you describe the download.

Every number was made by somebody, somehow

3 min

The commonest weakness in a secondary investigation is treating the values as facts that simply existed, waiting to be found. They are the output of a process: someone chose a method, an instrument, a sampling frequency and a set of definitions, usually for a purpose that was not yours.

Read the methodology or metadata page: it is where most of your Criterion F (Evaluation) material comes from.

How the number was madeWhat it means for your claim
Measured directly, on siteThe strongest, and still has an instrument and a detection limit
Reported by an organisation or a countrySomeone had a reason to report it, and sometimes a reason not to
Modelled or estimated from other variablesYou are analysing a model's output, not an observation
Remotely sensed from satelliteExcellent coverage, but it classifies rather than sees: plantations can show up as forest
Imputed to fill a gapThe value exists because the gap did, not because anyone measured it

Two more common traps. Values get revised: the figure you downloaded in September may not be the figure published in March, which is why the date accessed matters. And a definition can change mid-series, so a step in your graph may be a change in what was being counted rather than a change in the world.

What a catch actually iswhite shark study

Not a count of sharks in the sea. It is a count of the animals the gear caught, so it depends on how many nets and drumlines were in the water, and for how long, and the file records neither. That is why my dependent variable is a share of each gear's own catch rather than a rate: the file cannot say which gear catches more white sharks a day.

It is not a death either. Caught means found on the gear when it was checked: 65 of the 147 white sharks in my four areas were alive when found. The file's Fate column says whether each animal was released, died or was killed.

And it is recorded by the programme being judged on it. The Shark Control Program sets the gear and publishes what it catches. That is not a reason to distrust the numbers. It is a reason to say in the method what kind of number they are, and to notice when the programme changes its own gear: from 2024 new locations, daily servicing and new drumlines came in, which is why my study stops at 2023.

For your own investigation

Find out what one value is, in the publisher's own words, before you average anything. Three properties matter most: what it is a measure of, who produced it, and what the units already have built into them. All three are usually one page away, and all three end up in step 7.

The CDN approach

Fix the rule before you look

4 min

The decision that matters is which rows to keep. Nothing stops you making it after you have seen which choice gives the nicer graph. The file will let you, and it leaves no trace.

That is the integrity problem specific to this route. Choosing your countries, years or stations after you can already see the relationship is the same as throwing away the readings that disagreed with you. It is just harder to notice, and nobody can tell from the finished report.

So write the rule down first, apply it, and say in the method what it removed, with the count: five rows caught on Other gear left out, one white shark among them, and here is why.

Decided afterwards

I analysed the areas where the difference between nets and drumlines was clearest.

Unrepeatable, and the finding is a product of the choosing.

Decided in advance, and reported

I kept the sharks caught in nets and on drumlines, 1996 to 2023, in the four areas that set both gears: Gold Coast, Sunshine Coast North, Sunshine Coast South and Rainbow Beach. From 2024 new locations, daily servicing and new drumlines came in, so 2023 is the last year before the gear changed. Five rows caught on Other gear were left out, one white shark among them.

A reader can re-run it and get your file back.

One choice, and it decides the answerwhite shark study

There is only one real decision in this method: which places to compare. Everything else follows. And it does not just change the answer, it reverses it. In the four areas that set both gears, white sharks were 2.2% of the nets' sharks and 3.8% of the drumlines'. Across the whole state, over the same years, they were 1.8% of the nets' sharks and 0.9% of the drumlines'.

So a student looking for a particular result could have picked either and given no reason for it. The reason used here was fixed before looking at the white sharks, and it is about the gear, not the answer: drumlines are also set far north, from Cairns to Bundaberg, where white sharks do not swim. A statewide comparison compares places as well as gear. Keeping to the four areas that set both takes place out of it.

Look closely at the statewide figure. It is not wrong: across Queensland, white sharks really were a larger share of the nets' catch. It answers a different question, about where each gear is set, and it would have seemed to support my hypothesis that nets catch the larger share. Both findings are true. Which one you report depends entirely on a choice you have to be able to defend.

For your own investigation

Where your method has one decision in it, make that decision for a reason you can state before you see the result, and show the reader what the other choice would have given. Putting both answers side by side is the simplest proof that you did not pick the places that gave the answer you wanted.

Rules worth considering
A minimum amount of underlying data behind each value
A fixed date window, chosen because of your strategy rather than your results
Only units present throughout, so early and late are the same set of places
A geographic boundary you can defend, such as one catchment or one country group
A habit worth copying

Write your inclusion rule at the top of the spreadsheet, in a cell, before you make a single chart. It takes twenty seconds and it turns a good intention into a record. When step 7 asks what you would do differently, that cell holds the honest answer.

Your controls are arithmetic now

3 min

You cannot hold things constant physically: no same time of day, same observer, same instrument. What you can do is hold them constant mathematically, which is genuinely powerful and almost never used.

There are three moves, and naming which one you used is what turns a comparison into a controlled one.

Normalise
Divide the difference out

Per person, per square kilometre, per unit of output. Comparing two countries' total emissions mostly compares their populations; comparing emissions per person compares what you actually meant to.

Restrict
Narrow until they match

Only inland cities, only sites below 200 m, only stations on the same river. You lose sample size and you gain a comparison where the confounding variable cannot vary.

Pair
Compare like with like

The same places early and late. Upstream against downstream on the same watercourse. Pairing removes everything that is constant within a pair, which is usually most of what worries you.

Normalising, and what it rescuedwhite shark study

The first comparison was of raw counts, and it was misleading. In the four areas the nets caught 73 white sharks and the drumlines 74, which looks like no difference at all. But the nets caught 3,296 sharks in all and the drumlines 1,932, so a count of white sharks mostly compares how much of everything each gear caught.

Normalise. Each gear is measured against its own catch instead: not how many white sharks it caught, but what share of its sharks were white sharks. That gives 2.2% for the nets and 3.8% for the drumlines, two numbers measured the same way, and the two gears can sit in the same test.

Restrict did the rest. Only the four areas that set both gears, because drumlines are also set far north where white sharks do not swim, and only 1996 to 2023, because new locations, daily servicing and new drumlines came in from 2024. Each restriction removes a difference between the two gears that is not the gear itself.

For your own investigation

Ask what else differs between the things you are comparing besides the thing you are studying, then divide it out rather than hoping it is small. Measuring each group against its own total is the simplest version: one division, and it works in almost every dataset.

Unidentified sharks need a rule too

Not every shark in the file is named to species. A name with a star is a group, not a species: HAMMERHEAD SHARK * is a hammerhead nobody identified further. There is also UNKNOWN SHARK, 18 animals in my four areas.

Filter them out and both gears' totals of other sharks shrink, so both white-shark shares rise. I kept them, because they are sharks and none is recorded as a white shark, and the method says so. What it cannot say is what they were: an unidentified shark could have been a white shark, and that goes to step 7 as a limitation.

Sort, filter and count in Google Sheets
How to filter a file to your inclusion rule and check what is left, and the trap that makes a filtered average wrong without an error.
Open the spreadsheet moves

How much data is enough, when you chose it

3 min

Here the question is different. Sample size is not a limit you ran into. It is a decision you made, so you have to justify it, not apologise for it, whether it is large or small.

The floor is set by the test you are going to run. Decide the test first, then count backwards to how many rows you need to keep.

The test you plan to runRows it needs
A correlation10 pairs at least, 30 to be comfortable
A t-test10 or more per group
Chi-squared5 expected in every cell
The trap that is specific to this route

There is such a thing as too much. Run a correlation on ten thousand rows and almost anything comes out statistically significant, including relationships far too weak to mean anything. A p-value answers whether an effect exists, not whether it is big enough to care about.

So when your n (sample size) is in the thousands, quote the strength (r or R squared) alongside the significance. Talk about the size of the effect, not how small the p-value is. A moderator who sees p less than 0.001 on a correlation of 0.04 knows exactly what happened.

Bigger is not automatically better in the other direction either. Adding thirty more countries to reach a threshold, when twenty of them are not comparable with your original ten, gives you a bigger number but weakens your argument. Sample size that comes from widening your inclusion rule has to be defended against the rule, not against the total.

What was enough here, and what was notwhite shark study

Twenty-three thousand rows sounds enormous. The analysis has four numbers in it: white sharks and other sharks, in nets and on drumlines. The unit being compared is an animal, not a row, so the 4,299 rows that pass the rule are added up through NumberCaught into 5,228 sharks: 3,296 in nets and 1,932 on drumlines.

Then the rare thing sets the limit, not the total. White sharks are 147 of the 5,228, and chi-squared needs 5 expected in every cell. Pooled across the four areas that is easily met: every expected count is above 50, the smallest about 54. Split by area it thins fast: Sunshine Coast South's nets caught 3 white sharks in the whole period. So the four areas are tested together. Each area's two shares can still be shown, and drumlines have the larger share in all four, but they are not tested one by one.

One exclusion, and it is worth its line: five rows caught on Other gear, one white shark among them, left out because the question compares two gears. Unidentified sharks stay in with the other sharks, and the method says so.

For your own investigation

Count the things your test actually compares, not the rows in the file. Then find the rarest thing you are counting: the smallest cell in your table, not the total, decides whether the test is valid.

Ethics, when nobody gets wet

1 min

There is no risk assessment to write here and no consent form to hand anybody, which tempts students to skip the section entirely. Two things belong in it instead, and one of them is the most serious integrity question on this route.

The first is quick: say whose data it is, under what licence, and credit them. Open data still has to be credited, and a dataset is cited like any other source, plus the date you accessed it.

The second matters more: selective use of data is misconduct, not just carelessness. The protection is the rule you fixed before you looked, applied to everything and reported with what it removed.

What a dataset citation carries
Who published it, and the dataset's own title
The version, release or last-updated date
The URL you actually used
The date you downloaded it
The licence, where one is stated
One more, if people are in your data

Some open datasets describe individuals rather than places: health records, survey microdata, individual incomes. Those come with conditions, and the conditions still apply even if the file downloaded easily. If a dataset would let you identify a person, it does not belong in a school investigation.

Using AI at this stepLevel 3 · Targeted AI

It can explain what an indicator measures, what a licence permits, or how to write the spreadsheet formula that turns your raw column into the one you need. Ask it as many times as you like.

It cannot give you the numbers, or tell you a dataset exists. An invented figure looks exactly like a real one once it is in your spreadsheet. Every value comes out of a file you downloaded, and every dataset is one you have opened.

What this level means

Ready for step 5?

Secondary data checklist0 of 18

The first two are the criterion itself. Fixing your rules before you looked, and checking the file's shape, are the two that make a secondary investigation defensible rather than just convincing, and neither takes more than a few minutes.

Next: step 5, treat your data

You have a file and a rule for what is in it. Step 5 is where you turn its columns into something that answers your question, and where the good news arrives: tables, graphs and calculations all sit outside the word count. Decide your statistical test before you start processing, because it sets how much data you needed to keep.

The white shark investigation used on the secondary-data route of this guide is the author’s own analysis of a published dataset: the Queensland Shark Control Program’s record of every animal caught on its nets and drumlines, published by the Queensland Government under CC BY 4.0 and downloaded on 26 September 2026. The choice of the four areas, the analysis and the conclusions are the author’s, not the Queensland Government’s.