Select and document your data
A research question, a place, a tension, and probably the file you opened back in step 1. Until now, that file showed the investigation was possible. Here it becomes the dataset you can defend: counted, documented, and cut down by a rule. It may not be the same file. Swapping it now is much easier than swapping it in step 6.
- The file you will actually use, opened, counted and understood
- A protocol precise enough to rebuild your exact file
- An inclusion rule you fixed before you looked
- A licence, a credit and a date accessed
Your method is a set of instructions for getting the same file back.
This criterion wants an account precise enough that a stranger could go to the same source, apply the same filters, and end up with your dataset row for row. That is a higher standard than it sounds, but an easy one to meet: everything you need to write down is on the screen when you press download.
Two questions before you go any further. Both have to be yes.
Not find similar data. Rebuild yours: same source, same version, same filters, same rows. If a reader would have to ask you which dataset you meant, the answer is no.
Enough to answer the question you asked, and enough for the test you intend to run. On this route you chose the amount, so you have to justify it.
An investigation that only reviews what other people have written is explicitly not a repeatable method. It is not a weak method: it is not a method at all. It fails the first gate outright and lands Criterion C (Method) in the 1 to 2 band however good the argument is.
The line between the two is not how much you read. It is whether you took unprocessed numbers and did something to them yourself. Quoting a conclusion from a paper is a literature review. Downloading that paper's dataset and reworking it is an investigation.
Everything below is how we suggest you actually do it.
Four tests, before you spend too much time on it
Most secondary investigations are decided in the first twenty minutes, by whether the thing you found is usable. These four questions take a couple of minutes each. They protect you from the one failure you cannot recover from: finding out at the end that your dataset could never have answered your question.
A chart on a page, a figure in a paper, a PDF of a table. None of those are data yet. You need values you can put in a spreadsheet and change. If the only way to get them is to read them off an axis, keep looking.
One station in one year. One country in one month. If you cannot finish that sentence, you do not yet know what you would be comparing, and neither will your reader.
A strategy introduced in 2015 needs data either side of 2015. A question about a river needs stations on that river. Check before you write the question, not after.
Some of the best environmental archives need an account and a rights request that takes two weeks to approve. That is fine for a researcher and far too slow for you. If you cannot have it on your screen now, it is not your dataset.
People argue most with the fourth, and it is the one I would insist on most. An investigation you cannot start is worth less than a smaller one you can finish, and the good archive will still be there next year.
Open it before you trust it
A downloaded file arrives looking authoritative: a government logo, a licence, a last-updated date. None of that tells you what is inside.
So before anything else: how many rows are there really, what is one row, which years are missing, and does the current year look complete? Ten minutes with a spreadsheet, and it changes what you can honestly claim.
Then find your two lines in it: the column for line 1, what people do, and the column for line 2, what happens to nature. If one of them is in neither this file nor a second one you can join to it by place and year, your investigation is not in this data yet.
The shark file has 23,118 rows and looks like a list of animals caught. It is not. One row is one kind of animal caught on one kind of gear at one beach in one month, and a column called NumberCaught says how many. Of the 4,299 rows my question uses, 591 record more than one animal, so counting rows undercounts the catch.
The zeros are missing in the same way. A month in which no white shark was caught has no white shark row at all, not a row with 0 in it. Nothing tells you this. There is no note and no blank row: the missing row is the zero, and only adding up NumberCaught handles it correctly.
And before either, the download. The portal only hands the file to a browser: a script that asks for it gets back an empty file. "The download worked" and "I have the data" are not the same thing when a server can do that, and opening the file to count its rows is the only way to tell them apart.
Before you analyse anything, write down what one row is, and ask what a missing row means. Neither comes with a warning. On this route, the commonest way to be wrong is to count something other than what you think you are counting.
An absence cannot be seen in a summary statistic. Every mean, every median and every test will run perfectly happily on a file that is missing a third of the world, and the output looks exactly like the output of a complete one.
It is one of the two errors on this route a reader cannot detect from your finished report. The other is duplicated rows. They leave your averages correct but multiply your sample size, so the statistics come back significant whatever is in the data. Both are found by counting before you calculate, and both are worth a sentence in the method saying you did.
Could a stranger rebuild your file?
This is what repeatable means on this route, and it is a higher bar than it sounds. The question is not whether someone could find similar data. It is whether someone could follow your written method and end up with the identical file, without asking you a single question.
Everything you need to record is on your screen at the moment you download. Write it down then. If you try to work it out three weeks later, the date accessed becomes a guess.
| Record this | Because without it |
|---|---|
| The publisher and the dataset's own name | a reader searches the wrong catalogue |
| The version, release or last-updated date | the file changes under you and nobody can tell |
| The indicator or layer code | one publisher has forty things with similar names |
| Every filter, query or search term you applied | the same page returns a different file |
| The aggregation you chose: hourly, daily, monthly, annual | a reader cannot tell whose mean they are reading, the portal's or yours |
| Any quality-control option you ticked | the same query with the box ticked returns fewer rows, and better ones |
| The date you downloaded it | the numbers may since have been revised |
| The file format and how many rows arrived | your reader cannot tell whether they got the same thing |
Where the data comes through a query rather than a download button, it is easier: paste the query. A URL that returns the data is the most repeatable method anyone in this course will ever write, because a reader re-runs it rather than reconstructing it.
What will not do is the name of a website. "Data from the Queensland Government Open Data Portal" is the kind of line where most secondary methods stop, and it is insufficient: one dataset page on that portal offers a catch file, a file of equipment locations and a file of biological information from caught animals. The source has to be precise and reachable. Two extra lines are the difference between passing and failing the first gate.
Hand your method to somebody in another class. If they can produce your exact file, you have passed the first gate. If they come back and ask you which of the four datasets you meant, you have not.
Where the data comes out of a portal rather than a link, photograph the route: the search page with your filters set, the station or indicator list with your selection showing, the download button you pressed. Number them and put them in the method.
It costs nothing against the word count and it proves the method is repeatable better than prose can, because a reader can follow the pictures without knowing the site. ⚠️ This is not the same as a screenshot of somebody else's chart, which is the worst habit in step 5: this is a picture of your procedure, not of their results.
Past tense throughout. This is a record of what you did, including the dataset you tried first and abandoned. That abandoned attempt is worth a sentence: it shows you chose your dataset deliberately, not by luck, and step 7 will want it.
The extraction protocol is a page, a file and four filters. The page is "Queensland's Shark Control Program Data and Information" on the Queensland Government Open Data Portal. The file is "Number caught by area, calendar year and species group" (scp_numbers-caught.xlsx), last modified 19 June 2026, downloaded in a browser because the portal answers a script with an empty file. The method records the date and what arrived, 23,118 rows on 26 September 2026, then the filters: Year 1996 to 2023, SpeciesGroup SHARK, Gear Net or Drum, and Area Gold Coast, Sunshine Coast North, Sunshine Coast South and Rainbow Beach. That leaves 4,299 rows. There is no address a reader can re-run, so this is where numbered screenshots of the filters earn their place.
The harder half was a column the file does not have. The dependent variable is the share of each gear's sharks that were white sharks, and no column holds it. So the method builds it in steps a reader can repeat. A helper column P sorts every row: =IF(J2="WHITE SHARK", "White shark", "Other shark"), where J is CommonName. Four SUMIFS formulas, on a tab holding only the filtered rows, then fill the 2 by 2 table: =SUMIFS(O:O, F:F, "Net", P:P, "White shark") adds up NumberCaught for white sharks in nets and gives 73. Each gear's percentage comes from that table. It is written into the method, not left for the results, because without it nobody could rebuild the number the whole study rests on.
Where the variable you analyse is not a column in the file, the steps that build it are part of your method and belong there, formula and all. A number a reader cannot rebuild makes the whole analysis unrepeatable, however precisely you describe the download.
Every number was made by somebody, somehow
The commonest weakness in a secondary investigation is treating the values as facts that simply existed, waiting to be found. They are the output of a process: someone chose a method, an instrument, a sampling frequency and a set of definitions, usually for a purpose that was not yours.
Read the methodology or metadata page: it is where most of your Criterion F (Evaluation) material comes from.
| How the number was made | What it means for your claim |
|---|---|
| Measured directly, on site | The strongest, and still has an instrument and a detection limit |
| Reported by an organisation or a country | Someone had a reason to report it, and sometimes a reason not to |
| Modelled or estimated from other variables | You are analysing a model's output, not an observation |
| Remotely sensed from satellite | Excellent coverage, but it classifies rather than sees: plantations can show up as forest |
| Imputed to fill a gap | The value exists because the gap did, not because anyone measured it |
Two more common traps. Values get revised: the figure you downloaded in September may not be the figure published in March, which is why the date accessed matters. And a definition can change mid-series, so a step in your graph may be a change in what was being counted rather than a change in the world.
Not a count of sharks in the sea. It is a count of the animals the gear caught, so it depends on how many nets and drumlines were in the water, and for how long, and the file records neither. That is why my dependent variable is a share of each gear's own catch rather than a rate: the file cannot say which gear catches more white sharks a day.
It is not a death either. Caught means found on the gear when it was checked: 65 of the 147 white sharks in my four areas were alive when found. The file's Fate column says whether each animal was released, died or was killed.
And it is recorded by the programme being judged on it. The Shark Control Program sets the gear and publishes what it catches. That is not a reason to distrust the numbers. It is a reason to say in the method what kind of number they are, and to notice when the programme changes its own gear: from 2024 new locations, daily servicing and new drumlines came in, which is why my study stops at 2023.
Find out what one value is, in the publisher's own words, before you average anything. Three properties matter most: what it is a measure of, who produced it, and what the units already have built into them. All three are usually one page away, and all three end up in step 7.
Fix the rule before you look
The decision that matters is which rows to keep. Nothing stops you making it after you have seen which choice gives the nicer graph. The file will let you, and it leaves no trace.
That is the integrity problem specific to this route. Choosing your countries, years or stations after you can already see the relationship is the same as throwing away the readings that disagreed with you. It is just harder to notice, and nobody can tell from the finished report.
So write the rule down first, apply it, and say in the method what it removed, with the count: five rows caught on Other gear left out, one white shark among them, and here is why.
I analysed the areas where the difference between nets and drumlines was clearest.
Unrepeatable, and the finding is a product of the choosing.
I kept the sharks caught in nets and on drumlines, 1996 to 2023, in the four areas that set both gears: Gold Coast, Sunshine Coast North, Sunshine Coast South and Rainbow Beach. From 2024 new locations, daily servicing and new drumlines came in, so 2023 is the last year before the gear changed. Five rows caught on Other gear were left out, one white shark among them.
A reader can re-run it and get your file back.
There is only one real decision in this method: which places to compare. Everything else follows. And it does not just change the answer, it reverses it. In the four areas that set both gears, white sharks were 2.2% of the nets' sharks and 3.8% of the drumlines'. Across the whole state, over the same years, they were 1.8% of the nets' sharks and 0.9% of the drumlines'.
So a student looking for a particular result could have picked either and given no reason for it. The reason used here was fixed before looking at the white sharks, and it is about the gear, not the answer: drumlines are also set far north, from Cairns to Bundaberg, where white sharks do not swim. A statewide comparison compares places as well as gear. Keeping to the four areas that set both takes place out of it.
Look closely at the statewide figure. It is not wrong: across Queensland, white sharks really were a larger share of the nets' catch. It answers a different question, about where each gear is set, and it would have seemed to support my hypothesis that nets catch the larger share. Both findings are true. Which one you report depends entirely on a choice you have to be able to defend.
Where your method has one decision in it, make that decision for a reason you can state before you see the result, and show the reader what the other choice would have given. Putting both answers side by side is the simplest proof that you did not pick the places that gave the answer you wanted.
Write your inclusion rule at the top of the spreadsheet, in a cell, before you make a single chart. It takes twenty seconds and it turns a good intention into a record. When step 7 asks what you would do differently, that cell holds the honest answer.
Your controls are arithmetic now
You cannot hold things constant physically: no same time of day, same observer, same instrument. What you can do is hold them constant mathematically, which is genuinely powerful and almost never used.
There are three moves, and naming which one you used is what turns a comparison into a controlled one.
Per person, per square kilometre, per unit of output. Comparing two countries' total emissions mostly compares their populations; comparing emissions per person compares what you actually meant to.
Only inland cities, only sites below 200 m, only stations on the same river. You lose sample size and you gain a comparison where the confounding variable cannot vary.
The same places early and late. Upstream against downstream on the same watercourse. Pairing removes everything that is constant within a pair, which is usually most of what worries you.
The first comparison was of raw counts, and it was misleading. In the four areas the nets caught 73 white sharks and the drumlines 74, which looks like no difference at all. But the nets caught 3,296 sharks in all and the drumlines 1,932, so a count of white sharks mostly compares how much of everything each gear caught.
Normalise. Each gear is measured against its own catch instead: not how many white sharks it caught, but what share of its sharks were white sharks. That gives 2.2% for the nets and 3.8% for the drumlines, two numbers measured the same way, and the two gears can sit in the same test.
Restrict did the rest. Only the four areas that set both gears, because drumlines are also set far north where white sharks do not swim, and only 1996 to 2023, because new locations, daily servicing and new drumlines came in from 2024. Each restriction removes a difference between the two gears that is not the gear itself.
Ask what else differs between the things you are comparing besides the thing you are studying, then divide it out rather than hoping it is small. Measuring each group against its own total is the simplest version: one division, and it works in almost every dataset.
Not every shark in the file is named to species. A name with a star is a group, not a species: HAMMERHEAD SHARK * is a hammerhead nobody identified further. There is also UNKNOWN SHARK, 18 animals in my four areas.
Filter them out and both gears' totals of other sharks shrink, so both white-shark shares rise. I kept them, because they are sharks and none is recorded as a white shark, and the method says so. What it cannot say is what they were: an unidentified shark could have been a white shark, and that goes to step 7 as a limitation.
How much data is enough, when you chose it
Here the question is different. Sample size is not a limit you ran into. It is a decision you made, so you have to justify it, not apologise for it, whether it is large or small.
The floor is set by the test you are going to run. Decide the test first, then count backwards to how many rows you need to keep.
| The test you plan to run | Rows it needs |
|---|---|
| A correlation | 10 pairs at least, 30 to be comfortable |
| A t-test | 10 or more per group |
| Chi-squared | 5 expected in every cell |
There is such a thing as too much. Run a correlation on ten thousand rows and almost anything comes out statistically significant, including relationships far too weak to mean anything. A p-value answers whether an effect exists, not whether it is big enough to care about.
So when your n (sample size) is in the thousands, quote the strength (r or R squared) alongside the significance. Talk about the size of the effect, not how small the p-value is. A moderator who sees p less than 0.001 on a correlation of 0.04 knows exactly what happened.
Bigger is not automatically better in the other direction either. Adding thirty more countries to reach a threshold, when twenty of them are not comparable with your original ten, gives you a bigger number but weakens your argument. Sample size that comes from widening your inclusion rule has to be defended against the rule, not against the total.
Twenty-three thousand rows sounds enormous. The analysis has four numbers in it: white sharks and other sharks, in nets and on drumlines. The unit being compared is an animal, not a row, so the 4,299 rows that pass the rule are added up through NumberCaught into 5,228 sharks: 3,296 in nets and 1,932 on drumlines.
Then the rare thing sets the limit, not the total. White sharks are 147 of the 5,228, and chi-squared needs 5 expected in every cell. Pooled across the four areas that is easily met: every expected count is above 50, the smallest about 54. Split by area it thins fast: Sunshine Coast South's nets caught 3 white sharks in the whole period. So the four areas are tested together. Each area's two shares can still be shown, and drumlines have the larger share in all four, but they are not tested one by one.
One exclusion, and it is worth its line: five rows caught on Other gear, one white shark among them, left out because the question compares two gears. Unidentified sharks stay in with the other sharks, and the method says so.
Count the things your test actually compares, not the rows in the file. Then find the rarest thing you are counting: the smallest cell in your table, not the total, decides whether the test is valid.
Ethics, when nobody gets wet
There is no risk assessment to write here and no consent form to hand anybody, which tempts students to skip the section entirely. Two things belong in it instead, and one of them is the most serious integrity question on this route.
The first is quick: say whose data it is, under what licence, and credit them. Open data still has to be credited, and a dataset is cited like any other source, plus the date you accessed it.
The second matters more: selective use of data is misconduct, not just carelessness. The protection is the rule you fixed before you looked, applied to everything and reported with what it removed.
Some open datasets describe individuals rather than places: health records, survey microdata, individual incomes. Those come with conditions, and the conditions still apply even if the file downloaded easily. If a dataset would let you identify a person, it does not belong in a school investigation.
It can explain what an indicator measures, what a licence permits, or how to write the spreadsheet formula that turns your raw column into the one you need. Ask it as many times as you like.
It cannot give you the numbers, or tell you a dataset exists. An invented figure looks exactly like a real one once it is in your spreadsheet. Every value comes out of a file you downloaded, and every dataset is one you have opened.
Ready for step 5?
Secondary data checklist0 of 18The first two are the criterion itself. Fixing your rules before you looked, and checking the file's shape, are the two that make a secondary investigation defensible rather than just convincing, and neither takes more than a few minutes.
You have a file and a rule for what is in it. Step 5 is where you turn its columns into something that answers your question, and where the good news arrives: tables, graphs and calculations all sit outside the word count. Decide your statistical test before you start processing, because it sets how much data you needed to keep.
The white shark investigation used on the secondary-data route of this guide is the author’s own analysis of a published dataset: the Queensland Shark Control Program’s record of every animal caught on its nets and drumlines, published by the Queensland Government under CC BY 4.0 and downloaded on 26 September 2026. The choice of the four areas, the analysis and the conclusions are the author’s, not the Queensland Government’s.
