What counts as a dataset
All four requirements, not three. Each one takes a couple of minutes to check. Together they save you from the one failure you cannot recover from: finding out at the end that your dataset could never have answered your question.
The four requirements
- 1
It has to be numbers you can change, not a picture of numbers.
A chart on a page, a figure in a paper, a PDF of a table. None of those are data yet. You need values you can put in a spreadsheet and change. If the only way to get them is to read them off an axis, keep looking.
- 2
You have to be able to say what one row is, in a sentence.
One station in one year. One country in one month. If you cannot finish that sentence, you do not yet know what you would be comparing, and neither will your reader.
- 3
It has to cover your years and your places. Both, not one.
A strategy introduced in 2015 needs data either side of 2015. A question about a river needs stations on that river. Check before you write the question, not after.
- 4
You have to be able to have it on screen in ten minutes. Today, not in two weeks.
Some of the best environmental archives need an account and a rights request that takes two weeks to approve. That is fine for a researcher and far too slow for you. If you cannot have it on your screen now, it is not your dataset.
The fourth is the one people argue with, and the one to follow most strictly. An investigation you cannot start is worth less than a smaller one you can finish, and the good archive will still be there next year.
Things that look like data and are not
The shortest way to show why the four requirements exist: each of these is a real published resource a student would reasonably click on, and none can carry an investigation.
Qualité de l'eau des plages genevoises du lac
It is a status map, not a series. Thirty-five rows, one per beach, each carrying a word such as “Bonne” and the date it was last refreshed. There is no number to process and no history to compare, and you cannot tell that from its title.
The CIPEL Limnothèque, fifty years of Lake Geneva monitoring
The obvious place to look, and the reason this entry exists. The commission that monitors the lake has measured phosphorus, oxygen, nitrate, chlorophyll and water transparency at the deep station since the 1950s, and it publishes all of it in annual scientific reports. Reports, not files: the numbers are inside a PDF, in tables, next to the graphs drawn from them. Transcribing a table you can see is legitimate work and slow work, so choose to do it on purpose rather than discovering it late in the day. The rivers feeding the lake are a different matter: those are six of the counted files above.
Climate-Data.org, the climate of any city in the world
It has no years in it. Every city page is the same twelve rows, one per calendar month, and each value is a single average of 1991 to 2021, modelled from Copernicus grid data rather than measured at that place. Nothing can be compared with anything, so no strategy can be tested against it. It looks more like data than the beaches map does, which is exactly why it is here: London's “rainy days” reads 8 in all twelve months.
The 2026 World Population Data Sheet
The most credible-looking thing a search will hand you, which is why it is here. Twenty-seven indicators for two hundred countries, every one with its definition beside it, from an institute that has published them annually since 1962. It fails on years. The sheet is a single cross-section, mid-2026, with no history inside it, and only 2024 and 2026 are in the explorer. PRB also rules out the comparison you would try next: their own Methods page says data sheets from different years “should not be used as a time series”, because a value that moves between editions usually means they revised an estimate rather than that the world changed. The spreadsheet is behind a form wanting your name, your organisation, your job title and what you intend to do with the data, after which the links arrive by email, so it is not a file you can have open in ten minutes either. Read the poster, then take the numbers from World Population Prospects in the grid above, which has 1950 to 2100 in it.
The IUCN Red List of Threatened Species
The authority on which species are threatened, and the source of every “endangered” in your background reading. It is a verdict, not a series. Each species carries one category from its latest assessment, reassessed every ten years or so, so there is nothing to track through time and no count of animals behind the word. Downloading the spatial or bulk data needs an account and a description of what you intend to do with it. Use it to choose and justify your species, then take the numbers from GBIF in the grid above, and if your question is about extinction risk over time, the Red List Index is the series, published as SDG indicator 15.5.1.
The Ocean Cleanup's river plastic emissions
A real download, openly licensed, with 31,819 river mouths each carrying a figure in tonnes of plastic a year. None of those figures was measured. They are the outputs of one model (Meijer and others, 2021), run once for a mid-point scenario, so there is a single snapshot with no years to compare and nothing a strategy could have changed. It also arrives as a GIS shapefile, not a table. Comparing rivers inside it is comparing the model's assumptions with themselves. It is good background for why rivers matter to ocean plastic; for measured concentrations, use the NOAA marine microplastics file in the grid above.
