What everything is called
Nobody warns you about this problem: one thing can have four names. You search for “raw data”, find nothing, and decide it does not exist. But the repository you were on calls it disaggregated, and the next one calls it microdata.
So this is a phrasebook, not a dictionary. Each entry gives the other names for the same thing, and says why it matters to an investigation. The steps do the teaching. This page is for when you are on a download page and a word stops you.
What a dataset is
The words for how a file is laid out. You need to know this before anything else.
One measured quantity, tracked over time or across places.
Why it mattersMost repositories sort their data by indicator, and your research question has to name one. Search for a topic and you find articles. Search for an indicator and you find data.
Broken down to the smallest unit the collector recorded, before any averaging.
Why it mattersThis is the form you want. A national yearly average has already removed the variation you wanted to analyse, and no processing can put it back.
What a single row of your file represents.
Why it mattersIf you cannot finish the sentence "one row is...", you do not yet know what you are comparing. It is the first of step 4's four tests. It is also the fastest way to find out that a dataset is not the one you thought.
How finely the data is cut, in space or in time.
Why it mattersMonthly or yearly? One station or a whole country? If your question needs finer detail than the data has, it cannot be answered. This mismatch ends more investigations before they start than any other.
Which places and which years are actually in the file.
Why it mattersIt is rarely what the title suggests. A dataset called "global" has gaps, and the gaps are usually not random. So coverage is a finding as well as a limit.
One thing measured again and again over time. A cross-section is many things measured once. A panel is many things measured again and again.
Why it mattersIt decides which question you can ask. Change over time needs a series. A comparison between places needs a cross-section. Comparing change between places needs a panel, the rarest and most powerful kind.
How the number was made
Every number was made by some process. These are the processes, and some are more trustworthy than others.
Measured directly at the place, by an instrument or a person.
Why it mattersThe strongest kind. It still has a detection limit, a calibration and a person's judgement in it.
Calculated from other measurements rather than observed.
Why it mattersYou are then analysing what a model produced. That is allowed, but it is a different claim from analysing an observation. Say which one you have.
Worked out from what a sensor detected, usually on a satellite.
Why it mattersIt covers more ground than any other source, but it sorts the image into classes rather than seeing what is there. Every classified product is known to mix up certain classes, and its documentation usually says which.
A value filled in by a rule because the real one was missing.
Why it mattersIt is not a measurement. Treat it as one and you claim to know more than you do. Good datasets mark imputed values in a separate column, so first check whether yours does.
Published quickly and expected to change.
Why it mattersThe figure you download today may be different in six months. That is why your method has to record the date you accessed it.
Which issue of a dataset you have.
Why it mattersTwo people can download the same indicator a year apart and get different numbers. Recording the release makes your method exactly repeatable, not just roughly repeatable.
The smallest amount the method can distinguish from nothing.
Why it mattersValues at the limit are often recorded as zero, but that means "too small to detect", not "none". It changes what your mean means, and noticing it is a limitation worth marks.
What you do to it
The steps that turn a download into evidence. Naming the step you used is half of explaining why you used it.
Divide out a difference you are not interested in.
Why it mattersComparing two countries' total emissions mostly compares their populations. Per capita (per person) compares what you meant to compare. It is the most common and easiest form of control on this route.
The fixed point that change is measured against. An anomaly is a departure from it.
Why it mattersClimate data is often published as anomalies rather than values, and the publisher chooses the baseline. You cannot compare two datasets on different baselines until you put them on the same one.
A stock is how much there is. A flow is how fast it changes.
Why it mattersForest area is a stock, deforestation rate is a flow, and a question that mixes them up cannot be answered. It is also step 1's systems vocabulary, now in your spreadsheet.
Put two sources onto the same units, categories and dates before combining them.
Why it mattersAs soon as you use a second dataset, this becomes most of the work. Skip it and you get a join that looks fine but compares nothing.
Bring two datasets together on something they share, usually a date or a place.
Why it mattersIt is the best way to show independent processing (work you did on the data yourself) on this route, and nobody does it by accident. Rainfall joined to almost anything else is a good example to copy.
The rule that decides which rows you keep.
Why it mattersOn this route it is both how you control variables and a test of your honesty. Set it before you look, report what it removed, and it is a method. Apply it after you can see the pattern, and you are choosing your result.
Only the units present in every period you compare.
Why it mattersIt answers the objection that your early and late groups are different places. It is easy to do. It turns the strongest criticism of a before-and-after comparison into a sentence saying you tested for it.
How it misleads you
The common errors, and their names. One limitation you can name is worth more than three you only hint at.
Something measurable that you are treating as a stand-in for something you cannot measure.
Why it mattersAlmost every environmental indicator is a proxy. Say what yours stands in for, and where the two differ. This is one of the most reliable ways into the top band of Criterion F (Evaluation).
Assuming that what is true of a group is true of the things inside it.
Why it mattersA national average can hide severe local damage. A site mean can hide the storm peaks that cause the problem. It is the typical error when you work with somebody else's summaries.
Something that moves both of the things you are comparing.
Why it mattersYou did not control anything in the field, so this is the main threat to any relationship you find. Naming your two most likely confounders is worth more than a general warning that correlation is not causation.
The data you have is not a fair sample of the thing you care about.
Why it mattersMonitoring stations go where somebody expected a problem, or where access was easy. Missing data is rarely missing at random. Where the gaps are is often a finding in itself.
Trying enough combinations that something comes out significant by chance.
Why it mattersWith forty columns and a spreadsheet this is very easy to do, and nobody can see it in the finished report. To protect yourself, decide your question and your inclusion rule before you look at the data, and say that you did.
Two series that cover the same period but were recorded at different intervals or dates.
Why it mattersA census every five years against monitoring every month. A financial year against a calendar year. If your sources do not line up in time, saying so is part of your analysis, not an apology.
Finding it, and crediting it
What the parts of a repository are called, and what a citation of data has to carry.
A place that publishes datasets, usually with a search and a licence.
Why it mattersThe word to use when searching. "Air quality repository" finds the archive; "air quality data" finds articles about air quality data.
The document explaining what each column is, how it was measured, and by whom.
Why it mattersOn this route, reading it is like calibrating an instrument, and it is where most of your evaluation material comes from. Fifteen minutes reading it saves you much more time later.
A way of asking a repository for exactly the rows you want, as a web address.
Why it mattersEasier than it sounds, and better than a download button. The query URL is itself a record of exactly how you got your data, so a reader can run it again instead of rebuilding it.
What you are permitted to do with the data, and what you must do in return.
Why it mattersOpen data almost always still has to be credited. Most licences require you to name the source, and naming it is the easy half of the ethics section on this route.
A permanent address for a specific version of a dataset or paper.
Why it mattersIf your download offers one, use it. It is the strongest form of traceability there is, because it still works after the publisher reorganises their website.
The day you took your copy.
Why it mattersData gets revised. Without this a reader cannot tell whether they are looking at the numbers you had, and it is the one part of a data citation people leave out.
Put the word for the kind of data next to the word for your subject. “Air quality repository” finds the archive; “air quality data” finds articles about air quality data. “Station-level” or “disaggregated” next to your topic is the fastest way to get past a page of national totals.
