Sign in
Secondary data

What everything is called

Nobody warns you about this problem: one thing can have four names. You search for “raw data”, find nothing, and decide it does not exist. But the repository you were on calls it disaggregated, and the next one calls it microdata.

So this is a phrasebook, not a dictionary. Each entry gives the other names for the same thing, and says why it matters to an investigation. The steps do the teaching. This page is for when you are on a download page and a word stops you.

What a dataset is

The words for how a file is laid out. You need to know this before anything else.

Indicator
Also calledvariableseriesmetricparametermeasure

One measured quantity, tracked over time or across places.

Why it mattersMost repositories sort their data by indicator, and your research question has to name one. Search for a topic and you find articles. Search for an indicator and you find data.

Disaggregated
Also calledstation-levelmicrodataunit-recordraw recordsgranular

Broken down to the smallest unit the collector recorded, before any averaging.

Why it mattersThis is the form you want. A national yearly average has already removed the variation you wanted to analyse, and no processing can put it back.

Unit of analysis
Also calledobservationrecordone row

What a single row of your file represents.

Why it mattersIf you cannot finish the sentence "one row is...", you do not yet know what you are comparing. It is the first of step 4's four tests. It is also the fastest way to find out that a dataset is not the one you thought.

Granularity
Also calledresolutionlevel of detailspatial resolutiontemporal resolution

How finely the data is cut, in space or in time.

Why it mattersMonthly or yearly? One station or a whole country? If your question needs finer detail than the data has, it cannot be answered. This mismatch ends more investigations before they start than any other.

Coverage
Also calledextentscopecompleteness

Which places and which years are actually in the file.

Why it mattersIt is rarely what the title suggests. A dataset called "global" has gaps, and the gaps are usually not random. So coverage is a finding as well as a limit.

Time series
Also calledlongitudinalpanelcross-section

One thing measured again and again over time. A cross-section is many things measured once. A panel is many things measured again and again.

Why it mattersIt decides which question you can ask. Change over time needs a series. A comparison between places needs a cross-section. Comparing change between places needs a panel, the rarest and most powerful kind.

How the number was made

Every number was made by some process. These are the processes, and some are more trustworthy than others.

In situ
Also calledground-basedfield measurementobserved

Measured directly at the place, by an instrument or a person.

Why it mattersThe strongest kind. It still has a detection limit, a calibration and a person's judgement in it.

Modelled
Also calledestimatedderivedreanalysisgridded

Calculated from other measurements rather than observed.

Why it mattersYou are then analysing what a model produced. That is allowed, but it is a different claim from analysing an observation. Say which one you have.

Remotely sensed
Also calledsatellite-derivedearth observationclassified imagery

Worked out from what a sensor detected, usually on a satellite.

Why it mattersIt covers more ground than any other source, but it sorts the image into classes rather than seeing what is there. Every classified product is known to mix up certain classes, and its documentation usually says which.

Imputed
Also calledinterpolatedgap-filledestimated valuemodelled fill

A value filled in by a rule because the real one was missing.

Why it mattersIt is not a measurement. Treat it as one and you claim to know more than you do. Good datasets mark imputed values in a separate column, so first check whether yours does.

Provisional
Also calledpreliminaryunvalidatedsubject to revision

Published quickly and expected to change.

Why it mattersThe figure you download today may be different in six months. That is why your method has to record the date you accessed it.

Release
Also calledversionvintageeditionlast updated

Which issue of a dataset you have.

Why it mattersTwo people can download the same indicator a year apart and get different numbers. Recording the release makes your method exactly repeatable, not just roughly repeatable.

Detection limit
Also calledlimit of quantificationbelow reporting limitcensored

The smallest amount the method can distinguish from nothing.

Why it mattersValues at the limit are often recorded as zero, but that means "too small to detect", not "none". It changes what your mean means, and noticing it is a limitation worth marks.

What you do to it

The steps that turn a download into evidence. Naming the step you used is half of explaining why you used it.

Normalise
Also calledper capitaper unit areastandardisescale

Divide out a difference you are not interested in.

Why it mattersComparing two countries' total emissions mostly compares their populations. Per capita (per person) compares what you meant to compare. It is the most common and easiest form of control on this route.

Baseline
Also calledreference periodindex yearanomaly

The fixed point that change is measured against. An anomaly is a departure from it.

Why it mattersClimate data is often published as anomalies rather than values, and the publisher chooses the baseline. You cannot compare two datasets on different baselines until you put them on the same one.

Stock and flow
Also calledlevel and rateamount and change

A stock is how much there is. A flow is how fast it changes.

Why it mattersForest area is a stock, deforestation rate is a flow, and a question that mixes them up cannot be answered. It is also step 1's systems vocabulary, now in your spreadsheet.

Harmonise
Also calledreconcilealignmake comparable

Put two sources onto the same units, categories and dates before combining them.

Why it mattersAs soon as you use a second dataset, this becomes most of the work. Skip it and you get a join that looks fine but compares nothing.

Join
Also calledmergematchlink on a keylookup

Bring two datasets together on something they share, usually a date or a place.

Why it mattersIt is the best way to show independent processing (work you did on the data yourself) on this route, and nobody does it by accident. Rainfall joined to almost anything else is a good example to copy.

Inclusion criteria
Also calledselection rulefilterexclusion criteriasubset

The rule that decides which rows you keep.

Why it mattersOn this route it is both how you control variables and a test of your honesty. Set it before you look, report what it removed, and it is a method. Apply it after you can see the pattern, and you are choosing your result.

Balanced panel
Also calledmatched samplepaired comparisoncommon set

Only the units present in every period you compare.

Why it mattersIt answers the objection that your early and late groups are different places. It is easy to do. It turns the strongest criticism of a before-and-after comparison into a sentence saying you tested for it.

How it misleads you

The common errors, and their names. One limitation you can name is worth more than three you only hint at.

Proxy
Also calledindicator ofstands in forsurrogate

Something measurable that you are treating as a stand-in for something you cannot measure.

Why it mattersAlmost every environmental indicator is a proxy. Say what yours stands in for, and where the two differ. This is one of the most reliable ways into the top band of Criterion F (Evaluation).

Ecological fallacy
Also calledaggregation biasthe aggregate is not the individual

Assuming that what is true of a group is true of the things inside it.

Why it mattersA national average can hide severe local damage. A site mean can hide the storm peaks that cause the problem. It is the typical error when you work with somebody else's summaries.

Confounder
Also calledlurking variablethird variablecommon cause

Something that moves both of the things you are comparing.

Why it mattersYou did not control anything in the field, so this is the main threat to any relationship you find. Naming your two most likely confounders is worth more than a general warning that correlation is not causation.

Selection bias
Also calledsampling biasnot missing at randomsurvivorship

The data you have is not a fair sample of the thing you care about.

Why it mattersMonitoring stations go where somebody expected a problem, or where access was easy. Missing data is rarely missing at random. Where the gaps are is often a finding in itself.

Data dredging
Also calledp-hackingfishingcherry-picking

Trying enough combinations that something comes out significant by chance.

Why it mattersWith forty columns and a spreadsheet this is very easy to do, and nobody can see it in the finished report. To protect yourself, decide your question and your inclusion rule before you look at the data, and say that you did.

Temporal misalignment
Also calleddifferent reporting periodslagmismatched frequency

Two series that cover the same period but were recorded at different intervals or dates.

Why it mattersA census every five years against monitoring every month. A financial year against a calendar year. If your sources do not line up in time, saying so is part of your analysis, not an apology.

Finding it, and crediting it

What the parts of a repository are called, and what a citation of data has to carry.

Repository
Also calledportaldata hubcataloguedata store

A place that publishes datasets, usually with a search and a licence.

Why it mattersThe word to use when searching. "Air quality repository" finds the archive; "air quality data" finds articles about air quality data.

Metadata
Also calledcodebookdata dictionarydocumentationmethodology notereadme

The document explaining what each column is, how it was measured, and by whom.

Why it mattersOn this route, reading it is like calibrating an instrument, and it is where most of your evaluation material comes from. Fifteen minutes reading it saves you much more time later.

API
Also calledquery serviceendpointweb serviceREST

A way of asking a repository for exactly the rows you want, as a web address.

Why it mattersEasier than it sounds, and better than a download button. The query URL is itself a record of exactly how you got your data, so a reader can run it again instead of rebuilding it.

Licence
Also calledterms of useconditions of useCC BYODbLEtalabopen access

What you are permitted to do with the data, and what you must do in return.

Why it mattersOpen data almost always still has to be credited. Most licences require you to name the source, and naming it is the easy half of the ethics section on this route.

DOI
Also calleddigital object identifierpersistent identifier

A permanent address for a specific version of a dataset or paper.

Why it mattersIf your download offers one, use it. It is the strongest form of traceability there is, because it still works after the publisher reorganises their website.

Accessed date
Also calledretrieveddownloaded onas of

The day you took your copy.

Why it mattersData gets revised. Without this a reader cannot tell whether they are looking at the numbers you had, and it is the one part of a data citation people leave out.

Using these to search

Put the word for the kind of data next to the word for your subject. “Air quality repository” finds the archive; “air quality data” finds articles about air quality data. “Station-level” or “disaggregated” next to your topic is the fastest way to get past a page of national totals.

What everything is called · ESS IA guide | Revise