An elderly woman in a pink shirt sorting through old black and white family photographs, evoking nostalgia. People by heritage list: a researched guide
Photo by SHVETS production on Pexels

Industry

Part of People by heritage: names, facts and dates

People by heritage list: a researched guide

People by heritage lists have four different owners, and the roll-up rule, not the category names, is what actually defines every published category.

A list of heritage categories is not a description of humanity. It is an instrument, built by a particular body, to do a particular job, and it stops making sense the moment it is lifted out of that job.

That single point explains most of what goes wrong when heritage categories are reused: a list designed to administer a program gets treated as a taxonomy of peoples, and a person's own answer gets replaced by a code somebody wrote for a different purpose.

What to take away

  • Every category list has an owner and a purpose. Both belong beside the data, and neither survives being copied into a different project.
  • The roll-up rule, which detailed answers get folded into which broad heading, carries more meaning than the headings themselves.
  • Lists from different countries are not comparable, because the countries are not asking the same question.

Four owners, four different lists

Who builds the list Why What the categories are optimized for
Statistical agencies To compile and publish population figures Stable tabulation, comparability with the agency's own past rounds
Administrative bodies To run a program with defined eligibility Clear boundaries that can be applied by a clerk
Communities themselves To define their own membership The community's own criteria, which no outsider sets
Commercial data products To segment an audience Whatever predicts a behavior the buyer cares about

These lists overlap in vocabulary and diverge in meaning. The same word can name a self-reported identity in one, an eligibility class in the second, a membership status in the third, and a marketing segment in the fourth. Joining two of them because the strings match is how a person ends up assigned to a category nobody ever placed them in.

The roll-up is the real definition

Most large collections accept a written-in answer and then code it to a category. Two documents govern that: the category list, and the coding manual that says which write-ins map to which category.

The manual is where the decisions are. It determines whether a regional identity is coded to a country, whether a religious identity is coded as an ethnicity, whether a compound answer is coded to one category or split, and what falls into the residual heading. It is revised between rounds, which means the same write-in can be coded differently in two collections that publish identical category names.

So when a category is quoted, quote the round and the manual. A published table showing a category rising over time may be showing a change in the coding rules, and the table itself cannot tell you which. How the instrument changed under this question over successive rounds is set out on the heritage guide.

The residual heading deserves its own attention. When "other" grows, it usually means the list has drifted away from how people describe themselves, and the interesting information is inside the bucket where nothing is tabulated.

Ascription and self-identification are different measurements

A category recorded by an observer and a category chosen by the person are two different variables that frequently share a column heading.

Observer-assigned categories measure how a person was classified by an institution, which is a real and important historical fact: it determined how they were treated, and it is documented in the institution's own records rather than in the person's account, which is why it is verified the way any other historical record is. Self-reported categories measure how a person describes themselves, which is a different real fact. Neither substitutes for the other, and a series that switches from one to the other partway through is not a series.

When you read a historical table, ask who supplied the answer before you read the number. When you build one, keep the two in separate columns even where the category names are identical.

Why cross-country comparison mostly fails

There is no international list because there is no international question.

Some states ask about ancestry, some about ethnicity, some about national identity, some about language, some about religion, some about several of these separately, and some ask none of them. National statistical offices publish their own wording and its limits, as with the guidance on ethnicity and the definitions behind migration statistics. A few prohibit collecting such data at all, for reasons rooted in their own history. The categories offered are drawn from each country's own social and political vocabulary, and they do not translate.

The practical consequence: a figure for a category in one country cannot be placed beside a figure with a similar name from another country. Any table that does this has aligned two different instruments by their labels. Where a cross-country comparison genuinely matters, compare the questions first and expect to conclude that they cannot be compared, which is a legitimate finding and a more useful one than a chart. The same reasoning applies to the country codes underneath such a table, described on the country data page.

Communities set their own membership

For many peoples, membership is determined by the community or the nation itself, under its own rules. An external list can record that a body recognized someone; it cannot confer or withhold membership, and an individual's claim of descent is not the same thing as membership.

This is the point at which a data project has to stop. A dataset can hold the fact that an authority recognized a person, with the authority named and the date recorded. It cannot hold an inference from a name, a document, a genetic estimate or a family story, and presenting one as membership misrepresents both the person and the community.

What a category list should carry

Alongside the categories themselves:

  • The owner and the purpose it was built for.
  • The version, and the date it took effect.
  • The coding manual, or a link to it.
  • The roll-up map from detailed answers to published headings.
  • Whether the value was self-reported or assigned, and by whom.
  • What the residual category contains.

Why there is no list of people by heritage here

Everything above is why. A list of people by heritage requires one category per person, taken from a list built by someone else for another purpose, applied to individuals who did not answer that question.

For living people it is simpler still: we do not assert anyone's ancestry, ethnicity, religion, nationality or birthplace. Where a person has publicly described their own background and it bears on the subject, we quote them, name the source and date it. And we publish no ranked tally of people by origin, because a count of that kind is driven by which archives were indexed rather than by the people, and because the framing invites a comparison between populations that we do not make. The wider version of that reasoning sits on the state cross-tab page.

Bottom line

Read the list as an instrument: who built it, to do what, in which version, with which coding rules. Do that and heritage categories become usable for the narrow, real jobs they can do. Skip it and you are publishing somebody else's administrative convenience as a fact about a person.

Common questions

Can two rounds of the same collection be compared?

Only after reading both questionnaires and both coding manuals. If the question wording, the answer options, the multiple-response rules or the coding changed, then any movement in the number is partly instrument and cannot be separated from the rest without the microdata.

Is a commercial ethnicity segmentation ever a usable source?

For the fact that a vendor assigned a label, yes, and it should be written that way. Such products commonly infer categories from names, addresses and purchase behavior. That is a model output about a household, not a statement by a person, and it should never be published as either.

More in Industry

Latest from Practice Desk