Reviews

Notability maps and data explained for 2027

Notability maps and data: notable is an inclusion rule rather than a property, and every step from a birthplace field to a dot on a map leaks.

A map of where notable people were born is not a map of where notable people were born. It is a map of where records were created, survived, were cataloged, were digitized, and could be matched to a coordinate. Every one of those five filters has a geography of its own, and the picture you end up looking at is mostly their combined shape.

This page is about what these datasets actually contain, so that the small number of things they genuinely support can be separated from the much larger number of things they are used for.

What to take away

  • Descriptions of the reference work's own coverage, including where it is thin.
  • Any claim about the distribution of talent, ability, or achievement across places or populations.
  • The source dataset, its version, and the extraction date.

"Notable" is an inclusion rule, not a property

Nothing in a person makes them notable. Notability is a decision made by a reference work about whether to have an entry, and every reference work uses a different rule.

The most common modern rule is derivative: a person qualifies if independent sources have already covered them at length. That is a defensible rule for an encyclopedia, because it outsources judgment to the wider record rather than to the editor. It also means the resulting dataset inherits, whole, the coverage biases of the press, the academy, and the publishing industry, including which languages they worked in, which countries they staffed, and which periods they cared about.

Other rules are positional: holding a particular office, playing at a particular level, receiving a particular award, appearing in a particular prior compilation. These are cleaner to apply and openly circular, since the qualifying institutions are themselves unevenly distributed.

Consequences worth stating plainly:

  • The threshold moves. It differs by subject area within a single reference work, and it shifts over time as coverage grows and as inclusion policy is rewritten.
  • The rule selects for documentation, not achievement. Work absorbed into an institution's output, work in unpaid or subordinate roles, and work in fields with no habit of crediting individuals all fail a coverage test regardless of what was accomplished.
  • Thin-press periods and places drop out. Where local newspapers did not exist, did not survive, or were never indexed, nobody clears a coverage bar.
  • Language decides visibility. Sources in languages the compiling community does not read are functionally invisible to the rule, whatever they contain.
  • Deletion is part of the process. Entries are removed as well as added, and the removals follow the same biases as the additions.

None of this makes the datasets useless. It makes them datasets about a reference work. Read that way, they are excellent. Read as a census of accomplished people, they are not a census of anything.

What happens between a birthplace field and a dot

The pipeline is short and every step leaks.

1. The source field is text. It holds whatever the cataloguer wrote: a village, a district, a country, a historical polity, a hospital, or a string like "near" somewhere else. It may hold a place that no longer exists under that name, or one of several distinct places sharing a name in the same state, as any gazetteer entry in the Geographic Names Information System will show.

2. A geocoder turns text into coordinates. Where the name is ambiguous, and common place names are ambiguous across continents, the geocoder picks one, usually by population or by a default region. Where the name is not in the gazetteer at all, most pipelines fall back to the centroid of the containing region or country. That fallback is the single most consequential step in the whole process: it silently invents a precise-looking point in the middle of a large area, and if a few thousand records fall back the same way, the map grows a dense cluster in a place where nobody was born. Centroid artifacts are common, they look like findings, and they are almost never annotated.

3. The point is assigned to modern administrative units. Historical boundaries are not the boundaries the map uses. A person born under one jurisdiction is counted under whichever present-day unit contains the coordinate, which is a defensible convention and a different question from the one readers think is being answered.

4. Points are aggregated into areas. Counts per region, per state, per country.

5. The counts are normalized and colored. Both choices change the map, and neither is neutral.

The custody rules behind that text field are set out on the US state pillar and the city pillar.

Place of the record is not place of the event

Registration happens where the registry is. Historically that meant the parish, the district town, or the county seat, so birthplaces cluster on administrative centers rather than on where people lived. In the modern period the clustering is even stronger and has a different cause: births increasingly happen in hospitals, hospitals are in larger towns, and a birth certificate records the hospital's location.

The result is a real, measurable pull of recorded births toward towns with maternity facilities, and away from the rural areas whose residents used them. On a dot map that reads as urban concentration. It is partly a fact about where obstetric services are sited.

The same logic applies to deaths, which concentrate at hospitals and care facilities, and to any "worked in" field, which usually records an employer's registered address.

Five ways an area map misleads without anyone lying

Unit choice changes the answer. Aggregate the same points into different sets of areas and the pattern changes, sometimes reversing. This is a well-known property of area data with a name, the modifiable areal unit problem, not a flaw in a particular map, and it means any single choice of units is one view among many. A map published without its unit rationale cannot be evaluated.

Denominators are estimates. Per-capita rates need a population figure, and the choice of which year to use is a choice about what the map means: population at the subjects' birth, at their death, or now. For historical periods, the population figure is itself a reconstruction with its own uncertainty, which never appears on the map. Small denominators then produce extreme rates from tiny counts, so sparsely populated areas dominate the top and bottom of any per-capita ranking almost automatically.

Area-level patterns say nothing about individuals. A region with a high rate tells you nothing about any person from it, and a difference between regions is not a difference between the people in them. Treating an aggregate as a property of individuals is the standard error in this format.

Blank is ambiguous. An empty region can mean no qualifying people, no digitized records, no matching gazetteer entries, or a source that was never ingested. Those are four completely different statements, and a choropleth renders all of them as the same pale color. Missing data must be styled differently from zero, and almost never is.

Classification and projection carry the story. Quantile, equal-interval, and natural-break schemes drawn from identical data produce maps that look like different findings. Meanwhile large, sparsely populated areas take up visual space out of all proportion to the counts they contain, so the eye reads geography as magnitude.

Time series of a dataset are not time series of the world

Plot records per year and you will see structure. Before interpreting it, account for:

  • Ingest events. A collection is added, a partner institution's records are imported, a language edition is merged. The series jumps on the day of the import, not on anything that happened in the world.
  • Rule changes. An inclusion threshold is revised and a whole class of entries appears or disappears.
  • Digitization campaigns. A funded program to scan one archive changes the shape of a whole region and period.
  • Documentation lag. Recent years look sparse because coverage accumulates for decades after the fact.
  • Boundary changes. A region redrawn between two time points cannot be compared with itself.

Any of these produces a step change that reads as a historical trend. The fix is to keep a change log for the dataset and to plot it alongside the series, so that the discontinuities are visible rather than interpretable.

What these datasets can support

  • Descriptions of the reference work's own coverage, including where it is thin. This is genuinely valuable and under-published.
  • Relative comparisons within one source, one period, and one collection method, stated as such.
  • Detecting the effects of digitization and cataloging programs.
  • Finding candidates for research, which is what authority-file and catalog data was built for.
  • Documenting gaps as evidence of who the record leaves out.

What they cannot support

  • Any claim about the distribution of talent, ability, or achievement across places or populations.
  • Comparisons between regions with different record-keeping histories, which is nearly all pairs of regions.
  • Per-capita claims for historical periods, where the denominator is a model with unstated uncertainty.
  • Inference from an area to an individual.
  • Ranked tallies of people by origin. These are refused here, and the reasoning is set out with the country field, the heritage variables, and the attention metrics. The short version: an unreliable grouping variable crossed with a coverage-driven count produces an ordering determined by archives, and the framing would not be one we would publish even if the data were sound.

If you publish a map, publish these with it

  • The source dataset, its version, and the extraction date.
  • The inclusion rule, in the reference work's own words.
  • The geocoding method, the gazetteer used, and the fallback behavior: including how many records fell back to a centroid.
  • The date basis: birth, death, or activity, and which historical or modern boundary set was applied.
  • The aggregation units and why those units.
  • The denominator, its source, and its year.
  • The classification scheme and the class breaks.
  • Missing data rendered distinctly from zero, with a legend entry that says so.
  • One sentence naming the inference a reader should not draw from the map.

Bottom line

These datasets describe catalogs. That is a real subject, and the honest version of this work is cartography of the record itself: where it is dense, where it is empty, and why. The moment a map of records is captioned as a map of people, it stops being evidence and starts being an illustration of something nobody measured.

Common questions

Is "notable" ever a property of a person?

No. It is an inclusion rule belonging to whichever work you are counting from, and different works apply different rules. Say which work, and the sentence becomes checkable.

What is the minimum a published map needs beside it?

The numerator, the denominator, the period of each, the boundary vintage, and the underlying table. Without those a reader cannot check a single cell.

Why not just use a dot map to avoid area problems?

Dot maps trade one distortion for another. Points placed at an area's centroid look like locations and are not, which readers rarely discount.

Can a map of people ever be honest?

As a map of a named record set, yes, with a coverage layer. As a map of where notable people came from, no, because the pattern is dominated by which archives survived.

More in Reviews

Reviews

People by country: a sourced reference guide for 2027

People by country: six different questions share that label, borders moved under the people in the records, and no single field can hold the answer.