Identidade falsa > Artigos > Krzysztof Peaked in 1971. Kacper Peaked in 2001. How Birth Cohorts Separate Migrants From Their Children

Este artigo ainda não foi traduzido para Português — você está lendo o original em English. Também disponível em:Deutsch, English, Українська

Krzysztof Peaked in 1971. Kacper Peaked in 2001. How Birth Cohorts Separate Migrants From Their Children

In Norway's name data, Krzysztof rises from zero to 1,148 people in the 1971 birth cohort — and then collapses to 14 by the 2001 cohort. Kacper, also Polish, does the exact opposite: 0 in every cohort through 1971, then 14, then 107, then 215 in 2001.

Two Polish names in the same country, moving in opposite directions across the same fifty years. Neither pattern makes sense as fashion. Together they make sense as one thing: Krzysztof arrived in Norway as an adult; Kacper was born there.

This is a mechanism worth naming, because it turns a single frequency table into a two-generation story. Give a name dataset a time axis and it stops telling you what parents liked and starts telling you who moved, roughly when, and whether their children kept the naming tradition or dropped it. Below we run that analysis on four national datasets, report what the numbers say, and — at some length — explain the two places where we nearly published a false result.

The two-layer signature

Take the Polish given names in Norway's cohort data. Split them by whether they peak early or late:

Name1941195119611971198119912001Peak
Andrzej192996046102602001971
Krzysztof01565911148850134141971
Grzegorz0532898056356301971
Agnieszka007769372812941981
Wojciech04122140433710351971
Jakub000863472902701981
Kacper0000141072152001
Wiktoria00000742442001
Oliwia00000401652001

The top block and the bottom block are different populations.

The top block is the migrants themselves. Krzysztof, Grzegorz, Andrzej and Agnieszka were fashionable in Poland in the 1960s and 1970s. People born in Poland in those decades, carrying those names, moved to Norway as adults — overwhelmingly after Poland joined the EU in 2004, which opened Norwegian labour markets to Polish workers. Their birth cohorts are the 1960s and 1970s because that is when they were born; Norway had nothing to do with it. By the 2001 cohort these names are near zero, because Krzysztof stopped being fashionable in Poland too.

The bottom block is their children. Kacper, Wiktoria and Oliwia are names that became popular in Poland in the 1990s and 2000s. In Norwegian data they appear only in the youngest cohorts — these are children born in Norway to Polish parents, given contemporary Polish names.

The shape difference is the whole finding. A first-generation layer shows up as a bulge at the migrants' birth decades that then decays, because migration to a country is a one-time event and the cohort ages out. A second-generation layer shows up as a monotonic rise into the newest cohorts, because births keep happening.

You can read the second generation's choices off the same table. Polish parents in Norway did not switch to Norwegian names. They picked Kacper and Wiktoria — current Polish fashion, not Norwegian fashion, and not the parents' own generation's names either.

Contrast Vietnamese names in the same dataset:

Name1941195119611971198119912001
Minh0215149223420
Ngoc095874601813
Phuong0433705170

First-generation bulge peaking at the 1961–1971 cohorts — Vietnamese refugees who arrived in Norway from the late 1970s, born in the 1950s and 60s — and then no second layer at all. Phuong reaches zero. Whatever Vietnamese families in Norway named their Norwegian-born children, it was not Phuong.

Same country, same decades, two migration streams, two completely different transmission outcomes. That comparison is invisible in any snapshot of current name frequencies and obvious the moment you add the time axis.

The Arabic-tradition layer across four countries

The largest measurable naming layer in Western European data is names from the Arabic and Islamic onomastic tradition. We tracked a strict, unambiguous core — the Muhammad spelling variants, Ahmed/Ahmad, Fatima, Hamza, Ayoub, Youssef, Bilal, Ibrahim, Mustafa, Khadija, Aisha, Maryam — as a share of each birth cohort:

Country1941195119611971198119912001Growth
France0.018%0.073%0.192%0.310%0.459%0.468%0.719%×39
Spain0.109%0.292%0.454%0.686%1.032%1.347%1.070%×10
Norway0.187%0.226%0.379%0.518%0.853%0.876%0.933%×5
USA0.001%0.001%0.005%0.041%0.057%0.102%0.157%×174

In raw births, France goes from 815 in the 1941 cohort to 44,130 in the 2001 cohort. Spain from 3,555 to 48,800. The United States from 156 to 48,397.

Each country's curve encodes its own history rather than a common one:

  • France shows steady growth from a near-zero 1940s base, consistent with post-war Maghrebi labour migration beginning in the 1950s–60s and family reunification from the 1970s.
  • Spain starts from a base ten times France's — 0.109% in the 1941 cohort, over 3,500 births. Spain had Moroccan territories and a substantial Moroccan-descended population long before it became a migration destination in the 1990s. It is also the only one of the four to have peaked and declined: 1.347% at the 1991 cohort, 1.070% at 2001.
  • Norway has the flattest growth, ×5, from an already non-trivial base — but see the warning below, because Norway's baseline is the number we got wrong first.
  • The United States shows the steepest multiple, ×174, precisely because its 1940s base was so close to zero (156 births nationally).

Individual names carry the same story more legibly than aggregates:

NameCountry1941195119611971198119912001
MohamedFrance2502135613510560158201232014500
RayanFrance000595356513480
AyoubFrance00105542021004240
OmarUSA283772241210272143422455022973
HamzaSpain003112088847763833
SalmaSpain00461062956695196
IlhanFrance00030501051400

Rayan in France is the second-generation signature in its purest form: zero, zero, zero, 5, 95, 3,565, 13,480. That is not migration — nobody migrates in that pattern. That is French-born children being named, and the name they were given was not the one their grandparents carried. Mohamed, meanwhile, has the first-generation shape: it peaks at the 1981 cohort and then falls, from 15,820 to 14,500 even as the overall layer grows. The tradition is expanding while its single most traditional name contracts — a generational shift happening inside a growing population.

Names that vanished entirely

The same time axis catches the opposite process. In Norway's data, these names have a positive count in the 1941 cohort and exactly zero in 2001:

Name1941195119611971198119912001
Britt132525111904767293810
Toril10871565999418143460
Bodil98017491095513167560
Rigmor73010504941491940
Dagfinn50479648322995140
Arnfinn43997952921689290

Britt accounted for 0.90% of Norwegian women in the 1941 cohort. Sixty years later, not one. These are complete extinctions inside a living register, and they are far more dramatic than anything on the migration side. The largest incoming layer we measured in Norway moved by less than a percentage point; Anne alone fell from 5.64% to 0.65%, and Jan from 4.34% to 0.43%.

That proportion deserves emphasis, because coverage of naming change tends to invert it. In every dataset we examined, native fashion churn is an order of magnitude larger than migration. The names replacing Jan and Anne in Norway are Sander, Tobias, Emma and Nora — not migration at all, just the ordinary, relentless turnover documented in Your Name Is a Birth Certificate.

Two mistakes we made

Both were caught before publication. Both would have produced a plausible-looking, entirely wrong article.

Mistake 1: Laila is a Norwegian name

Our first pass at Norway's Arabic-tradition layer returned a 1941 baseline of 0.709% — implausibly high for wartime Norway. The cause was a single entry:

Name1941195119611971198119912001
Laila144222412593194950720399

Laila is an Arabic name. It is also, and in Norway overwhelmingly, a Nordic one — it entered Norwegian and Sámi usage in the nineteenth century and was a top-ranking Norwegian girls' name through the 1950s and 60s. Its curve is the classic native fashion shape: peak in the middle, decline to the present. It has nothing to do with migration.

Laila alone contributed 0.48 percentage points to the false baseline. Removing it and other host-established names dropped Norway's 1941 figure from 0.709% to 0.187% — and dropped the measured growth from ×2.4 to ×5, because the denominator had been inflated by a name that was already there.

The general failure: etymology is not usage. A name's origin says nothing about which population currently uses it. Every cluster in this article was cleaned by removing names long established in the host language, and the removals were substantial:

Removed from clusterWhy
Jasmine (S. Asian)Persian etymology, but an English fashion name since the 1970s — 103,726 US births in the 1991 cohort alone, which would have swamped every genuine signal
Laila (Arabic, Norway)Established Nordic name, peaked 1961
Amelia (Polish)English name; UK top-10 for reasons unrelated to Polish migration
Natalia, Igor (Polish, Spain)Established Spanish names
Magdalena (Polish, Spain)Spanish name — with it included, Spain's "Polish" layer appeared to be declining since 1941
Mai (Vietnamese)Also a Nordic and French name
Anita (S. Asian)Established in Norwegian and Spanish

Every one of these produced a wrong number before removal, and several produced wrong numbers with the right sign, which is the dangerous kind.

Mistake 2: the data doesn't go back that far — and it isn't British

Our first run reported that in our en_GB cohort file the Arabic-tradition layer rose from 0.000% in the 1941 cohort to 0.949% in 2001 — an infinite multiple, and by a wide margin the most quotable number we generated.

It is an artefact. Counting how many distinct names have a non-zero entry in each decade of that file:

DecadeDistinct male namesSmallest recorded count
194111632
195112138
196112244
19714,2591
19816,1021
19917,5941
20019,3921

Before 1971 the source publishes roughly a top-100 list. From 1971 it publishes everything down to a single birth. The zeros in the early decades are not measurements of absence — they are the absence of measurement. Any name outside the top 100 reads as zero, and every migration-associated name is outside the top 100 in 1941 by construction.

There is a second error in that paragraph, and it took longer to surface: the source was never British. The file was compiled from National Records of Scotland — a registrar covering roughly 8% of UK births — and had been labelled en_GB end to end. The counts and the 1971 break are correct; the country on the label was not. Decades 1941–1991 in that file are Scottish still, because no UK-wide source publishes them; the 2001 decade has since been rebuilt from all three UK registrars and now carries 14,094 distinct male names instead of 9,392. How that mislabelling was caught is a separate story: Is Your National Dataset Actually Regional?

The check that catches this is trivial and should be automatic: count the distinct entries per period and look at the smallest non-zero value. If either jumps, your series has a break in it. Our four reported countries all pass:

CountryNames per decade (1941 → 2001)Reporting thresholdVerdict
Norway465 → 725 (male)4, constantusable
USA5,934 → 23,079 (male)5, constantusable
France1,796 → 10,039 (male)5, constantusable
Spain5,000 every decade6–32, driftingusable with care
Scotland (labelled en_GB)116 → 9,39232 → 1excluded pre-1971
New Zealand239 → 56510, constanttoo shallow

The Scottish file and New Zealand were dropped from every headline figure in this article. The temptation not to drop them was real: the Scottish file produced our largest, cleanest-looking effect. It was entirely manufactured by a change in publication policy around 1971.

A note on what these cohorts count

One further subtlety, and it changes the interpretation.

Norway's 1941 cohort totals 297,799 people. Norway actually recorded roughly 620,000 births in 1941–1950. The cohort is about half the births — because this is a register of current residents by birth year, not a register of births. People who died are not in it; people who moved to Norway later are, filed under the year they were born abroad.

That is exactly why the Krzysztof analysis works. A man born in Poland in 1971 who moved to Norway in 2006 appears in Norway's 1971 cohort. In a pure birth register he would appear nowhere at all, and the first-generation layer would be invisible.

The four datasets are not the same kind of object:

  • Norway — resident register. Shows first and second generation. This is why the two-layer signature is visible there and hard to see elsewhere.
  • United States — social-security records tied to births. Shows people named in the US; a first-generation migrant who arrived as an adult is largely absent.
  • France — births registered in France, excluding people born abroad. First generation excluded by construction. Every French figure in this article is therefore an undercount of the tradition's real presence, in a known direction.
  • Spain — births by decade, truncated to the top 5,000 names per decade.

France's exclusion is worth restating because it inverts the usual worry: the ×39 growth in France is a lower bound. The real footprint is larger than the file can show.

Conclusion

A frequency table tells you what names exist. A frequency table with a time axis tells you who arrived, roughly when, and what they named their children. Krzysztof and Kacper are the same migration seen twice, twenty years apart, and neither name means anything without the other.

The method is not hard, but it has sharp edges, and both of ours drew blood. A single Nordic name inflated Norway's baseline nearly fourfold. A change in Scottish publication practice around 1971 manufactured an infinite growth rate out of nothing. Both errors would have survived any amount of proofreading, because the outputs looked entirely reasonable.

If you take two habits from this article: check what your zeros mean before believing them, and never assign a name to a population by its etymology.


Methodology and sources

What was computed. Cohort files give birth counts per name across seven decade buckets (1941, 1951, 1961, 1971, 1981, 1991, 2001). The analysis was run against share 1.1.7, which was current at the time and carried cohort files for six locales; the shipped corpus is now share 1.1.8 and carries them for ten, so the selection and exclusions below describe the files as they stood at that run. Male and female files were merged per locale. Cluster shares are cluster births divided by all births recorded in that decade for that locale. Analysis run 2026-07-18.

Locales used and excluded.

LocaleCohort depth (1941→2001, male)ThresholdUsed
no_NO465 → 7254, constantYes
en_US5,934 → 23,0795, constantYes
fr_FR1,796 → 10,0395, constantYes
es_ES5,000 flat (top-5000 cap)6→32, driftingYes, with caveat
en_GB116 → 9,39232→1, breaks at 1971No — and the file was Scottish, not UK-wide
en_NZ239 → 56510, constant but shallowNo

Cluster definitions. Clusters group names by linguistic and onomastic tradition — the language and naming stock a name comes from — and by nothing else. They are not and cannot be proxies for the ethnicity, nationality, religion or ancestry of the people carrying them, and no such inference is drawn anywhere above. A name is a naming choice; it identifies a tradition a family drew on, not the family. Cluster membership was assigned by etymology and then filtered by host-language usage, because etymology alone produces false positives (see Laila, Jasmine, Magdalena above). Clusters are hand-built and non-exhaustive; they under-count by design, since ambiguous names were removed rather than kept. Reported growth multiples are therefore conservative.

Register semantics — the numbers mean different things per country.

LocaleWhat a cohort entry countsConsequence
no_NOCurrent residents by birth yearIncludes adults born abroad; first-generation layers visible
en_USNames on birth-linked social-security recordsFirst-generation adult arrivals largely absent
fr_FRBirths registered in France, people born abroad excludedFirst generation excluded by construction; all French figures are lower bounds
es_ESBirths by decade, top 5,000 names per decadeTail truncated; small clusters may be cut off

Evidence for the Norwegian resident-register reading: the 1941 cohort totals 297,799 against roughly 620,000 actual Norwegian births in 1941–1950, and Polish-origin names peak at the 1971 birth cohort — a pattern impossible in a Norwegian birth register, since Poles were not being born in Norway in 1971 in those numbers. All four locales show a 1941 cohort at 0.43–0.50 of their peak decade, so the 1941 column is systematically the thinnest everywhere and comparisons anchored on it should be read as indicative.

Historical context — the migration histories referenced (Polish migration to Norway after EU accession in 2004; Vietnamese refugee arrivals in Norway from the late 1970s; Maghrebi labour migration to France from the 1950s–60s; Spanish–Moroccan historical ties) is standard historiography, provided to interpret the curves. It is not derived from our data and is not presented as a finding. What our data shows is the shape of the curves; the attribution of a shape to a historical cause is an interpretation, and where a curve admits more than one explanation we have said so.

What we do not claim. We do not claim these figures measure migration volumes, population composition, or ancestry. They measure the frequency of names in registers. A name layer can grow because people arrived, because resident families changed naming practice, because reporting improved, or because a name became fashionable independently — and cohort shape distinguishes some of these but not all.

Related reading: How Migration Plants Surnames runs the equivalent analysis on surnames; Your Name Is a Birth Certificate covers ordinary generational turnover, which is the larger effect; Name Data Pitfalls (companion article, not yet published) and How to Audit Name Frequency Data cover verification method.

← Artigos