Identidade falsa > Artigos > One French Boy in Ten Was Called Jean. Today's Number One Gets 1.96%. The Collapse of Name Concentration

Este artigo ainda não foi traduzido para Português — você está lendo o original em English. Também disponível em:Deutsch, English, Українська

One French Boy in Ten Was Called Jean. Today's Number One Gets 1.96%. The Collapse of Name Concentration

Among French boys born in the 1940s, 10.22% were named Jean. One in ten. The ten commonest boys' names covered 46.79% of the entire cohort — you could name half the boys in France from a list of twelve.

Among French boys born in the 2000s, the number one name is Lucas, at 1.96%. The top ten cover 14.62%. To reach half of that cohort you need 70 names, not twelve.

This is not French. It is happening in every registry we can measure, at roughly the same rate, over the same six decades. Given-name concentration in the developed world has collapsed by a factor of two to four within living memory, and it has done so without anyone deciding it should.

Surnames, meanwhile, have barely moved. We measured their concentration in Surname Concentration Curves and the numbers there are structural facts about states and legal systems, changing on a timescale of centuries. Given names are the opposite: chosen fresh every time, by millions of independent decisions, and the aggregate of those decisions has shifted dramatically inside one lifetime.

Here is the shift, and — as ever — the part of it we had to throw away.

The collapse, four registries

Share of each birth cohort covered by that cohort's ten commonest given names:

CountrySex1941195119611971198119912001
FranceM46.79%37.83%32.61%29.89%21.91%18.79%14.62%
FranceF32.23%29.76%34.69%25.71%19.97%16.91%14.47%
USAM34.12%31.29%28.17%25.21%21.86%15.68%10.94%
USAF24.81%20.67%14.82%16.50%16.88%11.51%8.83%
NorwayM28.74%25.73%22.69%16.82%16.64%17.10%15.86%
NorwayF27.48%24.78%21.86%17.82%14.92%14.75%16.24%
SpainM37.88%32.54%26.72%20.45%19.05%19.53%20.81%
SpainF28.79%25.89%23.03%17.08%17.07%20.70%23.61%

France's male column falls from 46.79% to 14.62% — a 3.2-fold dilution. The United States falls 3.1-fold for boys and 2.8-fold for girls. Norway falls 1.8-fold.

The same collapse, read as "how many names do you need":

Country / sexNames for 50% (1941 → 2001)Names for 95% (1941 → 2001)
France M12 → 70104 → 2,021
France F24 → 90211 → 2,848
USA M23 → 95514 → 3,703
USA F42 → 200964 → 8,058
Norway M27 → 51206 → 365
Norway F28 → 55229 → 428

The American girls' row is the most extreme measurement in this article. In the 1940s, 964 names covered 95% of American girls. In the 2000s it takes 8,058 — an eightfold expansion of the effective repertoire in sixty years.

And the top names themselves have shrunk to a degree that makes "most popular name" almost a category error:

Country / sex1941 number one2001 number one
France MJean 10.22%Lucas 1.96%
France FMarie 7.11%Léa 2.06%
USA MJames 5.06%Jacob 1.36%
USA FLinda 4.96%Emily 1.23%
Norway MJan 4.28%Jonas 1.74%
Norway FAnne 5.08%Emma 2.13%
Britain MJohn 11.47%*Jack 2.64%

The British 1941 figure is inflated by a data artefact — see the exclusions section.

Jean at 10.22% and Lucas at 1.96% are both "the most popular boys' name in France." They describe populations five times different in size. Any sentence comparing a modern number-one name to a historical one without the percentages attached is close to meaningless.

Why it happened

The mechanism is not mysterious, but it is worth separating from the folk explanation.

The folk explanation is that parents became more creative. That is at most half of it, and it is the less important half.

The larger driver is the collapse of the naming pool's gatekeepers. In 1941, French given names were legally restricted: the loi du 11 germinal an XI confined legal first names to saints' names and figures from ancient history, and it stayed in force until 1993. Catholic naming practice narrowed the field further, and godparent and grandparent naming conventions narrowed it again. A French parent in 1941 was choosing from a list of perhaps a few hundred socially available names, and realistically from a few dozen.

France's curve is the steepest in our data, and France is the country that had the most explicit legal restriction and repealed it inside the measured window. That is not a coincidence, though it is not proof either — the 1993 repeal comes late in a decline that was already well under way by 1971.

Three further mechanisms, none of which require anyone to become more creative:

  • Religious naming conventions weakened. Marie at 7.11% of French girls in the 1940s was not a fashion; it was a norm. Its decline is the decline of the norm.
  • Naming after relatives declined. Repeating grandparents' names is a powerful concentrating force, because it re-samples from the previous generation's already-narrow pool. Break the loop and the pool widens by default.
  • The information environment inverted. Parents in 1941 knew the names in their village and parish. Parents in 2001 had baby-name books, then the internet, and — crucially — published popularity rankings, which let them actively avoid the top of the list. Published statistics on name popularity are themselves a de-concentrating force.

That last point is the loop worth noticing: the better a country measures its name distribution, the more its parents can act on that measurement, and the flatter the distribution becomes.

Spain runs the other way

Every country in the table falls from 1941 to 1981. Then Spain turns around.

Spain19411961198119912001
Male37.88%26.72%19.05%19.53%20.81%
Female28.79%23.03%17.07%20.70%23.61%

Spanish girls' names hit their diversity peak in the 1980s and have been re-concentrating ever since — from 17.07% back up to 23.61%, recovering more than half the ground lost since 1941. Spain is the only country in our set that reverses, and the female reversal is stronger than the male one.

Norway shows a faint hint of the same thing (female 14.75% in 1991 → 16.24% in 2001), but Spain's is unambiguous and sustained across two decades.

We can measure this cleanly. We cannot explain it from our data, and we are not going to pretend otherwise. What the top-three lists show is a shift in kind:

Spain, femaleTop three
1941Maria Carmen 5.39%, Carmen 3.80%, Maria 3.77%
2001Maria 3.73%, Lucia 3.66%, Paula 3.07%

The 1941 list is dominated by Marian compound names — the old Catholic naming system, which concentrated enormous mass onto María and its compounds. The 2001 list is short, modern and secular in form, yet more concentrated at the top than 1981's list was. This is not the old system returning. It looks like a new, narrow fashion consensus — but "looks like" is as far as our data licenses us to go. A registry that counts names cannot tell you why parents chose them.

One caveat specific to Spain: the Spanish source publishes the top 5,000 names per decade and no more. That cap does not affect the top-10 column, which is what the reversal is measured on, but it does mean Spanish tail figures are truncated and Spain is absent from the "names for 95%" table above.

Boys used to be more predictable than girls. Not any more.

A consistent secondary pattern: in 1941, male names were more concentrated than female names in every country measured. By 2001 the gap has closed, and in two countries inverted.

Country1941 M − F gap2001 M − F gap
France+14.56 pp+0.15 pp
USA+9.31 pp+2.11 pp
Spain+9.09 pp−2.80 pp
Norway+1.26 pp−0.38 pp

The traditional asymmetry is well documented in onomastics: boys carried patrilineal and religious naming obligations — the father's name, the grandfather's name, the saint's name — while girls' names were freer and turned over faster. Jean, James, Jose, Jan: the 1941 male number-ones are exactly the names those obligations produce.

What our data adds is that the asymmetry has essentially disappeared within the measured window, and in Spain and Norway has flipped sign. Boys' names in Spain are now measurably more varied than girls'. Whatever obligation was concentrating male names in 1941 is no longer operating.

What our given-name corpus can and cannot tell you

Everything above comes from four national birth registries with real counts. We also hold a given-name corpus for every locale — 63 of them when this audit was run, 64 today — and the obvious move is to rank them all by concentration. We tried. Most of the result is unusable, and the way it fails is instructive.

First: the top-line comparison is a build artefact. Ranked on merged male+female corpus weights, the extremes look dramatic — Georgian at 25.02% against Swedish at 7.96%, a threefold gap. It is not a fact about Georgia or Sweden. Auditing the male given-name files by scale ceiling:

Scale ceilingLocalesEntries (male file)Mean top-10Range
406253–30024.60%23.05%–25.63%
10056250–71124.61%20.22%–33.14%
25512,01513.21%

Swedish was the only given-name corpus on the 255 scale when this was measured, and it held 2,015 male entries against 250–711 for everyone else. It was a later build generation with a deeper tail and an uncapped head. Its low concentration measures our build pipeline, not Swedish parents. Comparing Georgia to Sweden compares corpus versions. (Swedish given names have since been rebuilt onto real registry counts, which removes this particular locale from the comparison but not the problem.)

Reassuringly, the ceiling-40 and ceiling-100 groups have almost identical means (24.60% and 24.61%), so those 62 locales are comparable with each other. Excluding Sweden and using per-sex figures, the honest range is:

Most concentrated (male top-10)Least concentrated (male top-10)
Russia33.14%Saudi Arabia20.22%
Georgia32.14%Norway20.68%
Portugal31.16%Germany20.73%
Nepal30.22%Finland20.76%
Kazakhstan29.48%South Africa20.89%

A 1.6-fold spread, not threefold. Real, but far less dramatic than the headline version — and note that Norway's corpus figure (20.68%) sits well above its actual 2001 registry figure (15.86%), because the corpus represents a living population spanning many cohorts rather than one birth decade.

Second: use per-sex figures, never merged. Merging male and female lists roughly halves the top-10 share for arithmetic reasons — the top ten of a merged list is drawn from two separate distributions. Every merged figure in our first pass was about half its per-sex counterpart, which is why the merged Georgian 25.02% and the per-sex Georgian 32.14% are the same country.

Third: there is a residual depth confound. Within the ceiling-100 group, correlation between log corpus size and top-10 share is r = −0.283. Deeper corpora score flatter, partly for real reasons and partly mechanically. It is weak enough not to invalidate the ranking and strong enough that adjacent positions in it mean nothing.

Fourth — and decisively: the ranking carries no country signal at all. Fifty-six locales spanning Russia, Japan, Saudi Arabia, Vietnam, Iceland and Uganda all land between 20.2% and 33.1%, most between 22% and 27%. Real naming cultures are not that uniform, and we were able to prove they are not, because six of these countries have real birth counts in the cohort files. Measuring the same statistic both ways:

SourceRangeSpreadSD
Real birth counts (6 locales)9.88 – 33.32%23.45 pp6.00
Our corpus weights (same 6)20.68 – 24.84%4.16 pp1.36

Real concentration varies 3.4-fold across those six countries. Our weights vary 1.20-fold. The spread is 5.6× narrower and the standard deviation 4.4× smaller — and worse, the Spearman rank correlation between our figures and the real ones is 0.056. Our corpus does not merely compress the differences between countries; it does not rank them correctly either. Individual errors reach 13 pp: real US female top-10 is 9.88% where our corpus says 23.01%.

A second signature confirms it. Real registries show a large male/female concentration gap — 7.41 pp on average across the six cohort locales, and 13.21 pp in en_IE, a locale built from real counts after our calibration pass. Across the 63 calibrated locales in that pass the gap averages 1.81 pp. The calibration flattened a real asymmetry by roughly fourfold — the very asymmetry this article documents from cohort data.

So the whole-corpus given-name ranking is not a weak finding about countries; it is not a finding about countries. It is reported here only to document what our corpus contains and why it should not be used for cross-country comparison. Every cross-country claim in this article comes from Source A, the birth registries.

This applies to given names only. Surnames are a separate system with separate evidence, and the defect does not carry over — see the note below.

Why this does not apply to surnames

It would be easy to read the section above as "the corpus is unreliable," and that is not what the evidence says. The defect is specific to given names, and it has a specific cause: the validator documented a single expected top-10 band for given names in every country, so calibration converged on it. For surnames the validator deliberately sets no shared target, on the explicit grounds that no such target exists — Italian surnames top out near 0.3% for the leading name while Vietnamese ones reach 30%.

Absent the causal agent, the effect is absent too. Three independent checks:

CheckGiven namesSurnames
Spread of top-10 across locales2.75× (CV 0.144)16.45× (CV 0.783)
Validator's cliff probe: w[10th]/w[11th] > 1.80 of 64 locales
Match against published registriesnot testable per-localezh_TW ±0.02 pp, vi_VN ±0.27 pp

The cliff probe is the direct test: fitting a curve to a top-10 target leaves an artificial step between rank 10 and rank 11. Across all 64 surname corpora the median step is 1.06, the maximum 1.30, and the same ratio measured at ranks 5, 15 and 20 is statistically indistinguishable. There is no step. The surname curves were not fitted to a top-10 number.

Surnames have a different defect — floor pinning, the same 8-bit compression described below for Korea, which in 43 of 64 locales parks more than 5% of the corpus's probability mass on entries all sharing the minimum weight. That distorts the tail, not the cross-country ranking of heads.

A note on the Korean discrepancy

In the surname article we flagged that our Korean top-10 computed to 60.6% against KOSTAT's 63.9%, and reported it as an unexplained gap. It has a known cause, and it is a scale limitation rather than a data error.

Until recently our weights were capped at 255. In Korea that ceiling is catastrophic, because the real Korean surname distribution spans a range no 8-bit scale can hold: 김 (Kim) has roughly 1,527,138 bearers against 1 for the rarest surnames — a ratio of over a million to one, compressed into 255:1. The compression does not hurt the head much; it destroys the floor. Roughly 130 ultra-rare surnames sat pinned at an over-valued minimum weight, contributing on the order of 8% of phantom mass that belongs to nobody, and that phantom mass dilutes the top-10 share.

The ceiling has since been lifted, and zh_TW and vi_VN have been converted to true registry counts — which is precisely why those two now match their registries to two decimal places while Korea does not. Korea has not been converted yet. The 3.3-point gap is the visible residue of an 8-bit scale meeting a millionfold distribution, and it should close when Korea is rebuilt.

This is worth stating because it is a general lesson about weighted corpora: a compressed scale distorts the tail far more than the head, and the damage surfaces as a deficit in the head's share. If your top-10 comes out low against a registry, suspect your floor before you suspect your leaders.

Exclusions

Two datasets were dropped from the headline figures, for the same reasons documented in Migration Layers in Birth Cohorts.

Britain publishes roughly a top-100 list before 1971 and everything down to a single birth afterwards:

DecadeDistinct male namesMale top-10 as published
194111656.88%
196112241.13%
19714,25933.94%
20019,39218.17%

The 56.88% is computed against a denominator of only 116 names. It is not comparable to the 18.17% computed against 9,392, and the apparent collapse between them is mostly a change in publication policy. Britain's figures appear above only where explicitly marked.

New Zealand publishes at a constant threshold but only 239–565 names per decade — deep enough to be internally consistent, too shallow for its top-10 share to be compared with America's 23,079-name files. Its trend (32.43% → 18.46% for boys) is directionally consistent with everyone else and is reported nowhere as a headline.

Conclusion

Given-name concentration has fallen by a factor of two to four in six decades across every registry that can measure it properly. The effective repertoire of American girls' names expanded eightfold. France went from twelve names covering half its boys to seventy. The number-one name, that perennial press-release statistic, has shrunk from one child in ten to one in fifty.

Two things resist the trend. Spain has been re-concentrating since the 1980s, and we can measure that without being able to explain it. And the old male/female asymmetry — boys named after fathers and saints, girls named freely — has not just weakened but inverted in two countries.

The finding we deleted is worth as much as the ones we kept. Ranked naively, our own corpora say Georgian names are three times more concentrated than Swedish ones. They are not. Sweden is on a different build of our pipeline, and the comparison measures us. The rule that catches it is the same one that has caught everything else in this series: before comparing two numbers, check that they were made the same way.


Methodology and sources

Two independent data sources, deliberately kept apart.

Source A — national birth registries with real counts. Cohort files give per-name birth counts across seven decade buckets (1941–2001) for six locales. All headline figures come from here. These are real counts, not relative weights.

LocaleDistinct names/decade (M, 1941→2001)Reporting thresholdUsed for headline
fr_FR1,796 → 10,0395, constantYes
en_US5,934 → 23,0795, constantYes
no_NO465 → 7254, constantYes
es_ES5,000 flat (top-5,000 cap)6→32, driftingTop-10 only; excluded from tail figures
en_GB116 → 9,39232→1, breaks at 1971No (marked where shown)
en_NZ239 → 56510, constant, shallowNo

Source B — our 63-locale corpora. Relative frequency weights, used only in the section explicitly about their limits. Only 5 of the 63 given-name corpora carry real registry counts (en_US, fr_FR, es_ES, no_NO, en_NZ); the other 58 use relative integer scales, 51 of them capped at 100. No headline claim in this article rests on Source B.

⚠ Two cautions on that count. First, it is specific to given names: our surname files are registry-derived in more locales, so a locale can hold real counts on one side and a relative scale on the other — zh_TW and vi_VN are exactly that case, and earlier versions of this note misattributed their surname scales to their given names. Second, "registry-derived" is not one thing: elsewhere in the corpus vi_VN stores shares (per-mille × 1,000) and uk_UA mixes real counts, a flat "≥20,000" band and a placeholder floor in a single column. Counting such locales together as "real data" flatters them. en_IE, rebuilt from real counts, is the 64th locale and postdates this run.

Computation. For each list: sort descending by count, compute cumulative share at ranks 10 and 100, and the rank at which cumulative share first reaches 50%, 80% and 95%. Male and female analysed separately throughout; merged figures are reported only to demonstrate why they should not be used. Analysis run 2026-07-18 against corpus version share 1.1.7. ⚠ The corpus held 63 locales when these figures were measured; en_IE was added on 18 July 2026 and it holds 64 today. Since then en_IE, en_GB and sv_SE have moved to registry counts as well, so the "5 of 63" figure above is now 8 of 64 (en_US, en_GB, fr_FR, es_ES, sv_SE, no_NO, en_IE, en_NZ), re-measured 22 July 2026.

Build-generation audit (Source B). Grouping male given-name files by scale ceiling:

CeilingLocalesEntry rangeMean top-10Verdict
406253–30024.60%comparable with ceiling-100
10056250–71124.61%comparable with ceiling-40
2551 (sv_SE)2,01513.21%different build; excluded

Depth confound within the ceiling-100 group: Pearson r(log N, top-10) = −0.283 (n = 56). Across all 63 without exclusions: r = −0.436, inflated by the sv_SE outlier.

Known limitations.

  1. Cohort registers count different objects per country. Norway's is a resident register (its 1941 cohort totals 297,799 against ~620,000 actual Norwegian births in 1941–50), so it includes people born abroad. France's counts births registered in France, excluding people born abroad. The US series is birth-linked social-security records. Concentration is less sensitive to this than migration analysis is, but the series are not identical objects.
  2. All four locales show a 1941 cohort at 0.43–0.50 of their peak decade's total. The 1941 column is systematically the thinnest and the most affected by mortality and record coverage.
  3. Spain's top-5,000 cap truncates its tail; Spain appears in top-10 comparisons only.
  4. The Source B ranking is descriptive of our corpora, not of countries, for the four reasons given in the body.

Historical context — the French loi du 11 germinal an XI (in force 1803–1993), Catholic and Marian naming conventions, and patrilineal naming obligations — is standard onomastic and legal history, provided to interpret the curves. It is not derived from our data. The correlation between France's legal restriction and France's steepest curve is noted as suggestive and explicitly not as proof; the decline was already under way before the 1993 repeal.

What we do not claim. We do not claim to explain Spain's reversal. We do not claim the 63-locale ranking measures national naming cultures. We do not claim causation for any of the four proposed mechanisms behind the collapse — they are consistent with the curves, and the curves alone cannot distinguish between them.

Related reading: Surname Concentration Curves runs the same measurement on surnames, where the numbers are structural rather than fashion-driven; Your Name Is a Birth Certificate covers which names replaced which; Comparing Name Corpora Across Locales covers the cross-corpus failure modes referenced above.

← Artigos