← The Constellation
The Seminar · one field, examined

A Homeland and a Date

Thursday · July 23, 2026 · How comparative linguists, computational phylogeneticists, and ancient-DNA labs came to blows over the homeland of the world's largest language family.
I · Seminar

Historical linguistics / whose method locates a lost homeland

A two-century-old craft for resurrecting dead languages, and its three-way war with gene-tree software and ancient DNA over where Indo-European began.

Half the planet wakes up speaking a descendant of one lost language, and its experts cannot agree on where the first speakers stood. Historical linguistics reconstructs that vanished mother tongue, Proto-Indo-European, from the fossil-like sound patterns shared by Hindi, Greek, Russian and English. For two centuries the only tools were a scholar's ear and the fine print of dead grammars, but evolutionary biologists have since arrived with gene-style family-tree algorithms, and geneticists with skeletons full of ancient DNA. The result is a three-way war over a homeland and a date, and its deepest fault line is not steppe versus mountain but whether an algorithm can out-argue a philologist. An outsider should care because it is a rare open window onto a field deciding, in public, what counts as proof.

The field in brief

The evidence a historical linguist trusts most is a coincidence too regular to be one. When the word for 'father' begins with an f in English but a p in Latin (pater), Greek and Sanskrit, and that same f-for-p swap recurs across hundreds of words, the pattern is a sound law, a systematic correspondence that lets scholars run the change backward and reconstruct an ancestral form no one ever wrote down. Words linked by such laws are cognates, and the method of stacking them to rebuild a lost parent language is the comparative method, the field's two-century-old engine.[10]

That engine reconstructs shapes, not dates. To pin an age, linguists once counted how much shared basic vocabulary two languages had kept, a kind of stopwatch called glottochronology that most of the discipline discarded as unreliable by the 1970s. The idea returned in borrowed armor once evolutionary biologists noticed that sets of cognates behave rather like genes, and that the software built to grow dated family trees of species could be aimed at words instead. This is Bayesian phylogenetics, and its arrival is what turned a quiet reconstruction problem into a border war.[1]

Two homelands have dominated the argument for decades. The steppe, or Kurgan, hypothesis places the first speakers among the Yamnaya, mobile herders of the grasslands north of the Black and Caspian seas around 3500 BCE, whose horses and wheeled wagons carried the language outward.[9] Against it stands the Anatolian hypothesis, which sends the language out of what is now Turkey two or three thousand years earlier, riding the slow spread of farming.[9] A quieter third idea shadows both: the Indo-Anatolian hypothesis, first floated in 1926, holds that the Anatolian branch known from Hittite split off first and deepest, so far back that its parent may predate 'Proto-Indo-European proper' by a millennium.[9]

Map of Indo-European migrations radiating from the Pontic-Caspian steppe
The steppe model of Indo-European dispersal from the Pontic-Caspian grasslands, the map ancient DNA has largely borne out for the Bronze Age. [7]

The fight

The first shots came from a laboratory, not a library. In 2012 a team led by the biologist Remco Bouckaert and the psychologist-turned-phylogeneticist Quentin Atkinson published a family tree in Science that mapped Indo-European's birthplace to Anatolia and dated the root to roughly 8,000 to 9,500 years ago.[8] To many historical linguists this read as an occupation. A book-length rebuttal accused the phylogeneticists of 'mismodeling' the past, of feeding a black box wordlists it could not actually read and then trusting whatever tree fell out.[11]

Then a different laboratory seemed to settle it the other way. Beginning in 2015, ancient-DNA teams reading whole genomes from Bronze Age skeletons found a massive pulse of steppe ancestry sweeping into Europe around 3000 BCE, exactly the Yamnaya expansion the Kurgan model had predicted.[7] For a while the steppe looked victorious and the phylogenetic trees looked like an embarrassment.

The truce broke in July 2023. Paul Heggarty, a linguist then at the Max Planck Institute for Evolutionary Anthropology, assembled more than eighty specialists to rebuild the raw data from scratch, a vetted database of core vocabulary across 161 languages, 52 of them ancient.[6] Run through a newer 'ancestry-enabled' phylogenetic model, it returned a root about 8,100 years old and, startlingly, a compromise map, with an ultimate homeland south of the Caucasus and a secondary staging ground up on the steppe.[1] Heggarty called it a hybrid hypothesis, and it conceded something to everyone while satisfying almost no one.[6]

The comparative linguists struck back hard. George Starostin, of the Moscow school of long-range reconstruction, circulated a withering review, and with Alexei Kassian published a formal critique in 2025 arguing the new tree was statistically thin, only 170 concepts stretched across 161 languages, and that it produced groupings no specialist could accept, including a Hittite-Tocharian pairing they called impossible.[5][2] Undetected loanwords, they wrote, had been mistaken for shared inheritance, and the honest reading of the sound laws still put the family's breakup in the first half of the fourth millennium BCE, far younger than Heggarty's date.[2]

While the linguists fought over trees, the geneticists returned with a map. In February 2025 David Reich's Harvard laboratory, with the co-lead authors Iosif Lazaridis and Nick Patterson and the archaeologist David Anthony, published two papers in Nature naming a specific ancestral population, the Caucasus-Lower Volga people, who lived between the north Caucasus and the lower Volga about 6,500 years ago.[3] From this single source, they argued, descended both the Yamnaya and the earlier-branching Anatolians, giving genetics what one author called the first unified picture behind all Indo-European languages.[7]

Heggarty refused to read the genetics as a defeat.[4] The DNA papers, he pointed out, present no language data at all, and a skeleton's ancestry is not its speech.[4] The trace of steppe ancestry in early Anatolia was too small, around a tenth, to plausibly carry a wholesale language shift, he argued, so the 2025 findings were less a vindication of the steppe than a quiet retreat toward his own southern homeland.[4]

What the fight reveals

Strip away the homelands and the quarrel is about three instruments measuring three different things while pretending to measure one. The comparative method reads words, ancient DNA reads genes, and archaeology reads pots and graves, and nothing guarantees that a people, their genome and their language ever traveled together.[4] That is the genuinely unresolved core, and every camp knows it, because no skeleton has ever been recovered with a recording of the language its owner spoke.[3]

On the narrower factual questions the evidence does lean. The steppe migration into Bronze Age Europe is now about as solid as deep prehistory gets, confirmed independently by hundreds of genomes, so Heggarty's near-total demotion of the steppe sits at the field's margin.[3] Most Indo-Europeanists likewise reject his specific tree, the Hittite-Tocharian clade above all, as an artifact of the algorithm rather than a fact of the languages.[2] The 2025 synthesis, a deep root near the Volga and Caucasus with the Anatolians branching early and the Yamnaya spreading the rest, is where the weight of current evidence rests.[3]

What remains open is smaller than the century of shouting suggests, and sharper for it. Whether the ultimate cradle sat just north or just south of the Caucasus, and how many centuries older than the Yamnaya the deepest split truly is, the Caucasus-Lower Volga population runs straight across that seam and the data cannot yet cut it in two.[3] The comparative linguists and the phylogeneticists are, for now, arguing over a border a few hundred kilometers and a few hundred years wide, each convinced the other's instrument cannot see it.[4]

Sources
  1. Heggarty et al., 'Language trees with sampled ancestors support a hybrid model for the origin of Indo-European languages,' Science (July 2023). The paper that reopened the war: a rebuilt core-vocabulary database run through ancestry-enabled Bayesian phylogenetics, yielding a ~8,100-year root and the 'hybrid' south-of-Caucasus homeland.
  2. Kassian & Starostin, 'Do language trees with sampled ancestors really support a hybrid model for the origin of Indo-European?,' Humanities and Social Sciences Communications (2025). The comparative linguists' formal rebuttal, naming the Hittite-Tocharian clade as impossible and arguing for a younger fourth-millennium breakup.
  3. Lazaridis et al., 'The genetic origin of the Indo-Europeans,' Nature (Feb 2025). The ancient-DNA paper naming the Caucasus-Lower Volga population as the common source of both the Yamnaya and the Anatolian speakers.
  4. Heggarty, 'Beating the retreat from the Steppe hypothesis' (blog commentary on Lazaridis et al. 2025, Feb 2025). Heggarty's direct response, arguing the DNA papers carry no language data and that genes are not languages.
  5. George Starostin, informal review of Heggarty et al. 2023 (starlingdb.org). The opening salvo from the Moscow school of reconstruction, cataloguing the tree's implausible nodes.
  6. University of Jena press release, 'New insights into the origin of the Indo-European languages' (2023). Accessible account of the Heggarty dataset (161 languages, 52 ancient), the ~8,100-year date, and the hybrid claim, naming the Jena and Max Planck researchers.
  7. Harvard Gazette, 'Landmark studies track source of Indo-European languages' (Feb 2025). Field-press account of the Caucasus-Lower Volga studies, the Reich/Lazaridis/Patterson/Anthony team, and the Yamnaya spread.
  8. Bouckaert et al., 'Mapping the Origins and Expansion of the Indo-European Language Family,' Science (2012). The earlier phylogeographic round that first placed the homeland in Anatolia and provoked the linguists.
  9. Wikipedia, 'Indo-Hittite languages.' Background on the Indo-Anatolian hypothesis (Anatolian splitting first, Sturtevant 1926, Kloekhorst, Kroonen) and the competing homeland framings.
  10. 'The Comparative Method,' in The Handbook of Historical Linguistics (Wiley). Reference for the field's core method of sound laws, cognates, and reconstruction.
  11. GeoCurrents, 'Mismodeling Indo-European Origin and Expansion' (2012). Contemporaneous linguists' critique of the 2012 phylogenetics, the source of the 'mismodeling' framing later expanded into the book The Indo-European Controversy.