Skip to main content
ExplainerLexical DiffusionExplainer· 6 min read· in Culture

The S-Curve Model: How Lexical Diffusion Maps the Gradual Spread of Sound Change Through a Language

The S-curve model demonstrates that sound changes do not alter a language simultaneously, but rather incubate slowly in a few words before rapidly sweeping through the lexicon.

By Chen Wang

Lexical Diffusion Proponents 45%Social Network Theorists 30%Neogrammarian Traditionalists 25%
Lexical Diffusion Proponents
Argue that sound change is phonetically abrupt but spreads gradually word-by-word through the lexicon.
Social Network Theorists
Focus on the community identity and social pressures that drive the rapid adoption phase of the S-curve.
Neogrammarian Traditionalists
Maintain the historical view that sound changes operate without exception across all applicable words simultaneously.

Perspectives this story doesn't cover

  • Computational linguists modeling future changes
  • Speakers of highly endangered languages where diffusion operates differently

Common questions

What is lexical diffusion?

Lexical diffusion is the theory that a sound change spreads gradually through a language's vocabulary word by word, rather than affecting all applicable words simultaneously.

Why does language change follow an S-curve?

The S-curve reflects social adoption: a new pronunciation incubates slowly among a small group, spreads rapidly once it hits critical mass, and plateaus as the last few holdout words resist the change.

What is linguistic residue?

Linguistic residue refers to the permanent irregularities left behind when a sound change loses momentum and stops spreading before it can affect the entire vocabulary.

Does this model apply to internet slang?

Yes. Modern studies of social media platforms like Twitter show that new slang terms and phonetic spellings follow the exact same S-curve trajectory as historical sound changes.

The short answer

  1. Lexical diffusion proves that sound changes spread word-by-word, not simultaneously across a whole language.
  2. The spread follows a mathematical S-curve: slow incubation, rapid adoption, and a flat plateau.
  3. The theory was formalized in the 1960s after early computers revealed pervasive irregularities in Chinese dialects.
  4. Words that resist the rapid adoption phase become permanent 'linguistic residue,' creating irregular verbs and spellings.
  5. Modern social media data shows that internet slang follows the exact same S-curve trajectory.

The S-curve model reveals that sound changes do not hit a language all at once; instead, they incubate slowly in a handful of high-frequency words before exploding into rapid, systemic adoption and eventually tapering off. Lexical diffusion maps this exact mathematical progression, proving that human speech evolves word by word rather than sound by sound. The traditional assumption among linguists was that when a vowel or consonant shifted, it changed everywhere simultaneously, like a software update applied to an entire operating system. The reality is far messier, resembling a viral trend that starts in a small neighborhood, suddenly sweeps through the city, and leaves a few stubborn holdouts untouched.[3][5]

To understand why this mathematical model upended linguistics, you have to look at the dogma it replaced. In the late nineteenth century, a group of scholars known as the Neogrammarians established a strict rule: sound changes operate without exception. If the Old English long 'a' shifted to a long 'o' (turning 'stan' into 'stone'), it had to happen to every single applicable word in the lexicon at the exact same time. Any exceptions were dismissed as dialect mixing or analogical borrowing. It was an elegant, frictionless theory that made historical linguistics look like a hard science, but it ignored the chaotic reality of how people actually talk.[1][5]

The breaking point for the Neogrammarian view arrived in the 1960s, driven by the sheer computational power of early mainframes. In 1962, Peking University published the Hanyu Fangyin Zihui, a massive dataset containing transcriptions of 2,444 morphemes across 17 modern varieties of Chinese. At the University of California, Berkeley, the linguist William S-Y. Wang and his team digitized this data to apply traditional comparative methods at a scale previously impossible. What they found was a landscape riddled with pervasive irregularities that the old exceptionless rules simply could not explain.[5]

Wang’s team identified stark contradictions in the data. For instance, Middle Chinese words in the third tone class with voiced initials had split into two entirely different forms in the modern Teochew dialect, with no phonetic factor to explain the divergence. They documented 12 specific pairs of words that were perfectly homophonous in Middle Chinese but had evolved completely different modern pronunciations. Under the old rules, this was impossible. Wang realized that these irregularities were not mistakes or dialect anomalies; they were the permanent residue of a sound change that had run out of momentum before it could infect the entire vocabulary.[1][5]

The foundational dataset that revealed the permanent residue of incomplete sound changes.

In his landmark 1969 paper, 'Competing Changes as a Cause of Residue,' Wang proposed a radical new framework: lexical diffusion. He argued that a phonetic change is phonetically abrupt but lexically gradual. When a new pronunciation enters a language, it does not instantly overwrite the old one across the board. Instead, it takes root in a small cluster of words—often those most frequently used by a specific subculture or demographic. From there, it spreads word by word, competing with the older forms until it either takes over the lexicon or stalls out, leaving a trail of linguistic fossils behind.[1][3]

In his landmark 1969 paper, 'Competing Changes as a Cause of Residue,' Wang proposed a radical new framework: lexical diffusion.

This word-by-word spread is not linear; it follows a distinct mathematical trajectory known as the S-curve. The horizontal axis represents time, while the vertical axis measures the number of words in the language that have adopted the new pronunciation. The curve begins with a long, flat latency phase. During this period, the innovation is confined to a tiny fraction of the vocabulary—perhaps 5 to 15 percent—and its growth is agonizingly slow. It is a private, localized quirk, hovering on the edge of extinction as speakers test the new variant in isolated social pockets.[4][5]

If the innovation survives this incubation period and hits a critical mass, the curve suddenly steepens. This is the period of rapid spread, where the new pronunciation cascades through the lexicon. Words fall like dominoes. The social pressure to adopt the new form accelerates, and what was once a marginal variant becomes the dominant standard. This explosive middle phase accounts for the vast majority of the vocabulary shifting in a relatively compressed window of time, creating the steep vertical spine of the mathematical model.[2][4]

The three phases of lexical diffusion: latency, rapid spread, and maintenance.

Finally, the curve flattens out again at the top, entering the period of maintenance or completion. The change has infected 85 to 90 percent of the applicable words, but the momentum slows drastically. The last few holdouts—often highly salient words, or conversely, extremely rare ones—resist the change. Some may eventually succumb, but others never do. These permanent holdouts become the residue that Wang identified in the Chinese dialects, the irregular verbs and bizarre spellings that frustrate language learners centuries later.[1][4]

While Wang built his theory on historical Chinese data, modern digital communication has allowed linguists to watch the S-curve unfold in real time. A 2013 study presented at the Berkeley Linguistics Society tracked the diffusion patterns of lexical innovations on Twitter. By analyzing massive datasets of social media posts, researchers could map exactly how a new slang term or phonetic spelling incubates within a dense, highly connected social network before breaking out into the broader platform. The digital data perfectly mirrored the S-curve trajectory that Wang had theorized four decades earlier.[2]

The mechanics driving this curve are fundamentally social. As the linguist William Croft noted in 2000, the time course of language change typically follows an S-curve because of 'the desire of hearers to identify with the community' to which a speaker belongs. In the early stages, adopting the new pronunciation is a marker of in-group identity. As it spreads, it becomes a marker of broader cultural relevance. By the time the curve hits its steep upward trajectory, the social cost of not adopting the change outweighs the effort of learning it.[4]

Early computational linguistics in the 1960s provided the processing power needed to identify patterns of lexical diffusion.

The S-curve is so robust that linguists now debate whether it applies beyond phonetics to syntax and grammar. Historical data on early Modern English morphosyntax shows the same jagged but recognizable S-shaped forms. For example, the spread of specific sentence constructions—like the shift in how negative imperatives were formed between the twelfth and fifteenth centuries—followed a similar path of slow incubation, rapid adoption, and eventual plateau. Whether this represents true lexical diffusion or a different mechanism entirely remains a point of contention, but the mathematical shape of the change is undeniable.[4]

The mathematical elegance of the S-curve lies in how it accounts for human stubbornness. It explains why a modern English speaker still says 'child' and 'children' instead of applying a uniform plural, or why a Teochew speaker maintains two distinct pronunciations for words that were identical a thousand years ago. The model proves that language does not upgrade all at once. It evolves exactly as people do: in fits and starts, adopting the new when it becomes socially necessary, but always keeping a few irregular fossils safely preserved in the plateau.[1][5]

Why it matters

Understanding lexical diffusion reveals that language is not a rigid set of rules, but a living social ecosystem. It explains why our languages are filled with frustrating irregularities and how modern viral trends follow the exact same mathematical patterns as ancient phonetic shifts.

Jargon, explained

Lexical Diffusion
The hypothesis that a phonetic change spreads gradually across the vocabulary of a language, word by word.
S-Curve
A mathematical graph showing slow initial growth, a period of rapid acceleration, and a final flat plateau.
Neogrammarian
A school of linguistics from the late nineteenth century which argued that sound changes operate mechanically and without exception.
Linguistic Residue
Words that permanently resist a sound change, creating irregularities in the language.
Morpheme
The smallest meaningful unit in a language, such as a root word, a prefix, or a suffix.

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Lexical Diffusion Proponents 45%Social Network Theorists 30%Neogrammarian Traditionalists 25%
  1. [1]Linguistic Society of AmericaLexical Diffusion Proponents

    Competing Changes as a Cause of Residue

    Read on Linguistic Society of America
  2. [2]Annual Meeting of the Berkeley Linguistics SocietySocial Network Theorists

    Language Change as a Social Process: Diffusion Patterns of Lexical Innovations in Twitter

    Read on Annual Meeting of the Berkeley Linguistics Society
  3. [3]BrillLexical Diffusion Proponents

    Lexical Diffusion

    Read on Brill
  4. [4]Cambridge CoreSocial Network Theorists

    Log(ist)ic and simplistic S-curves

    Read on Cambridge Core
  5. [5]WikipediaLexical Diffusion Proponents

    Lexical diffusion

    Read on Wikipedia
  6. [6]Factlen Editorial TeamLexical Diffusion Proponents

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Culture stories with full source coverage and perspective breakdowns delivered to your inbox.