Skip to content
Menu
The Relational Economy
  • Recent Research
    • Analysis of Networks & Games
    • Game Theory & Applications
    • The Social Division of Labour
    • Potential Game Theory project
  • Publications
  • About Rob Gilles
    • Professional profile
    • My former PhD students
  • Other
    • Trip to Bosnia, Serbia and The Netherlands
The Relational Economy

On the difference between sounding alike and thinking alike

Posted on 2026-09-07

Nature has just published a long feature warning that generative artificial intelligence is flattening our collective imagination, and has done so without printing a single dissenting voice — an irony I leave the reader to enjoy.

I want to take the feature by conceding its central empirical claim. The most useful number comes from Holzner and colleagues, pooling twenty-eight studies and 8,214 participants. People working with generative AI outperform those working without it on creative performance (Hedges’ g = 0.27); the diversity of the ideas such collaborations drops hard (g = −0.86). An effect of that magnitude is not an artefact of one laboratory’s protocol. Doshi and Hauser found the same structure in the canonical experiment: writers given model-generated story ideas produced stories judged more creative, yet more alike. And the most consequential paper is not about creative writing at all: Hao and colleagues, analysing 41.3 million papers, find that scientists doing AI-augmented research publish 3.02 times more, receive 4.84 times more citations, and lead projects 1.37 years earlier, while the collective volume of topics studied contracts by 4.63 per cent and engagement between scientists drops by 22 per cent.

The private return to AI adoption is enormous and concentrated; the collective cost is modest in percentage terms and diffuse.

So much for the premise. My quarrel is with the step that follows it. What these studies measure is a reduction in the dispersion of expressed form: lexical diversity, syntactic variety, stylistic markers, the semantic spread of ideas. What the coverage claims, however, is that from this follows a reduction in the dispersion of thought. The step between the two is nowhere defended. Padmakumar and He located the mechanism: in co-written essays it is the model’s text that is less diverse, while the human’s contribution is unaffected — direct evidence that form and thought can move separately, at least over the horizons these experiments observe. Their identification is an assumption doing the work of a finding, and the policy conversation about disclosure, watermarking and journal AI rules rests on it.

Style was a signal

Read as a story about culture, this literature invites lamentation. Read as information economics, it makes predictions. Beginning with Spence’s theory, a signal separates types only if its cost differs systematically by type as defined in the concept of the separating equilibrium. Distinctive, fluent, well-organised academic prose has functioned for a century as precisely such a signal — expensive to produce for the writer, because the principal input to good prose is having understood the subject of one’s writing. Nobody wrote this down as a rule. It did not need writing down.

Large language models collapse the cost of producing that signal, and — this is the crucial part — they collapse it more for the writer who previously found it expensive. Doshi and Hauser found the largest gains among the least creative writers; Sourati and colleagues find that polishing preserves core content while homogenising style and amplifying dominant stylistic characteristics at the expense of others. The tool is, by construction, a variance-reducing transformation applied hardest at the lower tail. A separating equilibrium does not survive that. The signal pools.

What, then, does a screener do when a signal stops separating? The screener falls back on whatever correlated observable remains — and here we need not speculate, because we have been running the experiment for decades under the heading of blind review. Tomkins and colleagues, in a controlled experiment at a computer science conference, found reviewers who could see author identities, are markedly more likely to recommend acceptance for papers from famous authors (odds multiplier 1.63), top universities (1.58) and top companies (2.10). Fox and colleagues, in a randomised trial at Functional Ecology, found a large positive bias towards authors in wealthier, more English-proficient countries when identities were visible. Seeber and colleagues, across 21,535 papers at 71 computer science conferences, found single-blind review associated with a smaller share of contributions from newcomers. Three literatures, one message: when reviewers cannot read quality off the work, they read it off the person.

Therefore, the homogenisation of prose does not flatten the profession’s hierarchy. It entrenches it, by removing the one channel through which an unknown author at an unknown institution could signal that they were worth to be taken seriously.

This interpretation is a conjecture with good structural warrant, rather than a result. That referees use identity cues when available is documented; that style itself functioned as a quality signal is an inference from the structure of the problem: the blind-review evidence is all pre-2023. Someone with a journal’s submission archive could settle that in an afternoon: sensitivity to author identity should have risen since 2022 in single-blind venues and stayed flat in double-blind ones.

Who was paying the style tax?

Here is the objection that will occur to a good reader quickly: If style was a costly signal, its cost was not distributed by ability but, in substantial part, by accident of birth.

  • Hengel’s work identifies that female-authored papers in top economics journals are between one and six per cent better written than equivalent papers by men; the gap widens during peer review; women improve their writing over their careers while men do not; and their papers spend around six months longer under review.
  • Card, DellaVigna, Funk and Iorio come at it from the other side: female-authored papers receive roughly 25 per cent more citations than observably similar male-authored ones, conditional on referee evaluation — which is to say the profession is not publishing the best available research, because the bar is higher for women.
  • Amano and colleagues estimate that being a woman, a non-native English speaker and from a low-income country is associated with up to a 70 per cent reduction in English-language publications; Ramírez-Castañeda found Colombian doctoral students paying a quarter to a half of a monthly stipend for editing, with 43.5 per cent rejected or asked to revise on grounds of English grammar.

If the style signal was substantially a tax on being foreign, female or poor, then collapsing it is not a loss of information but the removal of a discriminatory levy, and the pooling equilibrium may be better than the separating one, because the separating one was separating on the wrong variable. The welfare sign is ambiguous, and turns on a question nobody has answered: how much of the variation in style cost tracked research quality, and how much tracked linguistic and demographic accident?

Two wrinkles complicate the optimistic reading. The equity gain is not itself distributed equally:
Agarwal and colleagues found AI assistance delivered greater efficiency gains to Americans than to Indians while pulling Indian writing towards Western norms, and Gao and colleagues find disciplines with more women or Black scientists reaping fewer benefits. And there is a second-order effect I find the darkest joke in the business. As prose converges, we have begun hunting for “AI tells”, using detectors that Weber-Wulff and colleagues find are not accurate enough to bear the weight placed on them. Whether they are themselves biased against non-native writers is contested — Liang and colleagues found they were, Jiang and Al Ali and their co-authors found they were not — but the point survives the dispute: a noisy classifier applied where prior suspicion is unequally distributed produces unequal false-positive burdens whether or not the classifier is biased.

Tails or middle?

Every paper in this literature reports that variance has fallen. Not one asks where in the distribution it fell. This matters, for a reason economists understand better than anyone else in the conversation: research returns are fat-tailed. Seglen’s classic result is that fifteen per cent of a journal’s articles collect fifty per cent of its citations; Bornmann and Leydesdorff, considering nearly three million articles, put close to half of all citation impact in the top decile. If the value of the enterprise sits overwhelmingly in the upper tail, compressing the middle may be an efficiency gain, in the way that removing noise from a signal is a gain. The finding is alarming if and only if it is thinning the tails.

There is evidence it may be doing the opposite. Shypula and colleagues propose measuring effective semantic diversity — diversity among outputs that clear a quality threshold — and find that preference-tuned models, precisely those scoring worst on naive diversity metrics, generate greater diversity once quality is held fixed. We have measured the variance of the thing we can measure and inferred the variance of the thing we care about, which is, of course, precisely the wrong direction of inference.

Blaming the mirror

Homogenisation is being discussed as a property of technology. The evidence increasingly says it is a property of the incentive structure surrounding technology.

  • Jo and colleagues ran a pre-registered randomised trial in which participants used AI interactively under different reward schemes: those rewarded for originality relative to peers wrote collectively more diverse prose than those rewarded for quality alone, and the divergence came not from abandoning the model but from using it differently — fewer verbatim adoptions, more selective use for brainstorming and targeted edits.
  • Wan and colleagues, rerunning Doshi and Hauser’s design with ten diverse AI personas, preserved story diversity outright — the trade-off, they conclude, comes from uniform deployment rather than any inherent limitation.
  • Raghavan gives the formal treatment: competition for attention selects for diverse models and mitigates monoculture — with the corollary that a model performing well on a benchmark in isolation may fail to deliver value in a competitive market.

The implication for academia is direct and uncomfortable. Academic selection does not reward originality relative to peers. It rewards conformity to a template: the referee-proof paper, the fundable grant, the tenurable portfolio. The tools did not create that incentive; they merely made it cheaper to satisfy. An academy that rewarded relative originality would get diversity out of the very same models. The homogenisation is at least partly a revealed preference of the research evaluation system, and blaming the model for it is blaming the mirror.

A final word of humility. The laboratory effect is robust; the field effect is not yet clearly detected. Fitterer and colleagues, comparing news corpora from 2018 and 2024, found the fingerprints of model-preferred vocabulary but no homogenisation on standard lexical measures; Xu and colleagues, across 24.3 million publications, found no significant change in topic variety, only a reallocation of attention within established domains; and Ashkinaze and colleagues found that heavy exposure to AI-generated ideas increased collective diversity.

Nor is the anxiety new. Park, Leahey and Funk dated the decline in disruptiveness to the 1940s, and Leibel and Bornmann have since shown that disruption scores do not match researchers’ own judgements of which of their papers were groundbreaking. Whatever began in 1945 cannot be blamed on a technology released in 2022.

Which is the point. We have become rather good at measuring how alike we sound. Whether we have begun to think alike is a different question, and one we have not yet answered.


References

Sources that inspired this blogpost:

  • Bland new world: is AI making us all think the same? Nature, 1 September 2026. https://www.nature.com/articles/d41586-026-02682-3
  • Braun, R. (2026). Who is responsible when AI helps to write science? Nature 657, 34–36. https://www.nature.com/articles/d41586-026-02686-z

Homogenisation: primary evidence

  • Holzner, N. et al. (2025). Generative AI and creativity: a systematic literature review and meta-analysis. DOI: 10.48550/arxiv.2505.17241
  • Doshi, A. R. and Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances. DOI: 10.1126/sciadv.adn5290
  • Hao, Q. et al. (2026). Artificial intelligence tools expand scientists’ impact but contract science’s focus. Nature. DOI: 10.1038/s41586-025-09922-y
  • Sourati, Z. et al. (2026). The shrinking landscape of linguistic diversity in the age of large language models. Nature Human Behaviour. DOI: 10.1038/s41562-026-02550-0
  • Padmakumar, V. and He, H. (2023). Does writing with language models reduce content diversity? DOI: 10.48550/arxiv.2309.05196
  • Agarwal, D. et al. (2025). AI suggestions homogenize writing toward Western styles and diminish cultural nuances. CHI 2025. DOI: 10.1145/3706598.3713564

Style as signal, and the screening fallback

  • Spence, A. M. (1973). Job market signaling. Quarterly Journal of Economics 87(3), 355–374. DOI: 10.2307/1882010
  • Tomkins, A., Zhang, M. and Heavlin, W. D. (2017). Reviewer bias in single- versus double-blind peer review. PNAS. DOI: 10.1073/pnas.1707323114
  • Fox, C. W. et al. (2023). Double-blind peer review affects reviewer ratings and editor decisions at an ecology journal. Functional Ecology. DOI: 10.1111/1365-2435.14259
  • Seeber, M. et al. (2017). Does single blind peer review hinder newcomers? Scientometrics. DOI: 10.1007/s11192-017-2264-7

Who paid the style tax

  • Hengel, E. (2022). Are women held to higher standards? Evidence from peer review. The Economic Journal 132(648). DOI: 10.1093/ej/ueac032
  • Card, D., DellaVigna, S., Funk, P. and Iorio, N. (2020). Are referees and editors in economics gender neutral? Quarterly Journal of Economics 135(1), 269–327. DOI: 10.3386/w25967
  • Amano, T. et al. (2025). Language, economic and gender disparities widen the scientific productivity gap. PLOS Biology. DOI: 10.1371/journal.pbio.3003372
  • Ramírez-Castañeda, V. (2020). Disadvantages in preparing and publishing scientific papers caused by the dominance of the English language in science. PLoS ONE. DOI: 10.1371/journal.pone.0238372
  • Gao, J. et al. (2024). Quantifying the use and potential benefits of artificial intelligence in scientific research. Nature Human Behaviour. DOI: 10.1038/s41562-024-02020-5

Detection

  • Weber-Wulff, D. et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity. DOI: 10.1007/s40979-023-00146-z
  • Liang, W. et al. (2023). GPT detectors are biased against non-native English writers. Patterns. DOI: 10.1016/j.patter.2023.100779
  • Jiang, Y. et al. (2024). Detecting ChatGPT-generated essays in a large-scale writing assessment. Computers & Education. DOI: 10.1016/j.compedu.2024.105070
  • Al Ali, A. et al. (2026). Different time, different language: revisiting the bias against non-native speakers in GPT detectors. DOI: 10.48550/arxiv.2602.05769

Tails versus middle

  • Seglen, P. O. (1992). The skewness of science. JASIS 43(9), 628–638. DOI: 10.1002/(SICI)1097-4571(199210)43:9<628::AID-ASI5>3.0.CO;2-0
  • Bornmann, L. and Leydesdorff, L. (2017). Skewness of citation impact data and covariates of citation distributions. Journal of Informetrics. DOI: 10.1016/j.joi.2016.12.001
  • Shypula, A. et al. (2025). Evaluating the diversity and quality of LLM generated content. DOI: 10.48550/arxiv.2504.12522

Incentives, design and competition

  • Jo, N. et al. (2026). Incentives shape how humans co-create with generative AI. Preprint. DOI: 10.48550/arxiv.2604.03529
  • Wan, Y. et al. (2025). Diverse AI personas can mitigate the homogenization effect in human-AI collaborative ideation. Computers in Human Behavior: Artificial Humans. DOI: 10.1016/j.chbah.2026.100289
  • Raghavan, M. (2024). Competition and diversity in generative AI. DOI: 10.48550/arxiv.2412.08610

Null results and contradictions

  • Fitterer, S. et al. (2025). Testing English news articles for lexical homogenization due to widespread use of large language models. ACL SRW. DOI: 10.18653/v1/2025.acl-srw.95
  • Xu, X. et al. (2026). The dual impact of generative AI on research: structural stability and attention reallocation. DOI: 10.47989/ir31iconf64268
  • Ashkinaze, J. et al. (2024). How AI ideas affect the creativity, diversity, and evolution of human ideas. ACM Collective Intelligence. DOI: 10.1145/3715928.3737481
  • Park, M., Leahey, E. and Funk, R. J. (2023). Papers and patents are becoming less disruptive over time. Nature. DOI: 10.1038/s41586-022-05543-x
  • Leibel, C. and Bornmann, L. (2026). Is groundbreaking research disruptive? Quantitative Science Studies. DOI: 10.1162/qss.a.462

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Search this site


Top Posts

  • What a game looks like from its second derivative
  • The Radical Irishman Who Got There First: Labour, Specialisation, and the Price of Things
  • Means-Testing in Reverse
  • Thoughts inspired by Rutger Bregman

Top Pages

  • My former PhD students
  • Publications

Pages

  • Recent Research
    • Analysis of Networks & Games
    • Game Theory & Applications
    • The Social Division of Labour
    • Potential Game Theory project
  • Publications
  • About Rob Gilles
    • Professional profile
    • My former PhD students
  • Other
    • Trip to Bosnia, Serbia and The Netherlands

Blog Stats

  • 26,333 hits

Categories

  • AI and its effects
  • Behavioral economics
  • Changes to site
  • Economic institutions
  • Economic theory of money
  • Game theory
  • History of Economic Thought
  • Methodological individualism
  • Methodology of economics
  • Networks
  • Political economy
  • Restructuring the global economy
  • Short poems
  • Social division of labour
  • Socio-economic embeddedness
  • State of economics
  • Theories of economic value
  • Trust
  • Uncategorized
©2026 The Relational Economy | Powered by SuperbThemes