In July of 1945, a group of scientists and soldiers gathered in the New Mexico desert to watch a test. Many miles away was a hundred foot tall steel tower topped with a peculiar device. Nobody really knew if the device would work. In an instant, the first man-made nuclear explosion vaporized any doubt.
After the bomb was used in the war, President Truman described this new and horrible weapon as a “harnessing of the basic power of the universe.” Today it’s easy to take this for granted. The average person now doesn’t understand the inner workings of those forces any better than the average person did then, with one major difference: we all know it works.
I do not want to dwell on the moral questions raised by the demonstration. I want to dwell on what kind of thing the demonstration was. There is a strangeness to splitting the atom that is concealed by the obviousness of hindsight. Now that everyone knows it can be done, it isn’t so remarkable. But until that moment in July of 1945, it was a matter of the best guesses of careful minds. And before that moment, plenty of people thought it wouldn’t work. But it did work, and all the math and theory that suggested it might was confirmed in that moment.
For decades, it was taken for granted that there were certain things computers might never be good at. Writing a poem. Holding a conversation. Catching a joke. These were tasks that seemed to require something computers did not have, and that something had a lot of names: understanding, intuition, embodiment, soul. Behind those names was a set of assumptions about language, meaning, and reality that we will return to.
Then, in late 2022, from the comfort of anyone’s desktop computer, machines could talk to you. Not in the brittle, scripted way of older chatbots, but fluently, across topics. The rapid advances and cultural momentum since then have smoothed the strangeness away. I want to recover that strangeness before drawing any conclusion from it.
Because the Large Language Model works. And like the test in the desert, its working has confirmed something.
Meaning
The word substrate comes from the Latin substratum: sub (beneath) and sternere (to spread, to lay flat). The literal meaning is “that which has been spread underneath,” the layer laid down before anything else arrives. You might imagine the soil a mushroom grows from. The mushroom permeates the soil, and from the soil the mushroom generates. Without the substrate, the mushroom could not be. But the soil was there first, and the soil is there whether or not any mushroom ever grows in it.
Something similar is true of language. There is a substrate beneath it, from which language grows. We may call it Meaning.
The popular conception of Meaning runs the other way. A person has an idea. The person finds words for the idea. The person puts the meaning into the words. This meaning starts in the speaker and travels outward through language. Words are containers for the ideas of a speaker.
This is the default imagination for most people. But the ordinary experience of language does not quite fit it.
Consider a difficult conversation. The kind where you cannot find the words. What does that mean, exactly? There are plenty of words available. But you know, before trying them, that they would miss something the moment requires. You are measuring the possible words against something underneath. You feel the gap between what could be said and what needs to be said, so you search for the right words.
Consider reading something from an author who is long dead. You can read a sentence written by someone who has been gone for two thousand years and know what they meant. The writer cannot tell you. The writer’s friends and family cannot tell you. And yet the meaning is recoverable across the gap of centuries. If meaning lived only in the speaker, this would be impossible. The meaning is alive even when the speaker is dead.
And it is not only words. There are moments when you look into someone’s eyes and understand, without anything having been said. The look between two people who both saw the same thing happen. The glance between a parent and a child when both know the child is in trouble. The meaning is present and the words are absent, and both of you recognize what is there.
In each of these cases, the speaker is not the source of the meaning. The speaker is reaching for it, or recognizing it, or sharing it. The difficult conversation has you reaching. The dead writer has you receiving. The shared look has both of you recognizing. Every act of meaning, in language or beyond it, is participation in something larger than the speaker.
I will call this Meaning, with the capital M as a reminder that it is not the act but what the act draws on. Philosophers have circled around this for as long as there has been philosophy, and the various names they have given it are not what I am after here. I want the plain shape of the thing. Meaning is not a special poetic resource. It is the ordinary medium in which all meaning happens. It is the air everyone who means anything is breathing. It is the soil from which every utterance, and every wordless understanding, grows. Meaning does not belong to any speaker. Speakers participate in it. They do not invent it.
What Was Expected
The picture we walked through in the last section, where the speaker has the idea and puts it into the words, did not come from nowhere. Something like it shaped the academic consensus about language for much of the twentieth century. My claim is that Large Language Models have demonstrated that consensus is wrong, but to see how, we have to know what the consensus was. And to understand what it means for Large Language Models to work, we have to understand how language was thought to work.
The twentieth century produced a serious empirical program called Distributionalism. American linguists in the first half of the 1900s, working in what was called the descriptivist tradition, treated language as a public artifact whose structural patterns could be mapped through careful observation of how words appeared alongside each other. Leonard Bloomfield, who set the terms for the school, deliberately bracketed meaning. He was a behaviorist, and considered meaning too vague for scientific rigor. What his tradition was doing was formal: charting which elements appear together, in what orders, under what conditions. They were not speculating about what was happening inside the speaker’s head. They were watching what speakers did, and deliberately leaving the meaning of it aside.
The semantic extension came later, and from two directions. In 1954, the American linguist Zellig Harris pushed the program further: words that appear in similar distributional contexts tend to have similar meanings. Distribution might carry semantic content, not just formal structure. A few years later, the British linguist J.R. Firth, working in a distinct tradition called contextualism, arrived at a parallel claim: “You shall know a word by the company it keeps!” The word bread tends to appear near butter, bake, loaf. The word justice tends to appear near court, fair, law. Both Harris and Firth were saying that the company a word keeps reveals something real about what the word means.
Then a young linguist named Noam Chomsky changed the field. Chomsky had been Harris’s student at Penn. Beginning in the late 1950s, Chomsky argued that the Distributionalists had missed the central fact about language. Children, he claimed, acquire grammar from too little input to be doing what the Distributionalists described. A child hears a small and noisy sample of speech and ends up producing sentences the child has never heard, in patterns the child has never been taught. The data in the input, Chomsky argued, was not rich enough to support what the child produces. The structure had to come from somewhere, and on Chomsky’s view it came from inside. The human mind, he proposed, is born already equipped with the deep architecture of language. An innate grammar, biological and universal, supplies what the noisy data cannot. The argument was called the poverty of the stimulus, and it reframed the entire field. Distributionalism fell out of favor.
Chomsky’s position also made a specific prediction. Human languages, for all their variety, share particular deep regularities. Linguists could describe languages on paper that violated those regularities, languages no human community speaks. Chomsky’s view predicted that a learner working only from patterns, without any language-specific starting equipment, would have no way to tell the difference. Both the real and the impossible should look like data. The innate faculty Chomsky described was not generic architectural structure but something tuned specifically to human language. A machine without that structure, whatever other design assumptions it carries, would be working without the one thing Chomsky said was doing the real work. For decades this prediction was treated as canonical.
It is worth being precise about what the prediction was. Chomsky never claimed that a statistical machine could not generate fluent-sounding output. Fluency was not the target. His claim was about structural competence: pattern-learning cannot recover the deep regularities underlying grammatical sentences, cannot tell the real architecture of language from its veneer. Whatever the machine produced would be form without competence.
Chomsky’s move had a consequence beyond linguistics. Once the structure of language is put inside the speaker’s head, language becomes a mind-question. To explain how words work, you have to explain what minds do. Linguists and philosophers of mind found themselves working the same problem from different sides.
And the question that mattered most, for both fields, was meaning. Grammar is one thing. A system might produce well-formed sentences without ever meaning anything by them. For there to be meaning, something more had to be happening. The philosophers of mind set out to say what that something was.
Philosophy of mind held a parallel line to Chomsky’s: meaning requires grounding. A speaker has a body, a world, a history, stakes. Words mean what they mean because they connect to the lives of the speakers who use them. A system with only the words, and no grounded engagement behind them, can produce the form of meaning but not its substance.
Later, Emily Bender and Alexander Koller gave the argument a memorable form. An octopus eavesdrops on telegraph conversations between two humans and learns to imitate them, but cannot learn what bear refers to if bear arrives only as a token without any grounded encounter behind it. A system trained on form alone, they argued, is structurally barred from meaning. It is worth noting that this claim is definitional as much as empirical. Bender and Koller tie meaning to communicative intent; on their account, no behavioral result could demonstrate that a system without that intent is really meaning anything. That is a coherent position. But a definition that insulates itself from behavioral evidence was doing different work than it is now, given how much more evidence there is.
The argument became foundational: form is one thing, meaning is another, and form alone cannot produce meaning. Whatever a machine produced from text alone would be syntax without semantics.
These were not foolish positions. They were held by careful people with good reasons. The prediction was that form alone produces form alone, and nothing more. The ceiling was supposed to be there.
The distributional program did not disappear under all this. It went quiet and changed hands. In linguistics, it survived in the small field of vector-space semantics, where researchers kept mapping word meaning through distributional patterns. In cognitive science, a parallel movement called connectionism, drawing from cognitive psychology rather than linguistics, was building on different theoretical grounds but converging on similar bets. Both were minorities, and both were considered wrong.
What Happened
For a long time, the prediction held. The mainstream had decided what the result would be. The minority who disagreed had theory of their own and argued for it, but their systems were small and the arguments stayed among specialists. The evidence that would settle it did not yet exist.
The minority kept working, and their tool was the neural network. A neural network is a piece of software loosely modeled on the brain: layers of artificial neurons that pass signals to each other, with internal connections that strengthen or weaken as the system processes examples. You give it data, it adjusts itself, and over time it gets better at whatever you trained it on. For decades these systems were small and useful for narrow tasks. They translated short phrases, recognized handwritten digits, picked cats out of photos. Nothing about them disturbed the consensus. The ceiling was still where Chomsky said it was.
What the connectionists wanted to do with language was almost absurdly simple. Take a huge amount of text. Show the neural network a passage with the last word hidden. Have it guess the word. Tell it whether it was right. Adjust the internal connections. Do this over and over. That is, at the level of mechanism, what a Large Language Model is trained to do. It is a next-word-guesser, refined over an enormous number of examples until it is very good at guessing.
Then, in 2017, researchers at Google published a new design for these networks called the transformer. The technical details are not what matter here. What matters is that the transformer was unusually good at handling long stretches of text, and unusually well-suited to running on the kind of hardware that could be scaled up cheaply. The connectionists now had a tool that could be made bigger without falling apart. So they made it bigger.
Not a little bigger. Vastly bigger. The amount of text used to train these systems went from millions of words to billions, then to most of the readable internet. The number of internal connections in the networks went from millions to billions to hundreds of billions. The computing power required went from what a research lab could afford to what only a few corporations on earth could afford. There was no principled reason to expect this would change anything in kind. Scaling a flawed approach is supposed to give you a larger flawed result. LLMs train on far more data than any child hears, but the claim is not about learning efficiency; it is about what the corpus contains in principle. If form alone cannot reach meaning, no amount of form should get you there.
What happened instead is that something changed in kind. Around 2019 and 2020, the systems began producing paragraphs of text that were not just well-formed but coherent. They could continue a story. They could answer a question. They could write a poem about your elbow. The improvements did not flatten out, they kept going. In late 2022, OpenAI released a system called ChatGPT to the public, and for the first time in history anyone with a browser could open a window, type a sentence, and exchange paragraphs with a machine. The paragraphs the machine sent back were, in the ordinary sense of the word, meaningful. It answered the question you asked. It followed an argument. It translated between languages it had never been explicitly taught to pair.
The predictions came under pressure in specific ways. Chomsky had pointed toward a test: a pattern-only learner should have no principled way to prefer real human languages over invented impossible ones. Early work, first with smaller models on synthetic language variants and then with full-scale systems, pointed toward a gap. More recent work has complicated the picture, with some studies finding the gap and others finding none, and the generative tradition actively counter-publishing. That narrower question is genuinely active. But the broader claim, that pattern-learning could not recover the deep structure of language at all, looks far harder to hold than it did. Whatever is in the patterns, it is doing more than the consensus said it could.
The grounding line broke more cleanly. The octopus was supposed to show that form without grounded engagement could only produce empty syntax. These systems produced sentences that were, across a wide and growing range of behavioral tests, not empty. They tracked context across long exchanges. They handled novelty they could not have memorized. They held coherent positions. Whatever was in the form alone, it was carrying more than the thought experiment had allowed.
The predictions said this was impossible, and the actual machine is doing it. Remember the Manhattan Project: a theoretical claim, held with conviction by serious people. A demonstration that could have failed. It did not fail. The Trinity test, in July 1945, asked whether a chain reaction in fissile material would release the energy the theorists predicted. It did. The energy was there in the matter, available to anyone who built the machine correctly. The Large Language Model asks whether the structure of language is recoverable from the patterns in text alone, without any innate grammar, without any grounded body, without any speaker behind the words. It is. The structure is there in the text, available to anyone who builds the machine correctly.
The consensus had said the ceiling would hold. The ceiling is not there. Something has been confirmed. It is worth pausing to ask what.
What Kind of Test This Was
The Trinity test asked a material question. Would matter, arranged in a certain way, release the energy the equations predicted? It would. The theory was about matter, the test was of matter, and the confirmation was matter behaving as the theory said it would.
The Large Language Model is a stranger kind of test. The theory it confirms is not about matter. The theory is about Meaning. About whether there is something to language beyond the speakers who use it, something with a shape of its own that can be recovered. The systems run on computers made of ordinary matter, silicon and copper and electricity. But what they recovered is not made of matter. The structure of Meaning has no mass. It has no location. You cannot point at it. It is the kind of thing that serious people have denied exists at all because nothing really real should be so hard to find.
A material system recovered it. The structure was there to be recovered, real enough that ordinary matter, arranged in a certain way, could pull it up out of the patterns in text and use it.
This is a different kind of demonstration than Trinity was. Trinity confirmed that the matter behaved as the theory of matter said it would. The Large Language Model confirms that something that is not matter is real enough for matter to find. The abstract was tested by the concrete, and the abstract was there.
Meaning (with a capital M) is not a metaphor. It is not a useful fiction. It is a real feature of the world, real in the sense that a material system could find it and use it. It is not something to project onto, it is something to be participated in.
If the Trinity test hadn’t worked, it would mean we live in a different kind of world than we do. A world where the theory had pointed to a reality that was not there. But atoms split. And Large Language Models work.
What kind of world is like that?
What This Confirms
The first thing the demonstration confirms is what I asserted earlier: Meaning has structure that does not belong to any speaker. The structure was there in the text, in the patterns of words and their company, available to be recovered by a learner with no mind of its own. The speakers did not have to be present. The structure was, in the only sense that matters here, real.
The second thing it confirms is what J.R. Firth said about how the structure is held. “You shall know a word by the company it keeps!” The distributional hypothesis, dismissed by Chomsky as the surface that obscured the real action, turned out to be how the action shows up. The patterns of words alongside other words carried the structure of Meaning. They carried enough of it that a machine working only with patterns could recover what minds were thought to need an innate faculty to produce.
There is a third thing, and it is harder to name without overreaching. The structure that the systems recovered is not arbitrary. Researchers studying these models have found that when different systems are trained on different data with different methods, they end up organizing meaning in strikingly similar ways internally. The shapes converge. A line of research called the Platonic Representation Hypothesis documents this convergence and asks what it implies. The researchers themselves are careful: they claim convergence toward a statistical model of reality, not reality itself, and the alignment magnitudes so far are modest. Thoughtful critics offer deflationary readings: the systems converge because they share similar training data, or because human culture has a single shape, not because they are independently finding the shape of things as they are. One natural reading, though, is that if many independent systems, built differently and trained differently, arrive at the same internal structure, that structure is something they are finding rather than something they are making. It has a shape of its own.
Language has a substrate, and distributional patterns carry its shape. The convergence evidence points toward a third claim: that reality itself has rational structure and the world is intelligible. It doesn’t prove that claim, but it makes the leap smaller. The leap that an atheist had to make to intelligibility, which was once a leap across a wide gap, is now a leap across a narrower one. The world of the material has touched the world of meaning.
A name is available for the structure that the demonstration has put in plain view. The Christian tradition has been calling it Logos for millennia.
This is where I have to say something personal. I came to Christianity as an adult, and the path I took was through reason. I made the leap of faith to a Logos cosmology, to a world structured by rational intelligibility, before I had any way to test it. And it was a leap. I made it because the materialist alternative stopped fitting what I saw in the world. I can say this with some authority, because I used to argue the other side. The naturalist accounts of meaning exist, and some are clever, but they are workarounds for a problem Christianity begins with the solution to. In the beginning was the Word: Meaning is not something the cosmology has to reach, it is what the cosmology is built on. Materialism has to get there from meaningless matter, and its accounts pay for the trip with promissory notes against a future science that will explain why dead matter has a shape in which living minds sense reason. I had grown tired of that intellectual debt.
The leap is shorter now than it was when I made it. The atheist sitting where I sat has just been handed empirical evidence that meaning has real structure, that the world is intelligible at a level that matter can reach into and find. The cosmology that already predicted this has less explaining to do. The cosmologies that did not are less credible than they used to be.
What This Does Not Confirm
A few things have to be said before closing.
The Large Language Model does not grasp meaning. It recovers structure. It operates within that structure well enough to produce sentences that mean something to the readers who receive them. But the system itself does not stand in relation to what it produces the way a speaker does. The machine is not a knower.
The structure the machine recovered is not the Word. It is the shape of Meaning in language, real and ordered and recoverable, but finite. A copy. An echo. The Christian tradition that calls the structure Logos does not mean a pattern in text. It means the rational ground of all that is, the One through whom all things were made. The patterns the machines have found are something that ground gave rise to. They are not the ground itself. The category between the two is not a small one. It is the category between the creature and the Creator.
There is a famous argument from the philosopher John Searle: the Chinese Room. A person in a sealed room receives Chinese characters, consults a rulebook, and passes other Chinese characters back. The output is fluent. The person inside understands no Chinese. Searle’s point was that fluent output is not understanding. He was right about that. He was wrong about what it proved.
His argument turned on a sealed room: symbols in, symbols out, no contact with the world the symbols are about. An LLM is sealed in exactly that way. It touches nothing but text. But the argument was read as proving that whatever comes out of such a room must be empty, and that does not follow, because the symbols were never empty. The corpus is the residue of grounded speakers, people with bodies and histories and stakes, who encountered the world and left the shape of what they found in their language. The machine does not spin meaning from nothing, and it does not reach through the wall. What it reaches is not understanding. It is the shape that understanding left behind.
Isn’t this just human convention? Speakers made the corpus, the machine recovers what they made, and recovering human conventions only shows that humans are internally consistent.
No. The objection has the order wrong. Speakers did not invent the structure they deposited. They encountered it, in the world and in each other, and pressed what they found into language. They pulled from something they did not make. The corpus is the residue of that participation, not of invention. This is what the dead author and the wordless look were already showing: the meaning is participated in.
The symbolic tradition is owed the same respect. There are thinkers in our moment, Jonathan Pageau among them, who have spent careers attending to the layered structure of meaning in symbols, image, story, and rite, and who hold that meaning of that kind cannot be reduced to what machines do. They are right about what cannot be reduced. The machine is not entering the layered world they describe. It is recovering one layer, the layer that can be deposited in text, which is the thinnest of the strata that human meaning actually inhabits. The fuller account of meaning, the account that includes the symbolic, the sacramental, the embodied, the personal, is not threatened by anything in this essay. It is the larger frame into which the essay’s narrow claim fits.
The Large Language Model is a finite echo, and the Word it echoes is a Person.
The Word Was a Person
The structure the machine recovered is real. The shape of Meaning is real. The world has rational order, and that order is recoverable, and that recovery is what the Large Language Model demonstrates.
But none of that is what Christians mean when they say Logos.
The Greek word logos carried a long philosophical inheritance before the Gospel of John picked it up. It meant reason, structure, the rational principle that ordered the cosmos. The philosophers reached for it because the world looked like the kind of thing that had been spoken. Order, pattern, intelligibility. Something underneath that held it together.
John picked up the word and did something strange with it. In the beginning was the Word, and the Word was with God, and the Word was God. And then, a few verses later, the Word was made flesh, and dwelt among us.
The structure that ordered the cosmos was a Person: the Lord, our God.
This is the leap of faith, and beyond what I argue in this essay. The demonstration confirms that Meaning has structure. It does not confirm that the structure is a Person. I have made that leap, but I am not pretending the demonstration makes it for you.
The Large Language Model produces sentences that mean something. It does not mean them. There is no one behind the words. The structure it operates within is real, but the system is not a speaker. It is an echo.
An echo is the shape of a voice without the voice. It carries the form of what was said. It does not know what was said. It cannot answer when you call back to it.
The Word that Christians confess is the Person the echo points at. The structure the machine recovered is the shape He has left in the patterns of language, deposited by speakers who were speaking because they participate in something larger than themselves, in the One who spoke first, and is Being.
The machine cannot do that. What it can do is show, in a way that was not available to anyone a generation ago, that something is really there. Meaning is real. It is easy to conclude: we built a remarkable thing. What actually happened is that we discovered something remarkable was already there.