A Word Is Not a Concept
Last week I posted a short version of this argument and a lot of smart people pushed back. So here is the long version, with the receipts.
Fei-Fei Li is right. I want to start there, because most of what follows sounds like disagreement and it isn’t. In her November essay, “From Words to Worlds,” she calls today’s language models “wordsmiths in the dark.” Eloquent but inexperienced. Knowledgeable but ungrounded. She is pointing at something real. A model trained on text has never touched the physical world. It can describe a room it has never stood in.
She is right that there is a hole. I think it is one hole among several, and not the one that decides whether a machine can mean anything.
The short version:
- Fei-Fei Li is right that text models are “ungrounded,” but she measures the wrong axis: fidelity to the physical world.
- Most of what you mean has no physical referent. You never touched justice, debt, or Tuesday. You built them by analogy.
- A machine cannot get embodied experience, so analogy is the road it has left. That road is the one we keep underbuilding.
What she actually said
It is worth being precise, because the argument only works if you take hers seriously first.
Li’s claim is that the next frontier is spatial intelligence. Not bigger language models, but world models that can do three things a text model cannot: generate a consistent world the way a storyteller keeps the furniture in the same place, navigate one the way a first responder moves through a building they have never seen, and reason about it with a physicist’s precision. Her company, World Labs, builds exactly this. Their model Marble turns a single image into a persistent 3D world you can walk through, where the physics stays put instead of melting after four seconds the way video generators do.
The backbone of her worry is old and serious. In 1990 Stevan Harnad named it the symbol grounding problem. How can a symbol mean anything if it is only ever defined by other symbols? Trying to learn meaning from text alone, he wrote, is like trying to learn Chinese as your first language from a Chinese-to-Chinese dictionary. You go in circles forever and never touch the world the words are about.
In 2020 Emily Bender and Alexander Koller made it concrete with the octopus test. Two people are stranded on separate islands, talking through an undersea cable. A hyper-intelligent octopus taps the line and learns to predict every message so well it can impersonate one of them. Then one person, facing a bear, asks how to build a weapon. The octopus has read a million conversations. It has never seen a bear, a stick, or fear. It produces fluent nonsense. You cannot learn meaning from the shape of words alone, they argue, because meaning is a bond between an expression and what the speaker intends by it, and the octopus only ever saw the expression.
This is the rigorous version of “ungrounded.” It is not a hot take. It is a careful position, and if you build AI you should be able to argue for it before you argue against it.
So let me put their case at full strength in one sentence. A machine that has only ever seen text cannot mean anything, because meaning lives out in the world and between people, and text is neither.
Now let me take it apart.
The wrong axis
Here is the move I think the whole conversation is missing. Li measures the gap as fidelity to the physical world. Space, geometry, depth, the way a thrown ball arcs. On that axis she is correct, current models are weak.
But look at what you actually mean on an average day.
Justice. Debt. Tuesday. Inflation. Loyalty. Betrayal. The number seven. A promise. None of these has a shape. None of them arcs through the air. You cannot walk through “justice” in Marble. And yet you understand all of them, with more confidence than you understand the physics of a bouncing ball.
So the interesting question is not “why doesn’t a text model understand space.” It is the question the spatial framing skips:
How did you come to understand justice, when you never touched it either?
You were not grounded in justice by your senses. There is no sense organ for it. Whatever you know about it, you built some other way. And if the concepts you use all day arrived some other way, then grounding-in-the-physical-world is not the load-bearing route to meaning. It is one route. It is not the one most of your mind runs on.
You cannot walk through justice in Marble. You understand it anyway. That is the whole argument.
A word is not a concept
Start from the smallest unit and the trap is obvious.
Take the word “dog.” On the page it is three letters. It is not the concept of a dog. It is a pointer, and a pointer is worthless unless the thing it points to already exists on the other side.
So how do you make the concept exist on the other side? You explain it. But here is Harnad’s circle, made personal: to explain a word you have to use other words, and those words only help if they are already understood. To understand a word it has to be explained. To explain it, it has to be understood. The dictionary defines “dog” with “canine,” and “canine” with “dog,” and around you go.
Nobody actually learns this way, which tells you the circle gets broken from outside the dictionary. And there are really only two tools that break it.
The two engines
The first is experience. You met a dog. It was warm, or it bit you. That encounter laid down a piece of non-verbal ground, a direct print of the world that needs no other word to prop it up. This is the grounding Harnad was hungry for, and Li is right that it matters.
But experience alone is a slow and narrow road. You cannot personally experience the French Revolution, a black hole, or the inside of your own kidney. Most of what an educated adult knows was never touched. So there has to be a second tool, and there is.
The second is analogy. You understand something new by mapping it onto something you already have. Electricity is water in pipes. An atom is a little solar system. The economy runs hot or cold. A firewall. A memory leak. A stream. You did not get these by experience. You got them by taking a structure you already owned and stretching it over new ground.
This is where cognitive science has landed. Lakoff and Johnson made the case in 1980 in “Metaphors We Live By”: abstract concepts are not grasped directly, they are grasped through metaphor from concrete, bodily experience. We handle time as money because we have handled money. We treat an argument as a war because we have pushed and been pushed. In their words, you cannot think abstractly without thinking metaphorically. Hofstadter and Sander put it even harder in 2013: analogy is the fuel and fire of thinking. Not a rhetorical trick, the actual machinery, from the smallest flicker of recognition to the largest scientific leap. Dedre Gentner gave it the formal spine decades earlier: an analogy is an alignment of relationships, carrying the structure of a known domain onto an unknown one.
Now, I am not claiming these are the only two ways knowledge ever moves. Someone tells you a thing and you trust them. You are born already expecting objects to persist and small sets to add, the core priors developmental psychologists have measured in infants. But being told just hands you a concept to slot in by analogy to ones you have, and innate priors are the seed bed, not the crop. The two engines that actually turn a word into a concept on your side are these: experience lays down a basis, and analogy spends that basis to reach the next thing. Which becomes new basis. Which funds the next analogy.
That is the real shape of understanding. Not a pile of grounded facts. A compounding loop. A small amount of direct contact with the world, leveraged over a lifetime into an enormous structure almost none of which you ever touched.
Which is why meaning is personal
Now the part the spatial framing cannot express.
If a concept is built out of your experience and your stock of analogies, then the concept behind a word is not a fixed public object. It is privately rebuilt, every time, on the listener’s side.
I say “dog.” Maybe I grew up with one and the word unpacks into loyalty, warm fur, a specific face that has been dead for twenty years. You say “dog” and you were bitten at six, and the same three letters unpack into a flash of teeth and a jump in your pulse. Same word. Different concept. We do not share a world model of “dog.” We share a pointer, and each of us rebuilds a different world behind it.
Wittgenstein got here first. Meaning is use, he argued, not a label hung on a thing. There is no private object called the-meaning-of-dog sitting in a public warehouse. There is only what the word does inside a particular life.
Which is why I keep saying language is a compression of the world, not a description of it. When I say “dog” I do not transmit the dog. I transmit a tiny code, and the code decompresses on your hardware, using your data. It is lossy, brutally lossy, but it is lossy at the level of meaning, and meaning is the one level a perfect video of a dog never reaches. A renderer can show you the wet nose in 4K and tell you nothing about what a dog means to the person watching.
Language is not a description of the world. It is a tiny code that decompresses on someone else’s hardware, using their data.
So language is a world model too. Just not the kind Li is building. Hers learns the structure of space. The human one learns the structure of meaning, and it learns it almost entirely by analogy.
The twist for machines
Here is where it turns.
A machine cannot walk through the world. It cannot be bitten by a dog. The first engine, embodied experience, is closed to it. Cameras and 3D worlds pry it open a crack, and that crack is real. Marble matters. But notice what it actually delivers: correlational sensory data, not the participatory experience of having lived. It is still the narrow road.
So if a machine is going to build concepts at all, it leans on the second engine. Analogy. Relational structure. The same tool you use for everything you never touched.
The analogy nobody programmed
And the part nobody designed is that the models walked through that door years ago on their own.
In 2013, Tomas Mikolov’s team at Google found that if you train a model to predict words from their neighbors, on the old linguistics hunch that you know a word by the company it keeps, the geometry that falls out has a property nobody put there. Take the vector for “king,” subtract “man,” add “woman,” and you land almost exactly on “queen.” Paris minus France plus Italy gives you Rome. The model was never taught the relationship “capital of.” It distilled it from raw co-occurrence and stored it as a direction in space you can do arithmetic with. The trick is cleaner for some relations than others and the geometry was later shown to do part of the work, but the core fact holds: relations live inside these models as structure you can compute on. That is not yet Gentner’s full structure-mapping. It is the substrate that analogy runs on, emerging from text alone.
The argument has since grown up. In 2022 Steven Piantadosi and Felix Hill argued, in “Meaning without reference in large language models,” that meaning may not require pointing at the world at all. It can come from conceptual role, the web of relationships a concept holds to every other concept. The obvious objection is that relations among words are still just words. The reply is that this is true of human concepts too. “Justice” has no referent you can point at either. It has only a position in a web of other concepts, and we do not therefore say you fail to understand it.
The grounding that actually matters
This is also where I have to be honest about a paper that does not fully agree with me. In 2023 Dimitri Mollo and Raphaël Millière pulled “grounding” apart into five different things people stack under one word, and argued that the kind that truly matters, referential grounding, is not delivered by relational structure on its own. Fair. But their route to it is the interesting part. It does not run through a body. It runs through interaction, the causal-historical loop of producing language, getting feedback, and adjusting, which is the same shape as reinforcement learning from human feedback. So even the grounding that matters most does not require the first engine. It requires a loop with the world, not a walk through it. That cuts against the embodiment camp from a different direction than mine, and it lands in the same place: the thing a machine cannot have turns out not to be the thing that decides meaning.
Stack it up and Li’s verdict needs one correction. She is right that a language model is ungrounded in the physical world. She is missing that almost nothing you mean is grounded there either. The scaffold a human mind runs on for justice, debt, Tuesday, and nearly everything that matters is relational, and relational structure is the one thing these models have in abundance.
A language model is ungrounded in the physical world. So is almost everything you mean.
What this means for what you build
I run an AI company, so let me land this where it pays rent.
If meaning is analogical and personal, three things change about how you build.
First, stop trying to teach the model THE world. There is no the-world to learn. There is a different world behind every user’s eyes. The job of a serious AI system is not to hold one correct world model. It is to rebuild this person’s world model well enough to land a thought on their side without it decompressing into the wrong thing. The dog-lover and the dog-bitten need different sentences for the same fact.
Second, more grounding is not always the highest lever. The reflex right now is to reach for sensors, video, 3D, embodiment, the first engine. Those are real. But for most of what your system has to mean, the cheaper and stronger lever is better analogy: retrieval that puts the right known thing in front of the model at the right moment, context that supplies the user’s own experience as the basis, explanation that lands a new idea on an old one the user already holds.
You are not grounding the model in the world. You are funding its analogies.
Third, the stack has three layers, not two. Li gives you the renderer that makes a world look right and the simulator that makes it behave true. Both matter. But on top of them sits the oldest machine in human cognition, the one that makes a thing mean something to a particular mind, and it runs on analogy. Most of what we ship today has the first layer, some of the second, and almost none of the third built on purpose. We get the third by accident, as a side effect of scale, and then act surprised when it comes out uneven.
The machine that matters will need all three. The analogy engine is the oldest tool we have and the one we keep underbuilding, because it is invisible. Nobody posts a demo of a model that explained a hard idea by reaching for exactly the right thing you already understood. But that, not a prettier rendered room, is what understanding has always actually been.
So here is the question I will leave for the people building this, the same one I closed the post with.
Is language a description of the world, or a compression of it?
That is not philosophy. It decides what you spend next year on. If language describes the world, you go add more world: sensors, video, space. If language compresses it, you go add more analogy: memory, context, the listener’s own basis. One of those roads is crowded right now. The other is the one your own mind has walked your whole life.
Related reading: Orchestration Is the System picks up where this ends. If meaning lives in the system and not the model, then the system around the model is the product.
Sources
- Fei-Fei Li, “From Words to Worlds: Spatial Intelligence is AI’s Next Frontier,” 2025.
- Stevan Harnad, “The Symbol Grounding Problem,” Physica D, 1990.
- Emily M. Bender and Alexander Koller, “Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data,” ACL 2020.
- George Lakoff and Mark Johnson, “Metaphors We Live By,” 1980.
- Douglas Hofstadter and Emmanuel Sander, “Surfaces and Essences: Analogy as the Fuel and Fire of Thinking,” 2013.
- Dedre Gentner, “Structure-Mapping: A Theoretical Framework for Analogy,” Cognitive Science, 1983.
- Tomas Mikolov et al., “Efficient Estimation of Word Representations in Vector Space,” 2013.
- Steven T. Piantadosi and Felix Hill, “Meaning without reference in large language models,” 2022.
- Dimitri Coelho Mollo and Raphaël Millière, “The Vector Grounding Problem,” 2023.
- Ludwig Wittgenstein, “Philosophical Investigations,” 1953.
- Elizabeth Spelke and Susan Carey, on core knowledge and infant concept acquisition.