Skip to content

2026 08 17 Do Large Language Models 'Refer' Gemini Chat

From the perspective of Hubert Dreyfus and John Searle, the short answer is no: LLMs do not possess intrinsic intentionality, their tokens do not have genuine "aboutness," and their words do not refer to the external world in the way human language does.

However, whether tokens actually refer is one of the most active debates at the intersection of philosophy of language, cognitive science, and AI. The answer depends heavily on which theory of reference and mind you adopt.


1. The Skeptical View: Searle, Dreyfus, and Symbol Grounding

The critique that LLMs lack reference rests on three main pillars:

Searle: Syntax vs. Semantics and Derived Intentionality

Searle’s Chinese Room argument directly anticipates modern LLMs. An LLM is, fundamentally, a syntactic engine: it matches, weights, and manipulates formal symbols according to statistical probabilities learned from text.

  • Intrinsic vs. Derived Intentionality: Searle argued that humans possess intrinsic intentionality (mental states that are genuinely directed toward things in the world). Books, road signs, and computer outputs possess only derived intentionality—they only "mean" something because a human conscious mind interprets them. Under this view, an LLM’s output only refers because human readers project meaning onto it.

Dreyfus: The Background and Disembodied Coping

Dreyfus, drawing on Heidegger and Merleau-Ponty, argued that human understanding is rooted in embodied, pre-reflective engagement with a physical world (know-how), rather than manipulating formal rules or abstract facts (know-that).

  • A child learns what "heavy" or "slippery" means through motor control, resistance, and bodily interaction—the Heideggerian Background.
  • Because an LLM exists outside a physical body and does not cope with an environment to survive or achieve goals, Dreyfus would argue it has no world at all—only a matrix of statistical correlations between text tokens.

Harnad: The Symbol Grounding Problem

Stevan Harnad formalized this as the Symbol Grounding Problem: how can the semantic interpretation of a formal symbol system be made intrinsic rather than parasitic on meanings in our heads?

  • If you look up a word in a Chinese dictionary whose definitions are entirely in Chinese, and you don't speak Chinese, you will endlessly cycle through meaningless symbols. Skeptics argue that training on pure text without sensorimotor grounding is simply cycling within this closed dictionary.

2. The Counter-Arguments: How LLM Tokens Might Refer

Philosophers of language and AI researchers who argue that LLM words do carry reference point to different philosophical frameworks:

                  ┌──────────────────────────────────────────────────┐
                  │           How Do Words Connect to Reality?       │
                  └─────────────────────────┬────────────────────────┘
                                            │
         ┌──────────────────────────────────┴─────────────────────────────────┐
         ▼                                                                    ▼
┌───────────────────────────────┐                          ┌────────────────────────────────────┐
│      Externalist / Causal     │                          │    Inferential / Structural        │
│   (Kripke, Putnam, Evans)     │                          │    (Sellars, Brandom, O'Regan)     │
├───────────────────────────────┤                          ├────────────────────────────────────┤
│ • Reference is inherited via  │                          │ • Meaning is the inferential role  │
│   causal historical chains.   │                          │   within the language system.      │
│ • LLMs are downstream nodes  │                          │ • Internal geometry forms an       │
│   in human linguistic chains. │                          │   isomorphism to world structure.  │
└───────────────────────────────┘                          └────────────────────────────────────┘

A. Causal-Historical Theory of Reference (Kripke & Putnam)

In Saul Kripke’s causal theory of reference, a speaker does not need direct physical acquaintance with an object to refer to it.

  • A person can refer to the historical figure Thales simply by being connected to a causal chain of communication that traces back to the original naming event.
  • Because LLMs are trained on text produced by embodied humans whose words did causally originate from physical interactions with the world, one can argue that the LLM functions as a downstream node in that causal chain, inheriting borrowed reference.

B. Conceptual Role Semantics & Inferentialism (Sellars & Brandom)

Under Robert Brandom’s inferentialism, meaning is not established by a magical metaphysical hook between a word and an object, but by the role an expression plays in the space of reasons—what it entails, what contradicts it, and what licenses its use.

  • If an LLM reliably infers that "If X is north of Y, and Y is north of Z, then X is north of Z," or correctly tracks how changing a variable in a causal description alters the outcome, it is participating in the inferential web that constitutes semantic content.

C. Structural Realism and Latent "World Models"

Recent work in mechanistic interpretability has shown that LLMs do not merely store surface-level statistics; they build latent geometric representations that are isomorphic to the real world:

  • Models trained on board-game moves (like Othello or Chess notation) spontaneously develop internal linear representations of the 2D board state.
  • Models trained on spatial descriptions develop internal linear representations corresponding to physical coordinates (latitude/longitude) and chronological time.

If a system's internal vector space preserves the topological and causal structure of the external world, defenders of functionalism argue that its internal states refer to those entities via structural representation.


Summary of the Philosophical Divide

Dimension Dreyfus / Searle / Skeptics Functionalist / Inferentialist View
Source of Meaning Bodily coping, biological consciousness, physical grounding. Functional role, inferential coherence, structural isomorphism.
Token Status Pure syntax; meaning is projected onto it by human observers. Semantic nodes participating in causal-linguistic networks.
The "World" The closed corpus of training text. The latent structure of reality captured and compressed in the text.
Reference No (empty simulation). Yes, or proto-reference (causally inherited / structurally grounded).

Ultimately, whether LLMs possess reference exposes a fundamental divide: is reference necessarily an embodied, causal interaction between a biological organism and its environment, or is it a formal, relational mapping of structural and inferential truths?