📚 Course
Intermediate
~2–3h

Can Machines Think?

Philosophy of Mind, Applied to Today's AI

Long before language models, philosophers built the arguments you need to evaluate them: what a behavioral test can and can't prove, whether symbol manipulation is understanding, and why subjective experience might be a separate question from intelligent behavior entirely. This course walks through those arguments and applies them directly to today's AI.
Intermediate
~2–3 hours (self-paced)
5 Modules

TL;DR:

Turing (1950) proposed a behavioral test to sidestep the unanswerable question “can machines think?” Searle's Chinese Room (1980) argued behavior alone can't prove understanding. Chalmers (1995) separated the “easy problems” of explaining behavior from the “hard problem” of subjective experience. Bender & Koller's Octopus Test (2020) applied the same logic directly to language models trained on text alone. None of these are settled — but together they give you a rigorous way to evaluate the next “this AI truly understands” claim you read.

Who this course is for

This course is for anyone who wants a rigorous, non-mystical way to think about claims like “this AI truly understands” or “this model is becoming conscious” — whether you encounter them in marketing copy, a philosophy seminar, or a late-night argument with a friend. No philosophy background is assumed; every argument is explained from scratch, with its original source.

Unlike most of this Academy, this course is deliberately built on decades-old, canonical sources rather than this year's headlines — the arguments don't expire when the next model ships.

What you'll learn

The Turing Test, Correctly Stated

What Turing's 1950 imitation game actually claims — and the common misreading that it 'proves' machine thought.

Searle's Chinese Room

Why symbol manipulation that passes every behavioral test still might not be understanding.

The Hard Problem

Chalmers' distinction between explaining behavior and explaining subjective experience — and why it may resist behavioral tests entirely.

The Octopus Test

Bender & Koller's modern, ML-specific argument that meaning can't be learned from text form alone.

A Working Framework

A practical way to separate "acts intelligent," "understands," and "is conscious" — three different claims routinely collapsed into one.

What Stays Unanswered

Honest boundaries: which questions this course resolves, and which remain genuinely open.

Module 1 — The Turing Test: what it claims (and doesn't)

In 1950, mathematician Alan Turing published “Computing Machinery and Intelligence” in the journal Mind. He opens by proposing to replace the question “can machines think?” — which he considered too ambiguous to answer — with a concrete alternative: the “imitation game.” A human judge holds text conversations with a human and a machine, without seeing either. If the judge can't reliably tell which is which, the machine passes.

The common misreading is that Turing was proposing a definition of thought: pass the test, and you've proven the machine thinks. Turing's own framing was more modest and more interesting — he was pointing out that convincing behavior is the only evidence we ever have for any mind, human or machine. You can't directly inspect another person's inner experience either; you infer it from what they say and do. The test asks whether machines can clear the same behavioral bar we already use for each other.

Why this still matters for LLMs

Modern language models pass informal, unstructured versions of the Turing Test routinely. Under Turing's own framing, that's evidence worth taking seriously — but it was never meant to settle deeper questions about understanding or experience. Module 2 introduces the argument that directly challenges what passing the test can prove.

Module 2 — Searle's Chinese Room

Thirty years later, philosopher John Searle published “Minds, Brains, and Programs” (1980), introducing the most-discussed thought experiment in the philosophy of AI. Full details are well documented in the Stanford Encyclopedia of Philosophy's Chinese Room entry. The setup: imagine Searle, who speaks no Chinese, locked in a room with a rulebook (in English) that tells him exactly how to manipulate incoming Chinese symbols to produce outgoing Chinese symbols. Follow the rulebook well enough, and the room produces answers indistinguishable from a fluent Chinese speaker — passing a Chinese-language Turing Test.

Searle's point: he still understands zero Chinese. He's doing pure syntax — symbol manipulation according to rules — with no access to semantics, to what the symbols mean. If the whole room (Searle plus rulebook) can pass the test while nobody and nothing inside understands Chinese, then passing a behavioral test cannot, by itself, prove understanding. Searle aimed this specifically at “strong AI” — the claim that running the right program is sufficient for a mind, not at AI capability in general.

What the argument does NOT claim

That machines can never understand anything, or that AI is fake or useless. Searle accepted that brains — which he considered biological machines — produce understanding somehow. His target is narrower: symbol manipulation alone, without more, isn't sufficient.

The strongest counter-reply

The “Systems Reply” (documented in the same SEP entry): maybe Searle-the-person doesn't understand Chinese, but the system (Searle + rulebook + room) does — just as no single neuron in your brain understands English, yet you do. The debate over this reply is still active and unresolved decades later.

Module 3 — The hard problem of consciousness

In 1995, philosopher David Chalmers published “Facing Up to the Problem of Consciousness”, drawing a distinction that reframes the whole debate. The “easy problems” of consciousness — explaining how a system discriminates stimuli, integrates information, reports internal states, or focuses attention — are hard engineering problems, but ones cognitive science can in principle solve by describing mechanisms and functions.

The “hard problem” is different in kind: why is any of that accompanied by subjective experience at all? Why does processing red-wavelength light produce the felt quality of seeing red, rather than happening “in the dark,” with no experience attached? Chalmers argues that fully explaining the mechanism — the easy problems — doesn't automatically explain why there's something it's like to be that mechanism.

Why this caps what any behavioral test can prove

The Turing Test and its descendants are behavioral — they test what a system does. The hard problem, by construction, is about something behavior can't settle: whether there's subjective experience behind the behavior. A system could ace every easy problem — language, reasoning, self-report — while the hard problem stays completely open. This is why “it passed the test” and “it is conscious” are different claims, not two ways of saying the same thing.

Module 4 — Today's LLMs through this lens

In 2020, computational linguists Emily Bender and Alexander Koller published “Climbing towards NLU”, updating the Chinese Room for the language-model era with the “Octopus Test.” Two people, stranded on separate islands, communicate over an undersea cable. An octopus secretly taps the cable — seeing only the signal patterns, never the world those signals describe. Bender and Koller argue the octopus could, in principle, learn to produce plausible-sounding responses purely from patterns in the signal, without ever grasping what any of it refers to.

Their core claim: a system trained only on the form of language (text, as LLMs are) has no built-in way to learn meaning — the connection between words and the world — from form alone. It's a technically specific, empirically grounded descendant of the Chinese Room, aimed directly at whether today's text-only language models can be said to understand in the fullest sense, independent of how fluent their output is.

Where does this leave large language models? Genuinely contested — Bender and Koller's critics argue that grounding can be approximated through the sheer scale and diversity of training data and multimodal input (text plus images, audio, video), while defenders maintain the form/meaning gap persists regardless of scale. That live disagreement, not a settled answer, is the actual state of the field — treat confident claims on either side with the same scrutiny you'd apply to any other unresolved philosophical debate.

Module 5 — A framework for “does it understand” claims

Four modules of philosophy converge on one practical habit: separate three claims that get routinely collapsed into one.

  1. Acts intelligently — produces outputs that look like the product of understanding, reasoning, or skill. Testable today, and current AI clearly clears this bar on many tasks.
  2. Understands — has genuine semantic grasp of what symbols refer to, not just facility manipulating them (the Chinese Room / Octopus Test question). Contested, and possibly untestable by behavior alone.
  3. Is conscious / has subjective experience — there is “something it is like” to be the system (the hard problem). Not resolvable by any test we currently have, for AI or arguably for anyone but yourself.

Most headline claims — “this AI truly understands you,” “this model shows signs of consciousness” — quietly use evidence for claim 1 to assert claim 2 or claim 3. Use the template below to pull these apart the next time you read one.

Mind-Claim Fact-Check Template:
I'm evaluating this claim about an AI system: [paste the claim or headline here]

Walk me through it against this framework:
1. What specific evidence is being cited — a benchmark score, a demo, a self-report from the AI itself, or something else?
2. Is that evidence actually about the system acting intelligently (producing good outputs), or is it being used to support a stronger claim about genuine understanding or subjective experience?
3. If it's a claim about "understanding": does the evidence rule out the Chinese Room / Octopus Test objection (that fluent output can come from form manipulation without meaning)?
4. If it's a claim about "consciousness" or "experience": what test, even in principle, could ever settle this either way — and does the claim acknowledge that limit?

Give me a one-paragraph honest assessment of which of the three claims (acts intelligently / understands / is conscious) this evidence actually supports.

Risks & Responsible Use

Know these before you go further.

Anthropomorphizing Fluent Output as Evidence of Understanding

Humans reliably project understanding, feelings, and intent onto anything that talks back fluently — a bias documented since ELIZA in the 1960s and still active with modern chatbots. Fluent, confident-sounding text is exactly the kind of evidence the Chinese Room and Octopus Test arguments warn can come apart from genuine understanding.

What this means for you

Treat fluency and coherence as evidence the system is useful, not as evidence it understands — verify factual claims and reserve judgment calls for humans regardless of how convincing the output sounds.

Treating "Passes the Turing Test" as Settling the Consciousness Question

Turing's test was explicitly designed to sidestep the question of machine consciousness with a behavioral proxy, not answer it. Chalmers' hard problem shows why a system can ace every behavioral test while whether it has subjective experience stays completely open.

What this means for you

Keep "acts intelligently," "understands," and "is conscious" as three separate claims — passing a behavioral benchmark is evidence for the first, not automatically for the other two.

Using Philosophical Uncertainty to Avoid Real Accountability

Because these questions are genuinely unresolved, it can be tempting to use that uncertainty to dodge concrete responsibility — e.g., "we can't know if it really understands, so we're not liable for what it said." The philosophical debate about understanding does not suspend ordinary accountability for a system's outputs and effects.

What this means for you

Separate the open philosophical question from the closed practical one: regardless of whether a system "truly understands," its deployer is accountable for what it does and says.

Overclaiming Certainty in Either Direction

Both "this AI is clearly just a stochastic parrot with zero understanding" and "this AI clearly shows signs of real understanding or consciousness" are stated with more confidence than the current state of the argument supports — these are live, contested academic debates, not settled science.

What this means for you

Match your stated confidence to the actual state of the philosophical literature — flag genuinely open questions as open, rather than picking a side and stating it as fact.

Test Your Knowledge

Complete this quiz to test your understanding of the Turing Test, the Chinese Room, the hard problem of consciousness, and the Octopus Test.

Loading quiz...

Frequently asked questions

Key Insights: What You've Learned

1

Turing's 1950 test proposed behavior as evidence of thought, not proof of it; Searle's 1980 Chinese Room argued behavior alone can't prove genuine understanding, since symbol manipulation can pass a test with nothing understanding anything.

2

Chalmers' 1995 hard problem separates explaining behavior (the "easy problems") from explaining subjective experience — a gap no behavioral test, including the Turing Test, can close by construction.

3

Three distinct claims — "acts intelligently," "understands," and "is conscious" — get routinely collapsed into one; separating them is the single most useful habit for evaluating any AI mind-claim, including the Octopus Test's specific challenge to language models trained on text form alone.