SHRDLU: Understanding Language Through Action

Michael BrenndoerferMarch 15, 202510 min read

Part of History of Language AI

In 1968, Terry Winograd's SHRDLU system demonstrated a major approach to natural language understanding by grounding language in a simulated blocks world.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

1968: SHRDLU

In the late 1960s, artificial intelligence researchers were asking whether computers could understand language or merely mimic conversation. Two years after ELIZA showed how far pattern matching could carry that illusion, Terry Winograd's SHRDLU pursued a model-based approach. It connected language to objects and actions in a simulated environment.

SHRDLU treated language comprehension as a problem of modeling the described world and acting within it. Winograd tested this idea in a simulated world of colored blocks, pyramids, and boxes that a virtual robot arm could manipulate. The narrow domain let SHRDLU connect commands and questions to an explicit representation of its environment.

What Does SHRDLU Stand For?

SHRDLU doesn't stand for anything—it's not an acronym. The name comes from the nonsense phrase "SHRDLU ETAOIN" which represents the most frequently used letters in English, traditionally used by Linotype operators and typesetters. Winograd chose this whimsical name to reflect the system's focus on language and text processing.

The choice was both practical and philosophical: just as "SHRDLU ETAOIN" represented the building blocks of written language, Winograd's SHRDLU represented an attempt to build language understanding from fundamental computational blocks.

Understanding Through Action

SHRDLU grounded language in action within what researchers called a "blocks world." This simplified virtual environment contained blocks, pyramids, and boxes of various colors on a flat surface. A simulated robot arm could pick up, move, and stack the objects in response to instructions. Language processing therefore operated against a world whose state the program could inspect and change.

SHRDLU connected several components that earlier dialogue systems had kept separate. It parsed sentences with formal grammar rules, interpreted commands and questions, and maintained a world state that tracked the positions, properties, and relationships of every object. It could then plan and execute requested actions, answer questions about the current arrangement, and recall earlier actions.

Unlike ELIZA's pattern-matching dialogue, SHRDLU used an explicit model of its world. It updated that model after each action and applied rules to spatial relationships and object properties. Within its simulated setting, the system showed how language could be interpreted against a changing context rather than handled as isolated text patterns.

The Blocks World: A Deliberately Narrow Test

Winograd chose the blocks world to make a difficult problem tractable with 1960s computing resources and AI techniques. The domain still required language interpretation and spatial reasoning, but its objects and permitted actions could be enumerated in advance.

In this simplified domain, the vocabulary was finite and unambiguous. Each object had clear, well-defined properties: color, shape, size, and position. There was no confusion about what "the red block" meant, no metaphorical language to interpret, no ambiguity about whether "pyramid" referred to an Egyptian monument or a geometric shape. This clarity allowed the system to focus on the structure of language itself rather than struggling with lexical ambiguity.

The semantics of actions were equally clear and well-defined. Commands like "pick up," "put down," and "stack" had precise, executable meanings. Each action corresponded to a specific sequence of operations the robot arm could perform, and the effects of these actions on the world state were completely predictable. This meant that SHRDLU could verify its understanding through action: if it correctly understood a command, executing that command would produce the expected result in the simulated world.

The blocks world also required spatial reasoning. To understand commands and answer questions, SHRDLU had to interpret relationships such as "on top of," "under," "beside," "taller than," and "inside." Doing so required a structured representation of the three-dimensional arrangement and an account of how actions changed it.

By avoiding open-domain conversation, Winograd could concentrate on sentence structure, reference, and the connection between language and a represented world. The constraint was central to both SHRDLU's success and its later limitations.

A Conversation with SHRDLU

The following exchange, adapted from Winograd's original demonstrations, shows SHRDLU handling ambiguity, resolving references, and reporting its interpretation of a pronoun.

Capabilities in the Conversation

The short exchange exercises several language-processing capabilities in sequence. The important details lie in how SHRDLU responds to ambiguity, a relative clause, and a reference to its current world state.

When the human asks SHRDLU to "Grasp the pyramid," more than one pyramid may satisfy the description. The system detects that ambiguity and asks, "I don't understand which pyramid you mean." This is reference resolution with an explicit clarification step: SHRDLU identifies that the noun phrase does not select a unique object.

The command "Find a block which is taller than the one you are holding and put it into the box" combines memory and comparison with reference resolution and planning. SHRDLU must represent the object it holds, compare that object's height with other blocks, identify a qualifying block, and plan the actions needed to place it in the box. The example shows how a nested relative clause can be translated into a sequence of operations.

SHRDLU also makes its handling of the pronoun "it" explicit. Its response, "By 'it', I assume you mean the block which is taller than the one I am holding," states the selected antecedent before acting. The user therefore has an opportunity to correct the interpretation.

The final question, "What does the box contain?" demonstrates yet another capability: the ability to query the current world state and report it in natural language. SHRDLU must identify the box (using reference resolution again), examine which objects are located inside it, and generate a grammatically correct response that enumerates those objects: "The blue pyramid and the blue block." This requires not just perception and knowledge representation but also natural language generation, the ability to produce appropriate linguistic responses to questions.

Together, these capabilities illustrated a model-based approach to language understanding. SHRDLU built representations, reasoned over them, and connected utterances to its model of the world.

The Limits of Understanding

SHRDLU's limitations proved as instructive as its successes. Its constraints exposed problems that would occupy artificial intelligence and natural language processing researchers for decades.

The most severe limitation was what researchers called domain brittleness. SHRDLU's understanding was entirely dependent on its blocks world. The system couldn't generalize beyond this narrow domain. If you wanted to add a new type of object, say a sphere or a cylinder, you couldn't simply tell SHRDLU about it. Instead, you had to modify the system's core knowledge representation, update its reasoning procedures, and potentially revise its grammar rules. Extending the system to handle a different domain entirely, such as a kitchen environment or a toolbox, would require rebuilding most of it from scratch. SHRDLU's knowledge was hand-coded and domain-specific, not learned or transferable.

This brittleness stemmed from a deeper problem: the knowledge acquisition bottleneck. Every piece of knowledge SHRDLU possessed had to be explicitly programmed by a human expert. The system knew that blocks could be stacked because Winograd had written code to represent that fact. It knew what "taller than" meant because that comparison had been explicitly defined. As researchers tried to expand such systems to handle richer, more realistic domains, they discovered that the amount of knowledge required grew exponentially. Each new sentence pattern demanded new grammar rules. Each new concept required new representations and reasoning procedures. The complexity quickly became unmanageable.

SHRDLU also confronted what philosophers and AI researchers called the frame problem. This is the challenge of determining which aspects of a situation are relevant when reasoning about actions and their consequences. In the simple blocks world, when you pick up a block, most things stay the same: other blocks don't move, colors don't change, the table remains where it is. But specifying all the things that don't change is surprisingly difficult, especially as domains become more complex. In richer environments, the frame problem makes complete symbolic representation intractable.

Related to this was the symbol grounding problem, a question that became central to cognitive science and AI: how do linguistic symbols connect to real-world meaning? SHRDLU's symbols were grounded in its simulated world, but that world was itself a set of symbols. The system manipulated representations of blocks, not actual blocks. It never confronted the challenge of connecting language to raw sensory data, images, or physical interactions. This left open whether its "understanding" amounted to anything beyond sophisticated symbol manipulation.

A Turning Point in the History of AI

Despite its limitations, SHRDLU influenced how researchers approached language understanding and artificial intelligence. Its technical design also fed longer-running debates about representation, grounding, and the nature of machine understanding.

SHRDLU offered an early demonstration of limited-domain language understanding built around an explicit world model. Where ELIZA produced dialogue through pattern matching, SHRDLU connected language to perception, action, and world knowledge. That contrast raised expectations for what natural language processing systems might do.

The system also made grounding central to its account of language understanding: meaning depended on connections between utterances and the world they described. This design anticipated later work on embodied cognition, situated language understanding, and the integration of language with perception and action. Modern systems that combine language models with vision, robotics, or interactive environments revisit the same connection with different methods.

The microworlds methodology that SHRDLU exemplified became a dominant approach in AI research for two decades. Researchers across many subfields adopted the strategy of studying intelligence within simplified, constrained domains where problems were tractable. While this approach had its limitations, it allowed researchers to make concrete progress on specific capabilities while building toward more general solutions. The microworlds tradition directly influenced research in planning, reasoning, knowledge representation, and human-computer interaction.

SHRDLU also foreshadowed work now described as embodied AI, which integrates language with perception and physical interaction. Robots that follow natural-language instructions, assistants that control devices, and multimodal models all connect language to states or actions outside the text itself.

SHRDLU represented both a high point and an early warning for rule-based natural language processing. Careful engineering, explicit knowledge representation, and hand-crafted rules produced an effective system within one domain. The same design was brittle, difficult to scale, and constrained by the knowledge acquisition bottleneck. Those limitations later helped motivate statistical and machine-learning approaches that acquired linguistic patterns from large text corpora.

In the years after SHRDLU, researchers struggled to move from microworlds to less constrained language. Progress required different mathematical frameworks, more data, and far more computational power. SHRDLU nevertheless established a benchmark for limited-domain language understanding that later systems would approach with very different methods.

The Paradox of SHRDLU's Success

SHRDLU's demonstrations created a paradox. The system worked because language was grounded in a world model, yet the simplified domain, hand-coded knowledge, and controlled environment that made this possible also prevented the approach from scaling to open-ended language.

This tension between narrow capability and poor generalization became a recurring pattern in AI research. Systems that worked within fixed constraints often failed when those constraints were relaxed. The difficulty helped push the field toward statistical and machine-learning approaches that could acquire patterns from data rather than requiring every rule to be programmed explicitly.

Quiz: Understanding SHRDLU

Test your knowledge of Terry Winograd's revolutionary language understanding system.

SHRDLU Quiz

Question 1 of 40 of 4 completed
What does 'SHRDLU' stand for?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025shrdluunderstanding, author = {Michael Brenndoerfer}, title = {SHRDLU: Understanding Language Through Action}, year = {2025}, url = {https://mbrenndoerfer.com/writing/history-shrdlu-language-understanding-blocks-world}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). SHRDLU: Understanding Language Through Action. Retrieved from https://mbrenndoerfer.com/writing/history-shrdlu-language-understanding-blocks-world
MLAAcademic
Michael Brenndoerfer. "SHRDLU: Understanding Language Through Action." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/history-shrdlu-language-understanding-blocks-world>.
CHICAGOAcademic
Michael Brenndoerfer. "SHRDLU: Understanding Language Through Action." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/history-shrdlu-language-understanding-blocks-world.
HARVARDAcademic
Michael Brenndoerfer (2025) 'SHRDLU: Understanding Language Through Action'. Available at: https://mbrenndoerfer.com/writing/history-shrdlu-language-understanding-blocks-world (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). SHRDLU: Understanding Language Through Action. https://mbrenndoerfer.com/writing/history-shrdlu-language-understanding-blocks-world

About the author

Continue with the full handbook

This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore History of Language AI
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.