Part of History of Language AI
In 2005, the PropBank project at the University of Pennsylvania added semantic role labels to the Penn Treebank.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
2005: PropBank
In the mid-2000s, researchers at the University of Pennsylvania began adding semantic role labels to the Penn Treebank's syntactic annotations. The resulting Proposition Bank, or PropBank, was developed under the leadership of Martha Palmer, Daniel Gildea, and Paul Kingsbury. Released in 2005, it paired a large semantic annotation resource with an established syntactic treebank, giving statistical systems training data for identifying "who did what to whom" while retaining the Penn Treebank's syntactic structure.
By 2005, the Penn Treebank was a standard resource for training statistical parsers, but syntax alone did not encode the roles that participants played in an event. FrameNet had demonstrated semantic role annotation with named roles such as Buyer and Seller, though its detailed frame-based approach covered a limited vocabulary. PropBank pursued broader coverage with numbered arguments (Arg0, Arg1, Arg2, and so on) rather than named frame elements. This simpler scheme allowed more predicates to be annotated on the existing Penn Treebank trees, placing syntactic and semantic information in the same resource.
PropBank's predicate-argument annotations made it possible to train and evaluate statistical systems for semantic role labeling (SRL). The resulting dataset helped turn SRL into a repeatable benchmark and showed that semantic annotation could be produced at a scale useful for machine learning.
Later semantic resources adopted PropBank's compatibility with existing syntax and its use of frame files for verb-specific roles. They also emphasized annotation consistency. PropBank supplied data for the CoNLL shared tasks and informed later projects such as OntoNotes. Its coverage made statistical evaluation possible across a sizeable portion of the vocabulary.
The Problem
By 2005, statistical NLP systems had made substantial progress in syntactic parsing, but syntax alone could not identify who did what to whom. Consider the sentences "John gave Mary a book" and "Mary received a book from John." Their syntax differs, yet both describe the same transfer of possession. A syntactic parser could identify grammatical relationships without recording that John is the agent, Mary the recipient, and the book the transferred theme. Those roles support tasks such as information extraction, question answering, and machine translation.
The problem extended beyond individual sentences. Different syntactic realizations could express the same semantic roles. In "The company bought the subsidiary," the subject "The company" is the buyer. But in "The subsidiary was bought by the company," the buyer appears in a prepositional phrase, while the object becomes the subject. Systems needed to recognize that despite different syntactic structures, the semantic roles remain the same. Traditional resources couldn't help with this. Syntactic parsers provided tree structures but not semantic roles. WordNet provided word relationships but not event structure. The gap between syntactic analysis and semantic understanding constrained NLP systems' capabilities.
This limitation became increasingly problematic as systems attempted more sophisticated applications. Information extraction systems needed to identify entities and their relationships, but without understanding semantic roles, they struggled to determine whether a person mentioned in a sentence was an agent, a patient, a beneficiary, or some other participant. Question answering systems needed to understand who performed actions, on what objects, for what purposes, but existing resources provided no framework for encoding this information systematically. Machine translation systems needed to preserve semantic relationships across languages, but word-level resources couldn't capture the event structures that verbs and other predicates evoked.
FrameNet, released in 1998, had demonstrated that semantic role annotation was valuable and feasible. FrameNet's frame-based approach with named roles like Buyer, Seller, and Goods provided detailed semantic structures, and it showed that large-scale semantic annotation of corpus data could produce useful resources. However, FrameNet's annotation process was labor-intensive, requiring detailed frame definitions for each lexical unit. By 2005, FrameNet had annotated around 13,000 lexical units, but English vocabulary includes hundreds of thousands of words. The coverage gap limited FrameNet's usefulness for unrestricted text processing, and its frame-based approach required substantial annotation effort for each new lexical unit.
Researchers at the University of Pennsylvania recognized that a different approach could provide broader coverage more efficiently. Rather than creating detailed named frames for each verb, a simpler scheme using numbered arguments could scale more effectively while still capturing the essential semantic information needed for semantic role labeling. This scheme could build directly on the Penn Treebank's existing syntactic annotations, avoiding the need to create a separate resource from scratch. The challenge was designing an annotation scheme that was simple enough to scale broadly but precise enough to capture meaningful semantic distinctions. PropBank addressed this challenge by combining verb-specific frame files with a general numbering scheme for arguments.
The Solution: Numbered Arguments and Frame Files
PropBank addressed this problem by creating a semantic annotation layer that worked directly with the existing Penn Treebank syntactic trees. The solution had two key components: numbered arguments that could scale across different verbs, and verb-specific frame files that defined what each numbered argument represented for particular verbs. This design enabled PropBank to provide broader coverage than FrameNet while maintaining compatibility with the most widely used syntactic resource in computational linguistics.
Numbered Arguments
PropBank's core innovation was using numbered arguments (Arg0, Arg1, Arg2, etc.) rather than named semantic roles. For verbs, these arguments had general semantic interpretations:
- Arg0: Typically the agent or proto-agent (the entity that performs the action)
- Arg1: Typically the patient, theme, or proto-patient (the entity affected by the action)
- Arg2: Typically the beneficiary, recipient, or instrument
- Arg3: Often the start point, benefactive, or attribute
- Arg4: Often the end point
These general interpretations provided consistency across verbs, while verb-specific frame files defined what each numbered argument meant for a particular verb. For "give," the frame file might specify that Arg0 is the giver, Arg1 is what is given, and Arg2 is the recipient. For "buy," Arg0 would be the buyer, Arg1 what is bought, and Arg2 the seller or source. The general numbering supported reuse, while the frame files preserved verb-specific distinctions.
The numbered argument scheme had several advantages. It was simpler to annotate than creating detailed named frames, enabling faster annotation and broader coverage. It was compatible with the Penn Treebank's tree structures, allowing annotators to mark semantic roles directly on existing syntactic nodes. It provided a framework that machine learning systems could learn from, with consistent argument numbering patterns across verbs. And it allowed verb-specific distinctions through frame files, avoiding the one-size-fits-all problem that purely general schemes might face.
Frame Files
Each verb in PropBank was associated with a frame file that defined the possible semantic roles for that verb. These frame files specified what each numbered argument (Arg0, Arg1, Arg2, etc.) represented for that particular verb. This ensured consistency in annotation across different occurrences of the same verb. Frame files also documented the syntactic realizations typical for each argument, such as whether Arg0 typically appears as the subject or whether Arg2 typically appears with a preposition.
Consider the verb "break." Its frame file might specify that Arg0 is the agent that causes the breaking (typically the subject), Arg1 is what gets broken (typically the object), and Arg2 might be the instrument used (typically a prepositional phrase with "with"). The frame file captures both the semantic roles and their typical syntactic realizations. This provided a complete picture of how this verb structures events. When annotators encountered "John broke the window with a hammer," they could consistently mark John as Arg0 (the agent), "the window" as Arg1 (what was broken), and "a hammer" as Arg2 (the instrument).
Frame files also handled verb alternations, where the same verb can appear in different syntactic frames while maintaining similar semantics. The verb "load" can appear as "John loaded the truck with boxes" (where Arg0 is the agent, Arg1 is the location, and Arg2 is what is loaded) or "John loaded boxes onto the truck" (where Arg0 is the agent, Arg1 is what is loaded, and Arg2 is the location). The frame file documents both patterns, showing how the same verb can realize its arguments differently while maintaining consistent semantic roles.
PropBank's frame files differ from FrameNet's frames in important ways. FrameNet creates named frames with descriptive element names like Commerce_buy with elements Buyer, Seller, Goods, and Money. PropBank uses numbered arguments that are general across verbs but specified in verb-specific frame files. This design choice makes PropBank's annotation faster and enables broader coverage, while FrameNet provides more detailed semantic structures. The two resources are complementary: FrameNet offers depth, while PropBank offers breadth.
Annotation Process
The PropBank annotation process involved several systematic steps. First, annotators identified the main predicate of each sentence, typically a verb. Then, they consulted the frame file for that verb to determine what semantic roles were possible. Next, they identified the arguments of the predicate in the Penn Treebank tree and determined which numbered arguments they represented based on the frame file. The annotation also included information about the syntactic realization of each argument, such as whether it was a noun phrase, prepositional phrase, or other syntactic category.
This dual annotation of semantic roles and syntactic realization made PropBank useful for both semantic and syntactic analysis. Systems could use the semantic role labels directly for semantic understanding tasks, or they could study the relationship between syntax and semantics by examining how semantic roles were realized syntactically. This combination proved valuable for research on argument alternations, voice constructions, and other phenomena where syntactic and semantic structures interact.
The annotation scheme was designed for the existing Penn Treebank trees. Semantic roles were marked on the same structures as the syntactic annotations, so researchers could use both levels in one analysis rather than maintain separate corpora.
Coverage and Scale
PropBank's simpler annotation scheme enabled broader coverage than FrameNet. By 2010, PropBank had annotated over 1 million words of text with semantic roles, covering substantially more lexical units than FrameNet's detailed frame annotations. This broader coverage made PropBank more useful for applications that needed to process unrestricted text, where encountering unannotated lexical units was common. The resource became the standard for semantic role labeling tasks precisely because of this coverage advantage.
The coverage advantage came from PropBank's efficient annotation process. Rather than creating detailed frame definitions for each lexical unit, annotators could use the general numbered argument scheme with verb-specific frame files. This process was faster than FrameNet's frame creation, enabling PropBank to annotate more text in less time. The trade-off was less detailed semantic structures than FrameNet provided, but for many applications, the broader coverage was more valuable than the additional detail.
Applications and Impact
PropBank's 2005 release supplied both a role-labeling scheme and annotated training data for statistical systems. Within a few years, the CoNLL shared tasks were using PropBank data to evaluate semantic role labeling systems.
Semantic Role Labeling
The most direct application of PropBank was semantic role labeling (SRL), the task of automatically identifying which phrases in a sentence fill which semantic roles. PropBank provided both a framework for specifying roles and substantial annotated training data, making SRL a practical task for statistical NLP systems. Early SRL systems trained on PropBank data could learn patterns like: when "give" appears as the main verb, the subject typically fills Arg0 (the giver), the direct object typically fills Arg1 (what is given), and a prepositional phrase with "to" typically fills Arg2 (the recipient).
These patterns enabled systems to label semantic roles in new text automatically. The CoNLL-2004 and CoNLL-2005 shared tasks used PropBank data, giving researchers shared training sets, test sets, and evaluation procedures for SRL.
Modern semantic role labeling systems, including neural approaches, continue to use PropBank annotations as training data. SRL remains useful for applications that represent event structure, including information extraction, question answering, and text summarization.
Information Extraction
Information extraction systems benefited substantially from PropBank's structured semantic annotations. Traditional information extraction focused on identifying entities and binary relations, but struggled with events that involved multiple participants with specific roles. PropBank provided a framework for representing these events systematically, enabling systems to extract structured information about who did what to whom, where, when, and why.
Consider extracting information from news articles about corporate events. A sentence like "Acme Corp acquired TechStart Inc for 50 million), and the date (2005). PropBank's annotations enable systems to recognize that "acquired" is the predicate, with the subject as Arg0 (the acquirer), the object as Arg1 (the target), and prepositional phrases providing additional roles. Systems trained on PropBank data could learn these patterns and apply them to extract structured information from unstructured text automatically.
PropBank's compatibility with syntactic trees also let extraction systems combine syntax and semantics. Syntactic information identified argument boundaries, while role labels described what each argument represented. This combination could outperform systems that used only one annotation layer.
Question Answering
Question answering systems used PropBank to represent what a question asks and where its answer might appear. In "Who gave the book to Mary?", the requested answer is Arg0 of a "give" event whose Arg1 is "the book" and Arg2 is "Mary." A system can search for a matching event and return the phrase that fills Arg0.
Structured roles also helped with questions involving several participants. "Who sold what to whom for how much?" requests multiple roles in a "sell" event. A system can represent those roles and match them against similarly structured passages instead of relying on keyword overlap alone.
Machine Translation
Machine translation systems, particularly statistical machine translation systems, used PropBank to preserve semantic roles across languages. Translation systems need to ensure that when they translate a sentence, the semantic relationships between participants remain consistent. If a source sentence has "John gave Mary a book," where John is Arg0 (the giver), Mary is Arg2 (the recipient), and the book is Arg1 (what is given), the translation should preserve these roles even if the target language expresses them differently syntactically.
PropBank provided a language-independent representation of these roles. Translation systems could map source language sentences to PropBank-style semantic structures, then generate target language sentences from those structures. This ensured that roles were preserved even when syntax differed. This approach was particularly useful for translation between languages with different syntactic structures, where direct word alignment might lose semantic information. Systems that preserved semantic roles produced more accurate translations, especially for sentences involving complex event structures.
Integration with Other Resources
PropBank's compatibility with the Penn Treebank supported research that combined syntactic and semantic analysis. Researchers could study how semantic roles were realized in argument alternations, voice constructions, and control structures. The shared annotation made both types of information available to the same models and analyses.
PropBank also influenced other semantic annotation projects. OntoNotes incorporated PropBank-style roles alongside layers for syntax, semantics, and discourse. These combined annotation layers supported analyses that crossed levels of linguistic structure.
Limitations
PropBank's design also imposed practical limits. Some followed from its choices about coverage and role labels; others reflected the difficulty of annotating semantic knowledge at scale.
Coverage Gaps
PropBank's most obvious limitation was its coverage. Even after substantial annotation efforts, PropBank covered only a fraction of the verbs and predicates that appear in real text. Many common verbs lacked frame files, limiting PropBank's usefulness for unrestricted text processing. The annotation process, while more efficient than FrameNet's, still required creating frame files for each verb and annotating sentences, a labor-intensive process that didn't scale easily to cover the full vocabulary.
The coverage problem was exacerbated by the fact that PropBank focused primarily on verbal predicates. While verbs are central to event structure, many sentences involve nominalizations, adjectives, and other predicates that also have semantic roles. A sentence like "The acquisition of TechStart by Acme Corp" involves an event with roles, but "acquisition" is a noun, not a verb. PropBank's focus on verbs limited its coverage of these alternative predicate types, constraining its usefulness for some applications.
Verb-Specific vs. General Roles
PropBank's use of verb-specific frame files created a tension between verb-specific distinctions and general role patterns. Some researchers argued that PropBank's verb-specific approach missed generalizations that could be captured with more general role labels. For example, many verbs have similar argument structures that could be grouped together, but PropBank's frame files treated each verb separately, potentially missing these patterns.
The verb-specific approach also created challenges for handling novel or rare verbs. If a system encountered a verb that didn't have a frame file in PropBank, it couldn't use PropBank's role definitions. While systems could fall back to general interpretations of Arg0, Arg1, etc., these interpretations were less precise than verb-specific definitions. This limitation affected PropBank's usefulness for domains with specialized vocabulary or for processing text with many rare verbs.
Annotation Consistency
PropBank's annotations, while valuable, reflected the inherent subjectivity in semantic annotation. Different annotators sometimes disagreed about which phrases filled which roles, especially for arguments that were less clearly defined or for verbs with multiple possible argument structures. This annotation inconsistency affected the quality of training data derived from PropBank. Machine learning systems trained on inconsistent annotations learned inconsistent patterns, reducing their accuracy.
The annotation process also faced challenges with ambiguity and underspecification in natural language. Consider "John broke the window." This sentence clearly involves a breaking event, but many semantic roles are unspecified: What instrument was used? When did it happen? Why did John break it? PropBank's annotations marked only the roles that were explicitly realized in the sentence, but this meant that many annotations were incomplete. Systems couldn't always distinguish between roles that were unrealized but implied versus roles that simply weren't relevant to the event.
Numbered Arguments vs. Named Roles
PropBank's use of numbered arguments rather than named roles was a design choice that enabled scalability but created limitations. Numbered arguments like Arg0 and Arg1 are less interpretable than named roles like Agent and Patient. Researchers and system developers had to consult frame files to understand what each argument represented, adding complexity to using PropBank annotations.
The numbered argument scheme also created challenges for applications that needed human-readable semantic representations. While Arg0 might consistently represent the agent across many verbs, the label "Arg0" doesn't convey this meaning directly. Named roles like Agent are more interpretable, but PropBank's design prioritized scalability over interpretability. This trade-off limited PropBank's usefulness for some applications that needed more transparent semantic representations.
Domain Specificity
PropBank's annotations were based primarily on general-purpose corpora like the Wall Street Journal texts used in the Penn Treebank. Many applications needed semantic role knowledge in specialized domains like medicine, law, or finance. PropBank's general-purpose annotations didn't always capture domain-specific event structures. A medical frame for diagnosis might require different elements than a general-purpose frame for observation, but creating domain-specific PropBank annotations required new annotation efforts.
As applications moved into specialized domains, they often needed additional annotations or adaptations of PropBank's general-purpose frame files. The original resource did not describe every domain-specific predicate or event structure.
Despite these limitations, PropBank remained widely used because it offered broad semantic annotations aligned with an existing syntactic resource. Its incomplete coverage and occasional annotation disagreements also documented the practical difficulty of building semantic resources at scale.
Legacy: Semantic Role Labeling in Modern NLP
PropBank helped establish semantic role labeling as a standard NLP task by supplying training data and an evaluation framework. Neural methods changed the models used for SRL, but many systems retained PropBank's account of predicates and arguments.
CoNLL Shared Tasks
The CoNLL-2004 and CoNLL-2005 shared tasks both used PropBank data for semantic role labeling. Their common evaluation procedures allowed direct comparison among SRL systems.
The CoNLL shared tasks also influenced how semantic role labeling systems are designed and evaluated. They supplied common metrics, data formats, and test sets, allowing researchers to compare systems and measure progress under the same conditions. PropBank-style annotations remain common in SRL evaluation.
Modern Semantic Role Labeling Systems
Modern semantic role labeling systems, including neural approaches, continue to use PropBank annotations as training data. The models have changed, but supervised SRL still depends on labeled examples of predicates and their arguments.
Recent neural SRL systems use PropBank as training data and as a source of structure for their predictions. Some explicitly model how predicates relate to their arguments and use PropBank frame files to constrain or organize the output. The framework therefore remains compatible with neural methods.
Abstract Meaning Representation
Abstract Meaning Representation (AMR), developed in the 2010s, represents sentence meaning as a directed acyclic graph. Its primitives differ from PropBank's numbered arguments, but its events also contain predicates connected to semantic roles. The two schemes therefore share a predicate-argument view of event structure.
Modern AMR parsers sometimes use PropBank information during parsing. AMR and PropBank differ in scope and notation, but both represent events through predicates and structured roles.
Event Extraction and Knowledge Graphs
Modern knowledge graphs and event extraction systems build directly on PropBank's insights about event structure. Knowledge graphs represent events with predicates and arguments, maintaining structured representations similar to PropBank's predicate-argument structures. When a knowledge graph represents "John gave Mary a book" as an event with roles for agent, recipient, and theme, it's using the same conceptual structure that PropBank formalized, just with different notation.
PropBank's influence is visible in event extraction systems that identify and structure events in text. These systems often use predicate-argument representations, assigning roles by event type rather than syntactic position alone. Identifying who did what to whom, where, and when remains a common NLP task, and PropBank supplied an annotation framework for it.
Neural Language Models and Semantic Roles
Research suggests that large neural language models learn some semantic-role structure during training. Models such as BERT and GPT can predict roles and predicate-argument relations without explicit PropBank-style training. PropBank's explicit annotations provide one way to analyze those learned patterns.
Some researchers combine explicit resources such as PropBank with neural language models. PropBank contributes structured role information, while neural models learn broader patterns from text. The pairing can make event structure explicit without limiting a system to the coverage of a hand-annotated resource.
Integration with Modern NLP Pipelines
PropBank's compatibility with syntactic trees allows modern pipelines to combine levels of analysis. A system can use a syntactic parser to identify argument boundaries and a PropBank-trained model to assign semantic roles, producing a representation that includes both kinds of information.
Modern NLP pipelines may include PropBank-style semantic role labeling alongside syntactic parsing and named entity recognition. Downstream applications can then draw on several levels of linguistic structure even when each component uses a different model.
Research on Syntax-Semantics Interface
PropBank's dual annotation of semantic roles and syntactic realization has enabled substantial research on the syntax-semantics interface. Researchers have used PropBank to study how semantic roles are realized syntactically, examining phenomena like argument alternations, voice constructions, and control structures. This research has advanced understanding of how syntax and semantics interact. This provided insights that inform both linguistic theory and computational applications.
The research enabled by PropBank's annotations has influenced how modern systems model the relationship between syntax and semantics. Some neural parsers now explicitly model semantic roles alongside syntactic structure, creating unified representations that capture both levels of analysis. This integration builds on PropBank's demonstration that syntactic and semantic annotation can work together effectively.
Today, PropBank continues to be maintained and expanded through new annotations and releases. Its data still trains modern SRL systems, while its predicate-argument scheme remains a common way to represent event structure computationally.
Quiz
The following questions review how PropBank added semantic structure to syntactic trees and supported semantic role labeling as a standard NLP task.
PropBank Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore History of Language AIStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!