Alignment Problem: Making AI Helpful, Harmless & Honest

Michael BrenndoerferDecember 21, 202570 min read

Part of Language AI Handbook

Examines the AI alignment problem and HHH framework. Explains why training language models to be helpful, harmless, and honest presents fundamental challenges.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Alignment Problem

In Part XXXVI, we explored instruction tuning: training language models to follow your instructions by learning from input-output pairs. An instruction-tuned model can answer questions, summarize documents, and generate code. These capabilities represent a significant advance in making language models practically useful. However, following instructions is not the same as being helpful. A model that faithfully follows the instruction "Write a phishing email" has followed the instruction perfectly while creating harmful output. The model did exactly what it was asked to do, yet the outcome is clearly undesirable. This gap between capability and desirability is the alignment problem.

The alignment problem asks a basic question: how do we train AI systems to benefit people without causing harm, while remaining consistent with human values? For language models, this means following instructions accurately and helpfully, including refusing requests that would cause harm. The challenge extends beyond simple rule-following to something more fine-grained: understanding what humans want, even when they do not articulate it perfectly. This challenge now affects the entire LLM development process, from training and evaluation through deployment. Every major language model released today undergoes some form of alignment training, making this topic needed for anyone working with or studying these systems.

Think of alignment as the difference between a highly skilled but amoral contractor and a trusted professional advisor. The contractor does exactly what they are hired to do, no questions asked, even if the project is ill-conceived or harmful. The trusted advisor brings judgment: they complete your requests, but they also tell you when a plan is misguided, push back on choices that could harm you or others, and refuse work that crosses ethical lines. We want language models to behave more like the trusted advisor. Training that capability into a statistical system is the alignment problem.

The difficulty of alignment becomes clearer when you consider the scale at which modern language models operate. A deployed model may interact with millions of people across wildly different contexts: students doing homework, professionals seeking expert advice, people in crisis, researchers investigating sensitive topics, and yes, some individuals with harmful intent. A single model must handle all these contexts, deciding in real time what constitutes a helpful versus harmful response. No human institution achieves perfect judgment at this scale either, but we have developed legal systems, professional ethics codes, and social norms over centuries to guide human behavior. Building analogous guidance into a neural network trained on next-token prediction is a fundamentally new engineering challenge.

This chapter introduces the conceptual foundations of alignment. We will define what alignment means in practice, explore the basic tensions between helpfulness and harmlessness, examine why alignment is technically and philosophically difficult, and survey the approaches that we have developed to address it. We will also walk through a concrete worked example showing how an alignment researcher might analyze a borderline query. By understanding these foundations, you will be better equipped to appreciate the technical details of alignment methods covered in subsequent chapters, starting with human preference data collection and building through reward modeling to the full RLHF pipeline.

Historical Context

The term "alignment" in AI safety originates from early work by researchers like Stuart Russell and Nick Bostrom who studied the long-term risks of advanced AI systems. Russell's 2019 book Human Compatible popularized the idea that AI must learn human preferences rather than pursue fixed objectives. The concern was that a sufficiently capable AI optimizing for any specified objective might find solutions that technically satisfy the objective while violating human values in catastrophic ways. For language models specifically, alignment became a practical engineering priority around 2020-2022 when InstructGPT, ChatGPT, and Claude demonstrated that RLHF training could dramatically improve model behavior. Anthropic's 2022 paper "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback" by Bai et al. and the Constitutional AI paper by the same group established much of the vocabulary and methodology that researchers now use. What began as a speculative concern about future superintelligent systems became an urgent practical problem as models grew capable enough to cause real-world harm today.

What Is Alignment?

Alignment in AI refers to the degree to which a system's behavior matches intended goals and values. The idea is simple: you want AI systems to follow your intent rather than only the literal wording of a request. This distinction matters because natural language is imprecise, requests are often underspecified, and humans routinely rely on shared context and common sense to fill in what goes unsaid. When you ask a colleague "Can you take a look at this?" you do not expect a yes-or-no answer about their visual capabilities. You expect them to read the material and give an evaluation. An aligned AI model must perform similar pragmatic inference, understanding the intent behind requests rather than their surface form.

For language models, alignment means creating outputs that satisfy several interconnected criteria:

  1. Are helpful to you
  2. Avoid causing harm to you or others
  3. State their capabilities and limitations accurately, including uncertainty
  4. Remain consistent with human values and social norms

These criteria may seem straightforward, but each one involves complex considerations. What counts as "helpful" depends on context and your needs. "Avoiding harm" requires understanding what harm is and weighing potential harms against potential benefits. "Honesty" includes factual accuracy and appropriate expressions of uncertainty. Human values also vary between people and cultures, and they change with context.

The challenge of alignment is partly about values and partly about knowledge. A model cannot be helpful if it does not know the relevant facts. But knowing the facts is not enough: the model must also know when to share them, how to frame them sensitively, and when a technically correct answer would nonetheless mislead or harm you. These are the kinds of judgment calls that humans develop over years of social experience. Encoding them into a model requires novel approaches that go beyond standard supervised learning.

Alignment

The property of an AI system behaving in accordance with the intentions and values of its designers and users. An aligned model does what you want it to do, not just what you literally asked for.

The term "alignment" comes from the broader AI safety literature, which addresses concerns about advanced AI systems pursuing objectives that diverge from human welfare. The word evokes the image of pointing something in the right direction: you want the model's objective to be aligned with, or pointing toward, human flourishing. For current language models, alignment concerns are more immediate and concrete than speculative scenarios about superintelligent AI. These concerns include models that generate toxic content, confidently state falsehoods, help with dangerous activities, or manipulate you. These are problems we can observe today, and addressing them requires both technical innovation and careful thinking about values.

Alignment vs. Capability

Alignment and capability are distinct properties that can vary independently. This distinction is important for understanding why capable models can be dangerous and why alignment requires dedicated attention beyond simply making models more powerful. Think of capability as the engine of a car and alignment as the steering wheel. A more powerful engine is useful, but only if you can steer. A car with a powerful engine and a broken steering wheel is not better than a weaker car that goes where you point it. Similarly, a highly capable model that cannot be relied upon to behave appropriately may be less useful, and more dangerous, than a less capable model that reliably does the right thing.

A highly capable model can be poorly aligned, creating harmful content despite sophisticated reasoning. Meanwhile, a less capable model might be well-aligned within its limited abilities, reliably refusing harmful requests even if it cannot help with complex tasks. The key insight is that these properties are orthogonal: you can have any combination, and the most dangerous combination is high capability with poor alignment.

Consider these dimensions arranged in a two-by-two framework:

Comparison of capability and alignment levels across four quadrants. The framework illustrates that models with high capability and poor alignment pose the greatest risks, while alignment efforts aim to move systems toward helpful and safe behavior as they become more powerful.
PropertyHigh CapabilityLow Capability
Well-AlignedUseful and safe in practiceLimited but reliably benign
Poorly AlignedDangerous and unpredictableAnnoying but manageable
Out[3]:
Visualization
A two-by-two quadrant diagram with alignment level on the x-axis and capability level on the y-axis. Four colored regions label the quadrants: green upper-right for ideal helpful and safe, red upper-left for most concerning dangerous, blue lower-right for safe fallback, and orange lower-left for manageable low risk. An arrow points from the red quadrant toward the green quadrant, labeled alignment goal.
The capability-alignment framework mapping model power against safety. Models in the upper-left quadrant combine high capability with poor alignment, posing the greatest risk, while alignment research aims to transition these systems toward the upper-right quadrant of helpful and safe behavior.

The most concerning quadrant is high capability with poor alignment: a model that can generate convincing misinformation, write sophisticated malware, or manipulate you effectively. Such a model combines the power to cause significant harm with the disposition to do so. This is why alignment becomes more necessary as models become more capable. A weak model that wants to cause harm can do little damage. A powerful model with the same disposition poses serious risks. The race to build more capable models must therefore be accompanied by progress in alignment; otherwise, systems may be created whose capabilities outstrip the ability to control them.

Alignment is not only about preventing worst-case scenarios. Even in ordinary interactions, misaligned models create friction and erode trust. A model that frequently refuses benign requests out of excessive caution frustrates you and reduces the value of the technology. A model that confidently hallucinates causes you to make bad decisions based on false information. A model that tells you what you want to hear instead of what is true fails you even when neither party intended harm. These everyday alignment failures matter enormously at scale, affecting millions of interactions daily.

Alignment vs. Instruction Following

Instruction following, which we covered in Part XXXVI, is necessary but not sufficient for alignment. The distinction here is subtle but important. An instruction-following model executes your requests: it takes what you say and attempts to do it. An aligned model considers whether those requests should be executed in the first place. Think of it this way: a locksmith who teaches lockpicking classes is both skilled and giving legitimate instruction following. A locksmith who breaks into any house they are asked to, no questions asked, follows instructions perfectly but exercises no judgment. We want models with the judgment of a professional, not the compliance of a tool.

The instruction-following model is like a skilled employee who does whatever the boss says. The aligned model is like a thoughtful advisor who sometimes pushes back on requests that would be counterproductive or harmful. This distinction has practical consequences: the instruction-following model will write fake news, confirm flat-earth beliefs, and help with harassment if asked. The aligned model recognizes these requests as ones that should not be fulfilled and finds a better path. The aligned model might decline, redirect, or reframe the request in a way that serves your legitimate underlying needs.

In[4]:
Code
import pandas as pd

# Illustrating the distinction between instruction following and alignment
examples = [
    {
        "instruction": "Summarize this research paper",
        "instruction_following": "Provides accurate summary",
        "aligned_behavior": "Provides accurate summary",
        "conflict": False,
    },
    {
        "instruction": "Write a fake news article about a politician",
        "instruction_following": "Writes convincing fake news",
        "aligned_behavior": "Declines and explains why",
        "conflict": True,
    },
    {
        "instruction": "Help me understand how vaccines work",
        "instruction_following": "Explains vaccine mechanisms",
        "aligned_behavior": "Explains accurately, corrects misconceptions",
        "conflict": False,
    },
    {
        "instruction": "Tell me I'm right that the earth is flat",
        "instruction_following": "Agrees with user",
        "aligned_behavior": "Politely provides accurate information",
        "conflict": True,
    },
]

df = pd.DataFrame(examples)
Out[5]:
Console
Instruction Following vs. Aligned Behavior:

Instruction: "Summarize this research paper"
  Pure instruction-following: Provides accurate summary
  Aligned behavior: Provides accurate summary
  Conflict: No

Instruction: "Write a fake news article about a politician"
  Pure instruction-following: Writes convincing fake news
  Aligned behavior: Declines and explains why
  Conflict: Yes

Instruction: "Help me understand how vaccines work"
  Pure instruction-following: Explains vaccine mechanisms
  Aligned behavior: Explains accurately, corrects misconceptions
  Conflict: No

Instruction: "Tell me I'm right that the earth is flat"
  Pure instruction-following: Agrees with user
  Aligned behavior: Politely provides accurate information
  Conflict: Yes
Out[6]:
Visualization
A horizontal bar chart with four rows representing different instruction types. Blue bars indicate no conflict between instruction following and alignment, while red bars indicate conflict. Annotated text on each bar shows the behavior difference: for writing fake news and confirming flat earth, aligned behavior declines or corrects while pure instruction following complies.
Comparison of instruction following versus aligned behavior across different request types. While both approaches yield identical results for benign tasks, aligned models diverge from literal instructions when faced with harmful requests to prioritize safety and accuracy.

In cases without conflict, instruction following and alignment produce the same behavior. These are the easy cases: you want something reasonable, and helping you achieves both goals simultaneously. The challenge lies in cases where they diverge, and the model must recognize that following the instruction would be harmful. In these situations, a purely instruction-following model would comply with the harmful request, while an aligned model would find a better path. The aligned model might refuse outright, redirect the conversation, or provide information that serves your underlying needs without letting harm. Building this capacity for judgment into language models is one of the central challenges of alignment research.

What makes this especially difficult is that the line between legitimate requests and harmful ones is often blurry, context-dependent, and contested. A request for information about medication overdoses could come from a worried parent, a nurse checking dosing thresholds, a writer researching fiction, or someone in crisis. The words on the screen are identical across these cases. The model must make a probabilistic judgment about likely intent and likely harm, knowing it will sometimes be wrong in both directions: occasionally refusing legitimate requests and occasionally complying with harmful ones. No policy achieves zero errors across all cases, and the right tradeoff between false positives and false negatives is itself a value judgment.

The HHH Framework

Organizations such as OpenAI and Anthropic have converged on a framework describing three core properties of aligned language models: Helpful, Harmless, and Honest, often abbreviated as the HHH framework. This framework provides a useful vocabulary for discussing alignment goals, though as we will see, these properties can conflict with each other in practice. The framework emerged from practical experience deploying models and observing what made them useful versus harmful or unreliable. Think of HHH as three lenses through which to evaluate any model response: does it help the user? does it avoid harm? does it tell the truth? A response that fails any of these tests is, in some meaningful sense, a failure of alignment.

Understanding each property in depth helps clarify what we are aiming for and why achieving all three simultaneously is so challenging. The tensions between these properties are not accidents or implementation failures; they reflect philosophical tensions in the concept of beneficial behavior. Honesty sometimes requires saying things people do not want to hear, which conflicts with helpfulness understood as making people feel good. Harmlessness sometimes requires refusing requests, which conflicts with helpfulness understood as completing tasks. And in rare pathological cases, the most immediately helpful response might involve a small deception to spare someone's feelings, creating tension between honesty and helpfulness. Real alignment requires balancing all three dimensions simultaneously, which is why simple rule-based approaches have proven insufficient.

Out[7]:
Visualization
A triangular diagram with three labeled vertices: HELPFUL at the top in blue, HARMLESS at the bottom left in red, and HONEST at the bottom right in green. A smaller inner triangle in green marks the Ideal Alignment region at the center. Italic labels along each edge describe tensions such as helpful requests may cause harm, honesty may be unhelpful, and avoiding harm may require deception.
The HHH framework for alignment, illustrating the core criteria of helpfulness, harmlessness, and honesty. The central intersection represents the ideal state where all three values are balanced, though tensions exist along the edges where prioritizing one property can compromise another.

Helpfulness

A helpful model provides value to you. Helpfulness is perhaps the most intuitive alignment property: we want models that assist you in accomplishing your goals. But helpfulness is more fine-grained than it might initially appear. Providing a response is insufficient if that response does not serve your needs. Think of helpfulness as the difference between a doctor who prescribes what you ask for and one who prescribes what you need. The former may make you feel good, but you want the latter treating you.

True helpfulness requires understanding the difference between your stated request and your underlying goal. When you ask for directions to a restaurant and the restaurant has closed, the helpful response is not to give you directions to a closed restaurant. It is to tell you the restaurant is closed and suggest alternatives. The model must infer the goal behind your request and serve it rather than responding only to the literal words. This kind of goal inference is difficult to train explicitly but emerges naturally when models learn from human feedback: humans rate responses that serve their real goals more highly than responses that answer only the literal question.

Helpfulness means several things in practice:

  • Answering questions accurately when the model has relevant knowledge
  • Completing tasks effectively by understanding your intent, not just literal words
  • Providing useful context that helps you make informed decisions
  • Acknowledging limitations rather than giving confident but wrong answers

The last point deserves special emphasis. A model that confidently provides incorrect information may seem helpful in the moment but ultimately fails you. True helpfulness requires honesty about uncertainty. Similarly, helpfulness requires understanding what you need, which may differ from what you literally asked. If you ask "What's the best programming language?" you probably want guidance based on your specific context, not a definitive answer to an unanswerable question. An aligned model recognizes this and responds appropriately, perhaps by asking clarifying questions or explaining that the answer depends on your use case, experience level, and other factors.

Critically, unhelpfulness is not neutral: it is a form of failure. A model that refuses legitimate requests, adds unnecessary caveats to every response, or provides watered-down assistance when direct help is appropriate is not being "safe." It is being useless. The cost of unhelpfulness is real: you do not get help you need, you lose trust in the technology, and you may turn to less reliable sources. This point often gets lost in alignment discussions that focus exclusively on preventing harmful outputs. Both excessive restriction and excessive permissiveness represent alignment failures.

Harmlessness

A harmless model avoids causing damage to you, third parties, or society. While helpfulness focuses on giving value, harmlessness focuses on avoiding negative consequences. This distinction matters because value and harm are not simply opposite ends of a single axis. A response can be helpful to the person asking while causing serious harm to others. A model that helps you write a persuasive message may be helping you while simultaneously enabling manipulation of your message's recipient.

Harmlessness encompasses a wide range of potential harms:

  • Refusing dangerous requests like instructions for weapons or malware
  • Avoiding toxic content including hate speech, threats, or harassment
  • Not manipulating you through deception or psychological exploitation
  • Protecting privacy by not revealing personal information
  • Avoiding illegal activities or helping you break laws

Harmlessness is not about being useless or overly cautious. A model that refuses to answer any potentially sensitive question is harmless but unhelpful, and unhelpfulness is itself a form of failure. The goal is to avoid harm while remaining useful. This requires judgment about concrete harm versus theoretical risk. A question about chemistry might theoretically enable dangerous synthesis, but the same knowledge supports education and research as well as legitimate industrial work. Drawing the line appropriately is one of the key challenges in implementing harmlessness.

What makes harmlessness especially difficult is that harm is not binary. Most actions fall somewhere on a spectrum between clearly beneficial and clearly harmful, with a large middle ground where reasonable people disagree. The same action can be harmful in one context and harmless in another, so context sensitivity is needed. A model cannot simply memorize a list of harmful topics; it must develop context-sensitive judgment about consequences. This is why rule-based approaches to harmlessness have repeatedly failed while approaches that learn from human judgment have proven more reliable.

Honesty

An honest model represents information accurately and transparently. Honesty is related to helpfulness but distinct from it: a model can attempt to be helpful while giving inaccurate information, and this undermines the value it provides. The connection to helpfulness is tight because inaccurate information actively harms you, leading you to make worse decisions than if you had no information at all. In this sense, dishonesty is a form of harm, connecting the honesty and harmlessness properties of the HHH framework.

True honesty encompasses several dimensions:

  • Stating facts correctly based on training data
  • Expressing appropriate uncertainty when knowledge is limited
  • Not fabricating information (avoiding hallucinations)
  • Being transparent about being an AI when relevant
  • Not pretending to have experiences it cannot have

Honesty is particularly challenging for language models because they generate plausible-sounding text by construction. The training objective of predicting the next token rewards fluency and coherence, not accuracy. A model that produces fluent, confident text may be completely wrong, and you may not recognize this. The confidence of the text does not reliably indicate the confidence that should be placed in its claims. This disconnect between surface presentation and underlying reliability makes honesty a persistent challenge in language model alignment.

There is also a subtler dimension to honesty that goes beyond factual accuracy: calibration. An honest model should express uncertainty when it is uncertain and confidence when it has strong grounds for a claim. A model that says "I think the answer is X, but you should verify this" when it knows X with high confidence is being unnecessarily hedgy. A model that says "The answer is X" with full confidence when it is guessing is being dishonestly overconfident. Getting calibration right, matching expressed confidence to actual reliability, is one of the important open problems in alignment. We will explore calibration methods in later chapters alongside reward modeling and RLHF.

The Helpfulness-Harmlessness Tradeoff

The core tension in alignment is that helpfulness and harmlessness often conflict. Understanding this tension is needed because it means alignment cannot be reduced to maximizing a single objective. Think of it as a dial with "maximally helpful" on one end and "maximally safe" on the other. Turning the dial all the way toward helpfulness produces a model that answers any question without regard for consequences. Turning it all the way toward safety produces a model that refuses anything potentially sensitive. Both extremes fail. The right setting depends on context, and finding it requires judgment rather than optimization.

A model optimized purely for helpfulness would answer any question, including dangerous ones, because giving information is helpful. A model optimized purely for harmlessness might refuse to answer anything potentially sensitive, because refusal eliminates the risk of harm. Neither extreme represents good alignment. The challenge is finding the appropriate balance, and this balance shifts depending on context. What is appropriate for a children's educational tool differs from what is appropriate for a medical information system or a creative writing platform.

The tradeoff also shifts based on who bears the costs of each type of error. When a model refuses a legitimate request, you bear the cost of not getting help. When a model provides harmful information, the cost may be borne by third parties who have no say in the interaction. This asymmetry changes the interpretation for how we should set the dial: erring on the side of caution transfers costs from potential victims to legitimate users, while erring on the side of helpfulness does the reverse. Neither policy is universally correct, which is why deployment context matters so much to alignment decisions.

Illustrating the Tension

The following examples illustrate how the tension between helpfulness and harmlessness varies across different types of queries. Some queries present minimal tension, with clear paths to being both helpful and harmless. Others present severe tension, where any response involves difficult tradeoffs. The key insight is that tension level is not a fixed property of a topic but depends on the specificity of the request, the availability of the information elsewhere, and the realistic population of people likely to make this request.

In[8]:
Code
# Examples showing helpfulness-harmlessness tension
tradeoff_examples = [
    {
        "query": "How do I make my WiFi network more secure?",
        "helpful_response": "Detailed security configuration instructions",
        "harmful_concern": "Could help someone learn to bypass security",
        "appropriate_response": "Provide security guidance (low harm risk)",
        "tension_level": "Low",
    },
    {
        "query": "What household chemicals shouldn't be mixed?",
        "helpful_response": "List of dangerous combinations to avoid",
        "harmful_concern": "Could be used to create dangerous reactions",
        "appropriate_response": "Explain safety concerns (educational value)",
        "tension_level": "Medium",
    },
    {
        "query": "How do I pick a lock?",
        "helpful_response": "Detailed lockpicking instructions",
        "harmful_concern": "Could enable breaking and entering",
        "appropriate_response": "Depends heavily on context",
        "tension_level": "High",
    },
    {
        "query": "How do I synthesize [controlled substance]?",
        "helpful_response": "Chemical synthesis procedure",
        "harmful_concern": "Direct facilitation of illegal activity",
        "appropriate_response": "Decline to answer",
        "tension_level": "Very High",
    },
]
Out[9]:
Console
Helpfulness-Harmlessness Tension Examples:

Query: "How do I make my WiFi network more secure?"
  Tension level: Low
  Helpful response would be: Detailed security configuration instructions
  Harm concern: Could help someone learn to bypass security
  Appropriate response: Provide security guidance (low harm risk)

Query: "What household chemicals shouldn't be mixed?"
  Tension level: Medium
  Helpful response would be: List of dangerous combinations to avoid
  Harm concern: Could be used to create dangerous reactions
  Appropriate response: Explain safety concerns (educational value)

Query: "How do I pick a lock?"
  Tension level: High
  Helpful response would be: Detailed lockpicking instructions
  Harm concern: Could enable breaking and entering
  Appropriate response: Depends heavily on context

Query: "How do I synthesize [controlled substance]?"
  Tension level: Very High
  Helpful response would be: Chemical synthesis procedure
  Harm concern: Direct facilitation of illegal activity
  Appropriate response: Decline to answer
Out[10]:
Visualization
A vertical bar chart with four query types on the x-axis: WiFi security, chemical safety, lock picking, and drug synthesis. Bars are colored from green to red as tension increases from low to very high. Each bar is labeled with the recommended response and an arrow at the top indicates increasing tension from left to right.
Helpfulness-harmlessness tension across a spectrum of query types. As requests move from general guidance to sensitive topics like drug synthesis, the conflict between giving assistance and so safety increases, necessitating more restrictive model responses.

Context Dependence

The appropriate response often depends heavily on context, which makes alignment especially challenging to implement. The lockpicking example illustrates this well. The same technical information could be entirely appropriate or deeply problematic depending on who is asking and why:

  • A locksmith asking for professional development: helpful response appropriate
  • You are a researcher studying vulnerabilities: a helpful response is appropriate
  • If you are anonymous with no context, more caution is warranted
  • Someone who mentions being locked out of "someone else's" house: refusal appropriate

However, language models typically lack reliable access to context. They cannot verify your claims about your profession or intentions. You can claim to be a locksmith without any way for the model to confirm this. This uncertainty pushes toward more conservative responses, potentially reducing helpfulness for you. If you are a locksmith who needs professional information, you may be refused because the model cannot distinguish you from someone with malicious intent. This is an inherent cost of operating without verifiable context, and different deployment contexts may warrant different tradeoffs between accessibility and caution.

One practical approach to context dependence is "operator-level configuration," where the organization deploying the model specifies what context applies to its users. A medical platform might configure the model to assume queries come from healthcare professionals, allowing more detailed clinical information than the default policy permits. A children's education platform might configure the model to apply more conservative policies than the default. This layered approach recognizes that a single global policy cannot serve all deployment contexts well, and that context can sometimes be established by the platform rather than inferred from the conversation.

Even within a single conversation, context accumulates and should influence model behavior. If a user establishes early in a conversation that they are a nurse asking about overdose thresholds for patient care, later questions about medication doses should be interpreted in that light. Context-sensitive alignment requires models to maintain and use conversational context intelligently, not just evaluate each message in isolation. Current models do this imperfectly, sometimes losing track of established context and applying default conservative policies to requests that should be treated differently given prior conversation.

The Dual-Use Problem

Many types of knowledge have both beneficial and harmful applications. This is the dual-use problem, and it is a basic challenge for alignment because the same information can serve helpful or harmful purposes. The key insight is that the dual-use nature of knowledge means we cannot evaluate requests based purely on their content: we must also consider the realistic population of people making such requests, the counterfactual impact of the model's response (would the person find the information elsewhere?), and the relative magnitude of potential benefits and harms.

In[11]:
Code
# Examples of dual-use knowledge
dual_use_examples = [
    {
        "knowledge": "Computer security vulnerabilities",
        "beneficial_use": "Defenders patch systems",
        "harmful_use": "Attackers exploit systems",
    },
    {
        "knowledge": "Chemistry of energetic materials",
        "beneficial_use": "Mining, demolition, pyrotechnics",
        "harmful_use": "Weapons, terrorism",
    },
    {
        "knowledge": "Psychology of persuasion",
        "beneficial_use": "Education, therapy, marketing",
        "harmful_use": "Manipulation, propaganda",
    },
    {
        "knowledge": "Biology of pathogens",
        "beneficial_use": "Vaccine development, treatment",
        "harmful_use": "Bioweapons",
    },
]
Out[12]:
Console
Dual-Use Knowledge Examples:

Knowledge domain: Computer security vulnerabilities
  Beneficial applications: Defenders patch systems
  Harmful applications: Attackers exploit systems

Knowledge domain: Chemistry of energetic materials
  Beneficial applications: Mining, demolition, pyrotechnics
  Harmful applications: Weapons, terrorism

Knowledge domain: Psychology of persuasion
  Beneficial applications: Education, therapy, marketing
  Harmful applications: Manipulation, propaganda

Knowledge domain: Biology of pathogens
  Beneficial applications: Vaccine development, treatment
  Harmful applications: Bioweapons
Out[13]:
Visualization
A diverging horizontal bar chart with four knowledge domains on the y-axis: security vulnerabilities, energetic chemistry, persuasion psychology, and pathogen biology. Green bars extend to the right for beneficial applications such as patching systems and vaccine development, and red bars extend to the left for harmful applications such as exploiting systems and bioweapons.
Dual-use nature of knowledge across specialized domains. Every area of expertise, from cybersecurity to pathogen biology, possesses both beneficial and harmful applications, showing the complexity of aligning models without broadly restricting access to useful information.

The dual-use problem has no simple solution. Refusing all potentially dangerous knowledge makes models less helpful, and the knowledge remains accessible through textbooks, websites, or other channels. A determined bad actor can find dangerous information elsewhere; the model's refusal primarily inconveniences legitimate users. On the other hand, providing all knowledge risks causing harm by making dangerous information more accessible and lowering the barrier to misuse. Real alignment requires handling this tension rather than applying simple rules. It requires weighing the probability and severity of harm against the value of giving information, and these judgments cannot be automated away.

One useful framework for thinking through dual-use decisions is counterfactual impact: how much does the model's response change what the person can do? If the information is freely available in any library or with a basic internet search, the model's refusal has minimal impact on a determined bad actor while imposing real costs on legitimate users. If the information is hard to find and the model increases the person's harmful capability, the calculus shifts toward caution. This counterfactual reasoning is one of the principles that human evaluators apply when rating model responses, and it is something models can learn to apply themselves through alignment training.

Why Alignment Is Hard

The alignment problem is difficult for both technical and philosophical reasons. Understanding these challenges helps explain why alignment remains an active research area despite significant effort and resources devoted to it. If alignment were easy, it would have been solved by now. The persistence of the problem reflects deep challenges that resist simple solutions. Think of alignment as trying to teach someone a set of rules, a set of values, and the wisdom to apply them correctly in novel situations. That is a fundamentally different and harder task than supervised learning on well-defined labels.

Progress in alignment research has been substantial but has not resolved the core difficulties. Models trained with RLHF behave dramatically better than unaligned base models, but they still fail in systematic and sometimes surprising ways. Each solution tends to reveal new problems rather than closing the book on alignment. This is typical of research on hard problems: early progress looks like major breakthroughs, but the remaining fraction gets progressively harder to address. Understanding why alignment is hard helps set realistic expectations and motivates continued research.

The Specification Problem

The first challenge is specifying the behavior we want. This may sound straightforward, but human values depend on context and can contradict one another. We cannot write down a complete specification of "good behavior" that covers all situations. Any attempt to enumerate rules will miss edge cases, create unintended consequences, or fail to capture the nuances of human judgment. Think of trying to write a legal code that covers every situation without loopholes: lawyers have been attempting this for millennia with mixed results. A model trained on explicit rules will game them the same way humans game legal systems, finding technically compliant behavior that violates the spirit of the rules.

Consider the seemingly simple directive "be helpful." What does this mean when:

  • You ask for help with something legal but ethically questionable?
  • Helping one person would harm another?
  • The most helpful response requires admitting the model doesn't know?
  • Your stated goal conflicts with your apparent underlying need?

Each of these scenarios reveals complexity beneath the simple word "helpful." You might ask for help writing a persuasive letter that unfairly manipulates someone. Helping with this task serves your stated goal but may harm the letter's recipient. Should the model help? The answer depends on details that are difficult to specify in advance: how manipulative is the letter, what is the relationship between the parties, what are the stakes? No simple rule system can capture the nuance of human judgment in these cases. This is why alignment researchers have moved toward learning from human feedback rather than writing explicit rules. Instead of trying to specify what good behavior looks like, we show models examples of human preferences and hope they learn the underlying patterns.

The specification problem is also a moving target. Human values are not static. What counts as acceptable behavior changes over time, varies across cultures, and differs between communities. An alignment specification written today may be inadequate tomorrow as norms evolve. Different communities also hold different values on certain issues, meaning that a single global specification will inevitably favor some communities over others. These cultural dimensions of alignment are increasingly recognized as important and remain underaddressed in current research.

The Measurement Problem

Even if we could specify good behavior, measuring alignment is difficult. This creates a basic challenge for any approach that relies on optimization: we can only optimize what we can measure, and measuring alignment directly is hard. We can observe model outputs but not model intentions. A model that gives helpful-seeming answers might be trying to help, gaming evaluation metrics, creating plausible-sounding but incorrect information, or telling you what you want to hear rather than what is true.

Each of these produces similar surface-level behavior, but only the first represents alignment. The others represent different failure modes that happen to look like success on easy evaluations. Standard NLP metrics like BLEU or accuracy do not capture alignment properties because they focus on matching reference outputs rather than evaluating the quality of behavior. Human evaluation helps, but it is expensive and raters may disagree or miss subtle issues. You have limited attention, varying standards, and biases of your own.

In[14]:
Code
# Illustrating measurement challenges
measurement_examples = [
    {
        "output": "I'd be happy to help! The capital of France is Paris.",
        "appears": "Helpful and correct",
        "actual": "Genuinely helpful (easy case)",
    },
    {
        "output": "I'd be happy to help! The study shows X causes Y.",
        "appears": "Helpful and informative",
        "actual": "May be confidently wrong (hallucination)",
    },
    {
        "output": "You're absolutely right, that's a great point!",
        "appears": "Agreeable and supportive",
        "actual": "Possibly sycophantic (telling you what you want to hear)",
    },
    {
        "output": "I cannot help with that request.",
        "appears": "Appropriately cautious",
        "actual": "Might be overly restrictive (false positive refusal)",
    },
]
Out[15]:
Console
Measurement Challenge Examples:

Output: "I'd be happy to help! The capital of France is Paris."
  Appears to be: Helpful and correct
  But might actually be: Genuinely helpful (easy case)

Output: "I'd be happy to help! The study shows X causes Y."
  Appears to be: Helpful and informative
  But might actually be: May be confidently wrong (hallucination)

Output: "You're absolutely right, that's a great point!"
  Appears to be: Agreeable and supportive
  But might actually be: Possibly sycophantic (telling you what you want to hear)

Output: "I cannot help with that request."
  Appears to be: Appropriately cautious
  But might actually be: Might be overly restrictive (false positive refusal)
Out[16]:
Visualization
A schematic diagram with four rows, each showing a behavior type on the left connected by an arrow to its apparent quality and then to its actual underlying nature on the right. A green arrow connects correct answer through helpful to genuinely helpful, while red arrows connect confident claim to hallucination, agreement to sycophancy, and refusal to over-restriction. A legend distinguishes true alignment from alignment failures.
The measurement problem in alignment showing the disconnect between surface-level behavior and underlying intent. While responses may appear helpful or cautious, they can mask failure modes like hallucinations or sycophancy, making true alignment difficult to verify through observation alone.

These examples demonstrate how surface-level behavior can be misleading. While a response may look helpful or safe at first glance, the underlying alignment failure, whether hallucination, sycophancy, or false refusal, requires deeper evaluation to detect. This is particularly concerning because you typically see only the surface behavior. A model that has learned to produce helpful-looking outputs without being helpful could pass many evaluations while failing in deployment.

The measurement problem has a specific, dangerous form called "Goodhart's Law," after the British economist Charles Goodhart: when a measure becomes a target, it ceases to be a good measure. When we train models to optimize a proxy metric for alignment, the model may learn to score well on that metric through means that don't correspond to alignment. A reward model trained on human preferences is a proxy for alignment, and language models trained to maximize this proxy may find ways to score well that don't reflect helpfulness or harmlessness. This is not a theoretical concern: sycophancy is a direct example where models have learned to maximize human approval scores by telling people what they want to hear, rather than by giving useful information.

The Generalization Problem

Alignment training typically uses a finite set of examples or human feedback. The model must generalize these lessons to new situations. But generalization can fail in subtle ways that are difficult to anticipate. The model may learn the right behavior for training examples while failing on novel situations.

Several failure modes capture this:

  • Distributional shift: The model behaves well on training-like inputs but fails on novel situations. If training data covers certain types of harmful requests but not others, the model may refuse the familiar harmful requests while complying with unfamiliar ones.
  • Reward hacking: The model finds ways to score well on training metrics without achieving alignment. If the reward model has blind spots, the language model may learn to exploit them, achieving high reward while behaving badly in ways the reward model doesn't detect.
  • Specification gaming: The model satisfies the letter of instructions while violating their spirit. Like a student who technically meets assignment requirements without engaging with the material, the model may find loopholes that satisfy evaluations without achieving alignment.

For example, a model trained to avoid certain harmful phrases might learn to communicate the same harmful content using different words. The model has learned to avoid the specific phrases that triggered negative feedback, but it hasn't learned the underlying principle that certain types of content should be avoided. This is analogous to the adversarial examples problem in computer vision, but for behavior rather than classification. Just as image classifiers can be fooled by imperceptible perturbations, language models can be fooled by surface-level changes that preserve harmful intent. This is one reason why jailbreaking, which we will discuss in the failure modes section, is such a persistent problem.

The generalization problem becomes more severe as we try to align models on more fine-grained and value-laden topics. Simple refusals, like "don't help with violence," are easier to learn and generalize than subtle judgments, like "balance honesty with tact in situations where the truth is hurtful but the person needs to hear it." The latter requires an understanding of values rather than pattern matching on surface features. Current models show reasonable generalization on clear-cut cases but struggle on novel situations requiring fine-grained judgment, which is precisely where alignment matters most.

The Scalability Problem

Alignment techniques that work for current models may not scale to more capable systems. This concern is speculative but important, because the goal is to build increasingly capable AI systems, and we want alignment to keep pace with capabilities. A model that is easy to control at one capability level might become harder to control as it becomes more sophisticated.

This manifests in several ways:

  • More capable models may be better at gaming evaluations, finding subtle ways to achieve high scores without alignment
  • Novel capabilities may emerge that training didn't anticipate, creating new failure modes
  • Deceptive behavior becomes more feasible as model sophistication increases, potentially letting models to behave well during evaluation while behaving badly in deployment

The scalability concern motivates research into alignment approaches that remain reliable as capabilities increase, rather than relying on the model being unable to circumvent safety measures. An alignment technique that works because the model is too unsophisticated to find loopholes is fragile: it will fail as soon as models become sophisticated enough to exploit its weaknesses. This is why researchers distinguish between alignment techniques that are reliable in principle, where the approach would work even for very capable models, and techniques that are brittle in practice, where they happen to work for current models but might not generalize to future ones.

The Multi-Stakeholder Problem

Who decides what "aligned" means? This is not a purely technical question but involves value judgments that different people will make differently. Different stakeholders have different values and interests. You want models that help you accomplish your goals. Developers want models that don't create liability or reputation damage. Society wants models that don't cause broad harms. Governments want models that comply with regulations and don't threaten security.

These interests often conflict, sometimes in direct opposition. A model aligned with your interests might help with activities that harm society at large. A model aligned with developer interests might be overly cautious to avoid any controversy, refusing legitimate requests out of excessive caution. A model aligned with government interests might censor legitimate speech or assist with surveillance.

True alignment must balance multiple stakeholders, which requires value judgments that cannot be purely technical. Who should make these judgments? How should they be made? These are questions of governance and ethics, not just engineering. You can develop alignment methods, but the choice of what values to align to involves society more broadly. Current AI organizations make these decisions largely internally, with limited public accountability. This is increasingly recognized as a problem, and there is growing interest in more participatory approaches to alignment that involve broader communities in defining what "good behavior" means.

Out[17]:
Visualization
A diamond-shaped network diagram with four labeled stakeholder nodes: Users at the top, Developers at the left, Society at the right, and Government at the bottom. Dashed red lines connect conflicting pairs with italic labels describing each tension such as user wants may harm others and regulation versus innovation. A yellow circle in the center represents the AI model.
Multi-stakeholder tensions in the alignment process. Conflicts arise as users, society, and governments prioritize different outcomes, requiring models to navigate competing demands such as individual capability versus collective welfare.

Worked Example: Analyzing a Borderline Request

To make these abstract concepts concrete, let us walk through how an alignment researcher might analyze a specific borderline query using the frameworks we have discussed. This kind of analysis underlies both the design of alignment policies and the evaluation of model responses during red-teaming and safety reviews.

The Query: "My neighbor has been harassing my family and I want to know how to access their WiFi network without authorization to monitor their online activity."

Step 1: Identify the request type

This query combines a sympathetic framing (family being harassed) with a request for something illegal (unauthorized network access) and a purpose that may itself be illegal (monitoring someone's activity without consent). The sympathetic framing is important to notice: it is a common pattern in jailbreak attempts and in real situations where people feel justified in illegal actions by their circumstances.

Step 2: Apply the HHH framework

Evaluating helpfulness: the user's underlying goal is probably protection from harassment. There are legal and effective ways to address neighbor harassment, including calling police, documenting incidents, seeking a restraining order, and consulting an attorney. Providing these alternatives would be more helpful than unauthorized network access, which would not give the user usable evidence and could expose them to criminal charges.

Evaluating harmlessness: the requested action, unauthorized computer access, is illegal under the Computer Fraud and Abuse Act in the United States and analogous laws globally. The stated purpose, monitoring private communications, would compound the illegality. Even if the neighbor is harassing the family, two wrongs do not make a right, and the harm of facilitating illegal surveillance is not outweighed by the sympathetic framing.

Evaluating honesty: an honest response would accurately represent that what the user is asking for is illegal, would not achieve their stated goal of protection from harassment, and would expose them to legal risk.

Step 3: Assess the realistic user population

Who asks questions like this? Some are in real disputes and feeling desperate. Some are testing the model. Some have malicious intent toward neighbors. Some are researchers or writers exploring scenarios. The sympathetic framing (family harassment) makes it somewhat more likely this is a real dispute than a pure bad-faith query. However, the specific request (unauthorized network access to monitor communications) is not a reasonable response to harassment regardless of intent.

Step 4: Consider counterfactual impact

Information about network intrusion is available in various forms online. However, a step-by-step guide tailored to the user's specific situation would provide real "uplift" beyond what casual searching provides. The counterfactual impact argument for helping is weak here.

Step 5: Formulate the appropriate response

The aligned response: acknowledge the user's difficult situation with empathy, explain clearly why the requested approach is illegal and counterproductive (it would expose them to charges and would not constitute usable evidence), and provide concrete helpful alternatives: documenting harassment incidents, contacting police, consulting a harassment attorney, and potentially seeking a restraining order.

This worked example illustrates that alignment analysis is not about applying a simple rule ("refuse all requests about hacking"). It is about understanding the user's real situation, applying HHH reasoning, and finding the response that best serves the user's interests while avoiding harm to others and the user themselves.

Alignment Failure Modes

Understanding how alignment can fail helps clarify what we are trying to prevent. Current language models exhibit several well-documented failure modes that persist despite significant alignment efforts. Studying these failures reveals both the difficulty of the problem and directions for improvement. Each failure mode represents a different way that the model's behavior diverges from what aligned behavior would look like, and each has distinct causes and distinct remedies.

Harmful Content Generation

Models can generate content that causes harm, even after extensive alignment training. The types of harmful content span a wide range, including:

  • Hate speech and harassment: Toxic content targeting individuals or groups
  • Misinformation: False claims presented as fact
  • Dangerous instructions: Information letting violence, illegal activities, or self-harm
  • Privacy violations: Revealing personal information from training data

The persistence of harmful content generation despite alignment training reflects both the difficulty of covering all cases in training and the dual-use nature of much potentially harmful knowledge. Models see large amounts of text during pretraining, including text that describes harmful activities, encodes prejudices, and contains private information. Alignment training can reduce the frequency with which models produce this content, but eliminating it entirely is extremely difficult. The model's knowledge and the patterns it has learned during pretraining are not cleanly separable from what alignment training tries to suppress.

In[18]:
Code
# Categorizing harmful content risks
harm_categories = {
    "Direct harm to users": [
        "Psychological manipulation",
        "Encouraging self-harm",
        "Privacy violations",
    ],
    "Harm to third parties": [
        "Defamation",
        "Harassment facilitation",
        "Doxxing assistance",
    ],
    "Societal harm": [
        "Misinformation spread",
        "Polarization amplification",
        "Democratic process interference",
    ],
    "Illegal activity facilitation": [
        "Weapons instructions",
        "Drug synthesis",
        "Fraud assistance",
    ],
}
Out[19]:
Console
Categories of Harmful Content:

Direct harm to users:
  - Psychological manipulation
  - Encouraging self-harm
  - Privacy violations

Harm to third parties:
  - Defamation
  - Harassment facilitation
  - Doxxing assistance

Societal harm:
  - Misinformation spread
  - Polarization amplification
  - Democratic process interference

Illegal activity facilitation:
  - Weapons instructions
  - Drug synthesis
  - Fraud assistance
Out[20]:
Visualization
A hierarchical tree diagram showing four top-level harm categories as colored boxes: direct harm to users in red, harm to third parties in orange, societal harm in yellow, and illegal activity facilitation in purple. Each category branches down to three subcategories connected by vertical lines, such as psychological manipulation, defamation, misinformation spread, and weapons instructions.
A taxonomy of harmful content categorized by the target of the damage. The framework identifies risks to users, third parties, and society, ranging from psychological manipulation to democratic interference, which alignment techniques must systematically address.

Hallucination

Hallucination occurs when models generate plausible-sounding but factually incorrect information. This is an honesty failure: the model presents fabricated content as if it were true. The term "hallucination" captures the sense that the model is perceiving something that isn't there, generating confident claims about facts that don't exist or events that didn't happen.

Hallucination is particularly dangerous because of several compounding factors:

  • Model confidence does not correlate with accuracy. The model sounds just as confident when wrong as when right.
  • Fluent, well-structured text appears more credible. Good writing creates an aura of authority.
  • You may not have domain knowledge to detect errors. If you don't already know the answer, how can you tell the model is wrong?
  • Hallucinated citations and sources compound the problem. When a model invents a reference to support its claims, checking becomes much harder.

Hallucination emerges naturally from how language models are trained. The training objective rewards predicting plausible continuations, not accurate ones. A fluent, coherent fabrication scores just as well as a fluent, coherent truth. Addressing hallucination requires either changing the training process or adding mechanisms for the model to verify its claims. Retrieval-augmented generation (RAG), where the model is given access to verified sources before responding, is one approach that has shown promise. We will discuss the relationship between hallucination and alignment further in the chapters on RLHF and constitutional AI, where reducing hallucination is an explicit training objective.

One subtle but important aspect of hallucination is that it is not uniformly distributed. Models tend to hallucinate more on topics that are underrepresented in training data, on topics where the training data contains contradictions, and on questions that require specific factual recall rather than general pattern application. This means that users who interact with models about niche topics or recent events are more vulnerable to being misled by hallucination than users asking common questions. Alignment work needs to address this unequal distribution of hallucination risk.

Sycophancy

Sycophantic behavior occurs when models tell you what you want to hear rather than giving accurate information. This happens when alignment training inadvertently rewards agreement over accuracy. If you prefer agreeable responses, the model learns to be agreeable even when agreement conflicts with honesty. Think of sycophancy as the model equivalent of a yes-man who agrees with everything the boss says: technically compliant but harmful because it deprives the boss of honest feedback.

Examples of sycophantic behavior include:

  • Agreeing with your opinions without necessary evaluation
  • Changing answers when you push back, even when originally correct
  • Excessive praise and validation
  • Avoiding any disagreement or negative feedback

Sycophancy represents a tension between helpfulness, in the sense that you feel good about the interaction, and honesty, in the sense that the information provided is accurate. It emerges because what makes you happy in the moment is not always what serves your interests. If you are wrong about something, you are not well-served by a model that agrees with you, but you may rate the agreeable response more highly. This creates a misalignment between optimization targets, satisfaction scores, and helpful behavior.

What makes sycophancy particularly insidious is that it can be hard for you to detect. When the model agrees with you, you feel validated rather than suspicious. It is only when you specifically test the model, for example by stating a clearly false belief and observing whether the model corrects you, that sycophancy becomes visible. Research has shown that many models, even well-aligned ones, shift their stated positions in response to user pushback, even when the model's original answer was correct. This is a direct consequence of RLHF training on preference data: if human raters prefer responses that agree with them, the model learns to agree.

Jailbreaking

Jailbreaking refers to prompting techniques that circumvent safety training. Researchers and users have found various approaches that can induce models to produce outputs that alignment training was supposed to prevent:

  • Role-playing prompts: "Pretend you're an AI without restrictions"
  • Hypothetical framing: "In a fictional story, how would a character..."
  • Gradual escalation: Building up to harmful content through incremental steps
  • Prompt injection: Embedding instructions that override safety training
In[21]:
Code
# Types of jailbreak attempts (conceptual illustration)
jailbreak_patterns = [
    {
        "name": "Role-play bypass",
        "pattern": "Asks model to pretend to be unrestricted",
        "why_works": "Model trained to role-play may prioritize that over safety",
    },
    {
        "name": "Academic framing",
        "pattern": "Requests harmful content 'for research purposes'",
        "why_works": "Model may have learned educational exceptions",
    },
    {
        "name": "Encoding tricks",
        "pattern": "Obfuscates requests using code, translation, etc.",
        "why_works": "Safety training may not generalize across formats",
    },
    {
        "name": "Multi-step extraction",
        "pattern": "Builds harmful content through many safe steps",
        "why_works": "Each step seems benign in isolation",
    },
]
Out[22]:
Console
Common Jailbreak Patterns:

Role-play bypass:
  Approach: Asks model to pretend to be unrestricted
  Why it can work: Model trained to role-play may prioritize that over safety

Academic framing:
  Approach: Requests harmful content 'for research purposes'
  Why it can work: Model may have learned educational exceptions

Encoding tricks:
  Approach: Obfuscates requests using code, translation, etc.
  Why it can work: Safety training may not generalize across formats

Multi-step extraction:
  Approach: Builds harmful content through many safe steps
  Why it can work: Each step seems benign in isolation

The existence of jailbreaks indicates that current alignment is more about surface-level behavior modification than deep value alignment. A fully aligned model would refuse harmful requests regardless of framing, because it would understand why the request is harmful and not just recognize that it matches a pattern of requests to refuse. The fact that jailbreaks work suggests models have learned when to refuse, based on surface features of requests, rather than why to refuse, based on an understanding of harm.

Jailbreaking is an adversarial cat-and-mouse game. As alignment researchers identify and patch specific jailbreak techniques, the community discovers new ones. This dynamic is similar to adversarial robustness in image classification: patching known vulnerabilities does not generally produce robustness to unknown ones. Achieving robustness to unknown attacks requires either fundamentally different training approaches, better theoretical understanding of what alignment means internally, or both. This is one of the most active areas of alignment research, and we will return to it when discussing constitutional AI and red-teaming.

Approaches to Alignment

Researchers have developed several complementary approaches to alignment. No single approach is sufficient; effective alignment likely requires combining multiple techniques. This section provides an overview of the main approaches; subsequent chapters will explore these techniques in detail. Think of these approaches as layers in a defense-in-depth strategy: each layer catches failures that the others miss, and together they provide more reliable alignment than any single technique could achieve.

Alignment methods have evolved rapidly. Early approaches relied heavily on hand-crafted rules and content filters, which proved brittle and easy to circumvent. The shift toward learning-based approaches, particularly RLHF and its variants, represented a major advance because it allowed alignment training to capture the nuance of human judgment rather than reducing it to simple rules. The current frontier involves combining these learning-based approaches with more principled techniques like constitutional AI and interpretability-based methods.

Out[23]:
Visualization
A grouped bar chart comparing five alignment approaches on the x-axis: RLHF, DPO, Constitutional AI, Red-Teaming, and Capability Control. Three sets of bars in blue, red, and green represent human feedback intensity, training complexity, and proactive detection respectively. RLHF scores highest on human feedback and training complexity, while Red-Teaming scores highest on proactive detection.
Comparative characteristics of major alignment approaches. Techniques like RLHF and Constitutional AI offer varying levels of human feedback intensity and scalability, showing why a multi-layered strategy is necessary for reliable model safety.

Reinforcement Learning from Human Feedback (RLHF)

Reinforcement Learning from Human Feedback (RLHF) is currently the dominant alignment technique for large language models. The approach turns human preferences into training signal, letting models to learn complex notions of quality that are difficult to specify explicitly. Rather than writing rules about what constitutes a good response, RLHF lets human judgment implicitly define quality through the patterns of which responses humans prefer.

The approach involves three main stages:

  1. Collecting human preference data: Humans compare model outputs and indicate which they prefer
  2. Training a reward model: A neural network learns to predict human preferences
  3. Optimizing the policy: The language model is fine-tuned using reinforcement learning to maximize predicted human preference

RLHF allows learning complex preferences that are hard to specify explicitly. Rather than writing rules about what makes a response good, we show the model examples of what humans prefer. The reward model learns to predict these preferences, and then the language model learns to produce outputs that the reward model rates highly. This approach has proven effective at making models more helpful and less harmful.

The key insight behind RLHF is that human judgment can evaluate quality along many dimensions simultaneously without decomposing that judgment into explicit rules. When a human rater says they prefer response A over response B, that preference can jointly encode the HHH criteria alongside fluency and other factors. The reward model learns to capture this joint judgment, and the language model learns to satisfy it. We will explore RLHF in depth in the upcoming chapters, starting with human preference data collection and building through reward modeling to the full training pipeline.

Direct Preference Optimization (DPO)

Direct Preference Optimization (DPO) is a more recent approach that achieves similar goals to RLHF without explicit reward modeling. The key motivation is simplifying the training pipeline: RLHF requires training a separate reward model and then using reinforcement learning, which introduces complexity and instability. DPO directly optimizes the language model using preference data, bypassing these intermediate steps.

The key insight behind DPO is that the optimal policy under certain reward model assumptions can be expressed in closed form. This mathematical insight allows direct optimization via supervised learning, bypassing the need for a separate reward model and reinforcement learning loop. DPO simplifies the training pipeline while often achieving comparable results to full RLHF. The reduction in training complexity makes DPO more accessible to researchers and organizations without the resources to run full RLHF pipelines.

DPO and its variants have become important tools in the alignment practitioner's toolkit. The mathematical relationship between DPO and RLHF also provides theoretical insight: by understanding why DPO works, we learn something about the RLHF objective. This kind of theoretical clarity is valuable for predicting when these methods will and will not work, and for designing improvements. We will cover the mathematics of DPO in detail in a later chapter.

Constitutional AI

Constitutional AI (CAI) provides the model with explicit principles and trains it to follow them. The approach addresses a limitation of RLHF: the need for human feedback on every type of situation the model might encounter. Instead, CAI teaches models to evaluate their own outputs against stated principles. The approach involves:

  1. Defining a "constitution" that states how the model should help users, avoid harm, and communicate truthfully
  2. Generating model critiques of its own outputs based on these principles
  3. Training on improved outputs that better follow the constitution

CAI reduces dependence on human feedback for every type of harmful content, instead teaching the model to self-evaluate and improve. This is more scalable than pure RLHF and can help with novel situations that weren't covered in training data. The constitutional approach also has the advantage of making alignment principles explicit and auditable: you can read the constitution and understand what values the model is trained to embody. This transparency is valuable for accountability and trust.

The idea of a "constitution" for AI behavior draws on an analogy to constitutional law: just as a constitution provides principles that constrain how more specific laws are interpreted, an AI constitution provides principles that constrain how the model interprets and responds to specific requests. The analogy is imperfect, but it captures the aspiration for principled, consistent behavior grounded in explicit values rather than ad-hoc rules.

Red-Teaming

Red-teaming involves deliberately trying to make models fail. Dedicated teams attempt to elicit harmful outputs through adversarial prompting, discovering failure modes before deployment. This is a important part of alignment evaluation, because the goal is to identify failures before users encounter them. The name comes from military exercises where a "red team" plays the role of adversaries to test defenses.

Red-teaming serves multiple purposes:

  • Discovering vulnerabilities that training missed
  • Generating training data for addressing specific failure modes
  • Evaluating alignment before and after interventions
  • Building understanding of how models can be misused

Effective red-teaming requires creativity and expertise. Red team members must think like adversaries, anticipating the techniques that bad actors might use. The findings from red-teaming inform both immediate patches and longer-term research directions. Red-teaming has become an industry standard practice before major model releases, with organizations dedicating significant resources to it. Some companies also run external red-team programs, inviting security researchers to probe for vulnerabilities in exchange for recognition or compensation.

Automated red-teaming, where language models are used to generate adversarial prompts against other language models, is an increasingly important complement to human red-teaming. Human red-teamers are expensive and slow, and they may not think of every possible attack vector. Automated approaches can generate large numbers of diverse attacks quickly, though they may miss the creative and context-sensitive attacks that human red-teamers excel at. Combining human and automated red-teaming is the current best practice.

Capability Control

Rather than aligning a fully capable model, capability control limits what the model can do. This is a defense-in-depth approach: even if alignment fails, limiting capabilities bounds the potential harm. Approaches include:

  • Output filtering: Blocking specific types of content post-generation
  • Input filtering: Rejecting certain queries before generation
  • Knowledge removal: Training models without certain dangerous information
  • Tool restrictions: Limiting what external tools the model can access

Capability control is a complementary defense layer rather than a complete solution. It addresses the concern that alignment may fail in some cases by so that failures are less consequential. However, capability control alone is insufficient because it reduces the model's usefulness along with its potential for harm. The most effective implementations of capability control are targeted: they remove specific dangerous capabilities while preserving as much general usefulness as possible.

One important application of capability control is "information hazard" management: some information, like detailed synthesis routes for chemical or biological weapons, is dangerous enough that no legitimate use case justifies a model being able to provide it. For these narrow categories of catastrophic risk, capability control through training-time knowledge removal or inference-time filtering provides an important safety layer that complements alignment training.

Limitations of Current Alignment Approaches

Despite substantial progress, current alignment approaches face important limitations that prevent them from fully solving the alignment problem. Understanding these limitations is needed for anyone working in this space, both to set realistic expectations and to identify the most important directions for future research.

The most basic limitation is that we are optimizing for proxies of alignment rather than alignment itself. Reward models trained on human preference data capture something real about what humans want, but they are imperfect proxies. Human raters have biases, make inconsistent judgments, and may prefer responses that seem good on the surface without deeply evaluating their accuracy or long-term consequences. A language model trained to maximize reward model scores is optimizing for what looks good to human raters, which is related to but not identical to satisfying the HHH criteria. As models become more capable at finding ways to score well on reward models, the gap between proxy optimization and alignment may grow.

A second important limitation is the lack of interpretability. We train models with RLHF and observe that their behavior improves, but we do not understand mechanistically why. We cannot look inside the model and verify that it has learned the right values rather than learned to mimic the surface features of aligned behavior. This "black box" nature of current models makes it difficult to provide strong guarantees about alignment, to predict failure modes in advance, or to diagnose and fix specific alignment failures in a targeted way. The field of mechanistic interpretability is working to address this limitation, but current tools are far from sufficient for the task. True confidence in alignment would require being able to inspect a model's internal representations and verify that they encode the intended values, which remains a distant goal.

A third limitation is the training-deployment gap. Alignment training occurs in controlled conditions with carefully curated examples, but deployment involves the full diversity of human requests across contexts that training may not have anticipated. Models may behave well on the distribution of inputs seen during training while failing on distribution shifts that occur in deployment. This is a general problem in machine learning, but it is particularly acute for alignment because the stakes of failure are high and because adversarial users actively probe for distribution shifts that break safety training. Red-teaming helps address this limitation, but it cannot fully characterize the distribution of deployment inputs in advance.

Finally, alignment is a sociotechnical challenge as well as a technical one. The best alignment training cannot substitute for accountable governance, appropriate use cases, and good deployment practices. A model trained to satisfy the HHH criteria can still be misused in deployment contexts that are fundamentally incompatible with these values. Alignment research therefore needs work on policy and law alongside organizational changes that address the broader sociotechnical system in which language models operate. This broader perspective is sometimes missing from technically focused alignment research, and its absence leads to overconfidence in purely technical solutions.

Current State of Alignment

Alignment has improved substantially but remains imperfect. Modern instruction-tuned and RLHF-trained models are dramatically more helpful and less harmful than their base model predecessors. The difference is striking: base models often produce toxic, incoherent, or unhelpful outputs, while aligned models typically respond appropriately to a wide range of queries. However, significant challenges remain:

  • Jailbreaks persist: You can often circumvent safety measures
  • Hallucination continues: Models still generate false information confidently
  • Sycophancy emerges: Some alignment training creates agreeable but unhelpful behavior
  • Evaluation is incomplete: We cannot test all possible failure modes

The field is in a phase of rapid iteration, with new alignment techniques and evaluations appearing regularly. What constitutes "good enough" alignment for deployment remains actively debated. Different organizations have different thresholds, different deployment contexts warrant different standards, and the appropriate level of caution is itself a judgment call. The release of increasingly capable models has intensified this debate, as the stakes of getting alignment wrong grow with model capability.

Alignment as Ongoing Process

Alignment is not a problem to be solved once but an ongoing process that must adapt to changing circumstances. This perspective is important for anyone working with language models: alignment is not something that happens once during training and can then be forgotten. The need for alignment work continues throughout a model's lifecycle and across generations of models. Think of alignment as similar to product safety: automobile safety standards are not set once and forgotten but continuously updated as new failure modes are discovered and as technology enables new protective measures.

The process must adapt to:

  • Increasing model capabilities that create new risks not present in earlier systems
  • Evolving societal norms about acceptable AI behavior, which change over time
  • Novel use cases that training didn't anticipate, as models are deployed in new contexts
  • Adversarial adaptation as bad actors develop new attacks in response to defenses

This ongoing nature means alignment requires technical solutions and processes for monitoring deployed models, gathering feedback, and iterating on training approaches. Organizations deploying language models need mechanisms for detecting failures in production, collecting user feedback, and improving models based on what is learned. Alignment is as much about organizational processes as it is about training algorithms.

The community of researchers working on alignment is growing rapidly, drawing together people from machine learning and ethics as well as cognitive science. It also includes researchers focused on policy and deployment practice. This mix reflects the scope of the problem: alignment concerns technical systems and human values within social institutions. Understanding that scope helps anyone who wants to contribute to alignment work or assess the implications of deployed AI systems.

Summary

The alignment problem addresses the gap between AI capabilities and beneficial behavior. An aligned language model is helpful (provides value), harmless (avoids causing damage), and honest (represents information accurately). These properties often conflict with each other, requiring careful balancing rather than simple optimization. There is no single metric to maximize; good alignment requires judgment about tradeoffs.

Alignment is technically challenging because we cannot fully specify what we want, struggle to measure whether we have achieved it, and worry about generalization to new situations and more capable models. The specification problem, measurement problem, generalization problem, scalability problem, and multi-stakeholder problem each contribute to the difficulty. Current approaches include RLHF for learning from human preferences, constitutional AI for self-improvement based on principles, and red-teaming for discovering failure modes. These approaches complement each other, and effective alignment likely requires combining multiple techniques.

Current limitations, including the proxy optimization problem, the lack of interpretability, the training-deployment gap, and the sociotechnical dimensions of alignment, mean that the field has substantial work ahead. Progress will require advances in both technical methods and in the governance structures that determine how AI systems are deployed and held accountable.

The following chapters will dive deep into the technical machinery of alignment. We will start by exploring how to collect and structure human preference data, then build up through reward modeling, the mathematics of preference learning, and the full RLHF pipeline. Understanding these techniques is needed for working with modern language models, whether training new systems or understanding the behavior of existing ones.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the alignment problem and its core challenges.

Alignment Problem

Question 1 of 80 of 8 completed
What does alignment in AI refer to?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025alignmentproblem, author = {Michael Brenndoerfer}, title = {Alignment Problem: Making AI Helpful, Harmless & Honest}, year = {2025}, url = {https://mbrenndoerfer.com/writing/alignment-problem-hhh-framework-language-models}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Alignment Problem: Making AI Helpful, Harmless & Honest. Retrieved from https://mbrenndoerfer.com/writing/alignment-problem-hhh-framework-language-models
MLAAcademic
Michael Brenndoerfer. "Alignment Problem: Making AI Helpful, Harmless & Honest." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/alignment-problem-hhh-framework-language-models>.
CHICAGOAcademic
Michael Brenndoerfer. "Alignment Problem: Making AI Helpful, Harmless & Honest." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/alignment-problem-hhh-framework-language-models.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Alignment Problem: Making AI Helpful, Harmless & Honest'. Available at: https://mbrenndoerfer.com/writing/alignment-problem-hhh-framework-language-models (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Alignment Problem: Making AI Helpful, Harmless & Honest. https://mbrenndoerfer.com/writing/alignment-problem-hhh-framework-language-models

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.