Part of History of Language AI
Covers continuous post-training, including parameter-efficient fine-tuning with LoRA, catastrophic forgetting prevention, incremental model updates.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
2025: Continuous Post-Training
A deployed language model does not learn from a conversation merely because it produced a response. Its parameters remain fixed until another training run changes them. By 2025, model developers were running those later training stages more often and treating them as a recurring maintenance process.
The term continuous post-training describes that process. It is an umbrella term, not one algorithm introduced in a single paper. A team starts with a pretrained model, collects a bounded update dataset, trains an adaptation, evaluates the candidate, and decides whether to release it. New cycles can address observed errors, a changed policy, or a capability needed by an application.
This practice combines ideas that predate 2025. Parameter-efficient fine-tuning reduces the number of trainable weights. Continual-learning methods address interference between old and new behavior. Model registries and staged rollouts provide the operational controls needed to ship repeated updates. Calling the process continuous does not mean that the model updates itself live; most systems still use discrete, supervised release cycles.
Continuous post-training also differs from retrieval. A retrieval system can place a recent document in the prompt without changing model parameters. Post-training changes behavior encoded in the model, but it is a poor substitute for a database when facts change often or must be cited exactly. Production systems may use both.
The Problem
A fixed model develops two kinds of mismatch. Its factual associations age, and its behavior may not fit the way people use it. For example, an old model may recommend a deprecated library API. Deployment logs may also reveal a recurring failure on a prompt pattern that was rare in the original evaluation set.
Updating all weights can be expensive. It can also damage unrelated behavior when the update set overrepresents one task. This failure is usually called catastrophic forgetting: training on the new distribution improves the target task while reducing performance on earlier ones.
The problem is not solved by training more frequently. Each candidate still needs traceable data, regression tests, and a release decision. Without those controls, frequent training can make failures harder to reproduce because two deployments may use different model versions.
Different defects also call for different interventions. Retrieval is often suitable for changing facts. A small adapter may teach a stable output format. A safety-policy change may require preference data and broader adversarial evaluation. Architectural limits cannot be repaired by attaching another adapter.
The Solution
Continuous post-training turns model updates into a pipeline with a narrow release unit. Each cycle begins with a written objective, such as reducing failures on a particular tool call. The team assembles examples for that objective and keeps separate regression data for behavior that should not change.
Parameter-efficient fine-tuning (PEFT) makes a small release unit practical. LoRA, for example, freezes the base model and learns low-rank updates for selected weight matrices. Training and storing those updates costs less than producing another full copy of every model weight.
A replay set can mix older examples into the update batch. Regularization can also discourage changes to parameters associated with earlier tasks. Neither technique guarantees preservation, so the candidate must be tested on both its target task and the existing release criteria.
If the candidate passes offline tests, it can move through a limited rollout before replacing the current version. The model registry should record the base checkpoint, adapter, data revision, training configuration, and evaluation result. Keeping the previous combination makes rollback possible.
This loop resembles continuous delivery in software, but there is an important difference: a model update may change behavior outside the examples used to train it. A passing target metric is therefore necessary but insufficient evidence for release.
Training Methodology
LoRA represents an update to a weight matrix with two smaller matrices. Let and , where . The adapted weight is
Only and are trained. This can reduce trainable parameter count and storage, although inference cost depends on whether the adapter remains separate or is merged into the base weights.
A practical update cycle has four boundaries:
- At the data boundary, record where each example came from, why it belongs in the update, and which model version produced any logged response. Remove private or low-quality feedback before training.
- At the training boundary, pin the base checkpoint and training configuration. Decide whether the update is an adapter, a merged checkpoint, or a full fine-tune.
- At the evaluation boundary, measure the stated target and rerun regression sets. Human review is useful when a numeric score cannot capture correctness or policy compliance.
- At the release boundary, assign a version, begin with limited traffic, monitor predefined metrics, and retain a rollback path.
Forgetting controls fit inside the training and evaluation boundaries. Replay exposes the candidate to selected examples from earlier tasks. Elastic Weight Consolidation (EWC) adds a penalty for changing parameters estimated to matter to those tasks. Separate adapters can isolate changes, but combining them may create interference that was absent when each adapter was tested alone.
An accumulating adapter stack needs an explicit policy. A system may select one adapter per request, compose several adapters at inference time, or merge updates into a new checkpoint. Each choice changes latency and storage requirements. It also changes what must be tested: two adapters that pass independently can still conflict when composed.
Applications and Impact
The clearest use case is a stable, repeated error for which a team can collect representative examples. Suppose a tool-using model repeatedly chooses the wrong API operation for one class of request. An adapter can target that decision, while regression sets check that other tool calls still work.
Post-training can also adjust output conventions or domain behavior. A support model might need to follow a revised escalation policy. A coding model might need examples from a newly adopted internal API. These cases have a bounded desired behavior, which makes update data and acceptance criteria easier to define.
Fresh factual lookup is less suitable. Encoding today's interest rate or package version into weights provides no guarantee that the model will recall it correctly tomorrow. Retrieval from a maintained source is usually easier to inspect and update. Post-training may still teach the model how to use that source or when to decline an unsupported answer.
Shorter training runs make experiments cheaper, but they do not make feedback automatically trustworthy. Usage logs overrepresent active users and frequent tasks. A team must account for consent, privacy, sampling bias, and malicious submissions before turning logs into training examples.
Limitations
The first limitation is regression risk. Evaluations sample behavior; they cannot prove that every unaffected capability stayed fixed. More releases create more chances for interactions between updates, especially after adapters are merged or composed.
The second is provenance. Training from deployment feedback can reproduce private data or reinforce the preferences of whichever users generate the most logs. Removing identifiers is not enough if an example itself contains sensitive context.
The third is operational cost. PEFT reduces training and storage costs, but a reliable pipeline still needs dataset review, reproducible jobs, evaluation capacity, version tracking, rollout controls, and monitoring. The surrounding system can cost more engineering time than the adapter run.
Repeated updates also fragment a model fleet. If customers pin versions or adapters differ by application, the organization must know which base and which updates served every response. That record is necessary for incident analysis.
Finally, an incremental update cannot change the base architecture. Large domain shifts may expose tokenizer, context-window, or objective limitations that require a new base model. Periodic consolidation or full retraining may also be simpler than maintaining a long chain of interacting adapters.
What Continuous Post-Training Changed
By 2025, post-training was no longer best understood as one final step between pretraining and release. Teams could maintain a base checkpoint and ship smaller behavioral revisions around it. PEFT made those revisions cheaper to train and store, while continual-learning research supplied ways to measure interference.
The lasting change was operational. Model development acquired recurring release concerns familiar from software systems: versioned artifacts, regression gates, staged deployment, and rollback. Models still differ from ordinary software because their behavior cannot be enumerated in advance, so those controls reduce risk without eliminating it.
The central design question is therefore not "How can this model learn continuously?" It is "Which change should enter the weights, and what evidence is enough to release it?" That distinction keeps rapidly changing facts in inspectable data systems while reserving post-training for behavior that merits a model update.
Quiz
The following questions review the update pipeline, LoRA, catastrophic forgetting, and the limits of incremental training.
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore History of Language AIStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!