Replay Methods: Buffer, Pseudo-Rehearsal & Generative Replay
Explains how experience replay buffers, pseudo-rehearsal, and generative replay prevent catastrophic forgetting in continual learning systems.
637 items · Page 3 of 14
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.

Explains how experience replay buffers, pseudo-rehearsal, and generative replay prevent catastrophic forgetting in continual learning systems.
Explains how transformer models power sentiment analysis, topic classification, and intent detection.
Explains how classification-based and rule-based content filters protect language model deployments.
Explains how Elastic Weight Consolidation and Synaptic Intelligence protect critical parameters to prevent catastrophic forgetting in continual learning.
Design reliable LLM evaluations with representative test sets, suitable metrics, and controls for contamination, overfitting, and misleading averages.
Covers the continual learning problem, why neural networks catastrophically forget sequential tasks, and the three canonical learning scenarios.
OpenAI Whisper's encoder-decoder architecture enables multilingual speech recognition. Explains how multitask training and special tokens unify transcription.
Explains how sparse autoencoders decompose language model activations into interpretable features using overcomplete dictionaries, sparsity constraints.
Explains how language models link generated claims to source documents, evaluate citation accuracy with NLI, and measure attribution precision and recall.
Explains how speech recognition systems transform raw audio waveforms into mel spectrograms using STFT, mel filterbanks.
Covers LLM cost management with token-based cost modeling, optimization techniques like caching and quantization, cost allocation.
Covers multimodal AI evaluation with VQA benchmarks, image captioning metrics (BLEU, CIDEr, CLIPScore), and text-to-image quality (FID).
Monitor LLM output quality in production: track BLEU, BERTScore, and semantic metrics, detect drift with KS tests and CUSUM, catch regressions with A/B testing.
Monitor production LLM systems with metrics collection, structured logging, dashboard design, alerting rules, and concept drift detection.
Flamingo adds visual input to frozen language models with a Perceiver Resampler and gated cross-attention. Covers its architecture and training design.
Explains how the logit lens projects transformer hidden states into vocabulary space to reveal how predictions evolve.
LLaVA connects a frozen vision encoder to a language model through a projection layer. Covers visual instruction tuning, architecture, data, and evaluation.
Extract attention weights from transformer models, visualize head patterns, measure head entropy, and understand the key caveats.
Covers LLM caching strategies including prompt caching with KV reuse, semantic caching with embedding similarity, cache invalidation policies.
Route LLM requests intelligently across model tiers using rule-based, difficulty-based, and cascade strategies, plus A/B testing to validate routing decisions.
Auto-scale LLM serving infrastructure using GPU utilization, queue depth, and TTFT metrics. Topics include horizontal vs vertical scaling.
Explains how projection layers bridge vision encoders and LLMs. Topics include linear transformation, MLP projection.
How $1 trillion vanished from software stocks in days. The mechanics of leverage, algorithms, and margin calls behind the Feb 2026 crash.
Covers LLM throughput optimization with continuous batching, batch size tuning, parallelism strategies, and measuring tokens-per-second under real workloads.
Covers LLM serving latency from TTFT to TPOT. Topics include KV cache management, continuous batching, speculative decoding, streaming responses.
Serve language models in production: weight loading strategies, tensor parallelism, request batching, rate limiting.
Explains how layer probing maps linguistic properties across BERT layers, from surface features in lower layers to semantic content in upper layers.
Measure code LLM quality using execution-based evaluation, the pass@k metric, and key benchmarks like HumanEval, APPS, and SWE-bench.
Explains how LLM agents safely execute generated code using sandboxed environments, capture execution feedback.
Transform raw text into structured data using named entity recognition, relation extraction, event extraction, and LLM-based JSON output generation.
Covers code generation with LLMs, covering docstring-to-code, test-driven generation, sampling strategies, self-repair loops, and generation quality metrics.
Explains how LLM-powered code completion works: context assembly with fill-in-the-middle, candidate ranking, latency optimizations like speculative decoding.
Explains why interpretability matters for language models, covering debugging, trust, safety alignment, scientific understanding.
Language models can explain code, detect bugs, support review, and run semantic code search. Covers dual encoders, contrastive training, and evaluation.
Explains how red teams systematically probe language models for safety failures, covering attack taxonomies, ASR metrics, RL-based automated attack generation.
Covers direct and indirect prompt injection attacks, how RAG pipelines create injection surfaces, and the layered defense strategies used to mitigate them.
Explains how code language models are trained: code data curation, code-specific tokenization, fill-in-the-middle objectives.
Covers systematic strategies for hyperparameter search, how muP enables transfer across model scales.
Explains how steering vectors and activation addition let you modify language model behavior at inference time.
Explains how Product Quantization compresses embeddings up to 100x using learned codebooks and asymmetric distance computation for scalable vector search.
Detect and prevent training instability in deep learning. Topics include loss spikes, gradient norm monitoring, gradient clipping.
Explains how gradient accumulation simulates large batch sizes on limited GPU memory by splitting batches into micro-batches.
Explains how weight decay regularizes neural networks, why AdamW decouples weight decay from adaptive gradients.
Explains how batch size affects gradient noise and generalization, apply the linear scaling rule for learning rates, identify the critical batch size.
Covers IVF indexes for scalable vector search. Topics include clustering-based partitioning, nprobe tuning, and IVF-PQ compression for billion-scale retrieval.
Examines harmful content categories, misuse scenarios, unintended harms, and structured threat models for reasoning about safety in language AI systems.
Covers cosine learning rate schedule used in GPT and LLaMA training: the decay formula, warm restarts, key parameters, and comparison with linear decay.
Covers Hierarchical Navigable Small World (HNSW) graphs for vector search. Topics include graph architecture, construction, and tuning for high-speed retrieval.
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.
