Data, Analytics & AI

Articles about data, analytics, and AI, including machine learning, data visualization, and AI applications.

637 items · Page 3 of 14

Community

Learn with the community

Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.

  • Keep your place across books and articles
  • Ask questions and take part in discussions
  • Read with fewer interruptions

Free to join · Takes less than a minute

Why

I created this space so readers can learn together, ask questions, and make sense of difficult ideas.

Michael Brenndoerfer
From readers1 / 12
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Replay Methods: Buffer, Pseudo-Rehearsal & Generative Replay

Feb 22, 2026·52 min read

Explains how experience replay buffers, pseudo-rehearsal, and generative replay prevent catastrophic forgetting in continual learning systems.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Classification Applications: Sentiment, Intent, and Scale

Feb 22, 2026·61 min read

Explains how transformer models power sentiment analysis, topic classification, and intent detection.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Content Filtering: Classification, Rules, and Evaluation

Feb 22, 2026·60 min read

Explains how classification-based and rule-based content filters protect language model deployments.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Regularization Methods: EWC and Synaptic Intelligence

Feb 22, 2026·48 min read

Explains how Elastic Weight Consolidation and Synaptic Intelligence protect critical parameters to prevent catastrophic forgetting in continual learning.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

LLM Evaluation Fundamentals: Goals, Design, and Metrics

Feb 21, 2026·59 min read

Design reliable LLM evaluations with representative test sets, suitable metrics, and controls for contamination, overfitting, and misleading averages.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Continual Learning: Catastrophic Forgetting and Scenarios

Feb 19, 2026·57 min read

Covers the continual learning problem, why neural networks catastrophically forget sequential tasks, and the three canonical learning scenarios.

Open notebook
Data, Analytics & AIMachine LearningLanguage AI Handbook

Whisper Architecture: Encoder-Decoder Speech Recognition

Feb 16, 2026·59 min read

OpenAI Whisper's encoder-decoder architecture enables multilingual speech recognition. Explains how multitask training and special tokens unify transcription.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Sparse Autoencoders: Dictionary Learning for LLM Features

Feb 16, 2026·59 min read

Explains how sparse autoencoders decompose language model activations into interpretable features using overcomplete dictionaries, sparsity constraints.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Attribution and Citation: Sourcing LLM Outputs

Feb 16, 2026·50 min read

Explains how language models link generated claims to source documents, evaluate citation accuracy with NLI, and measure attribution precision and recall.

Open notebook
Language AI HandbookMachine LearningData, Analytics & AI

Speech Representations: From Waveforms to Mel Spectrograms

Feb 15, 2026·60 min read

Explains how speech recognition systems transform raw audio waveforms into mel spectrograms using STFT, mel filterbanks.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

LLM Cost Management: Modeling, Optimization & Allocation

Feb 14, 2026·40 min read

Covers LLM cost management with token-based cost modeling, optimization techniques like caching and quantization, cost allocation.

Open notebook
Data, Analytics & AIMachine LearningLanguage AI Handbook

Multimodal Evaluation: VQA Benchmarks, Metrics

Feb 14, 2026·54 min read

Covers multimodal AI evaluation with VQA benchmarks, image captioning metrics (BLEU, CIDEr, CLIPScore), and text-to-image quality (FID).

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Quality Monitoring: Drift Detection

Feb 14, 2026·55 min read

Monitor LLM output quality in production: track BLEU, BERTScore, and semantic metrics, detect drift with KS tests and CUSUM, catch regressions with A/B testing.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Monitoring LLM Systems: Metrics, Logging, Alerting

Feb 13, 2026·46 min read

Monitor production LLM systems with metrics collection, structured logging, dashboard design, alerting rules, and concept drift detection.

Open notebook
Language AI HandbookMachine LearningData, Analytics & AI

Flamingo Architecture

Feb 12, 2026·57 min read

Flamingo adds visual input to frozen language models with a Perceiver Resampler and gated cross-attention. Covers its architecture and training design.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Logit Lens: Reading Transformer Hidden States

Feb 12, 2026·45 min read

Explains how the logit lens projects transformer hidden states into vocabulary space to reveal how predictions evolve.

Open notebook
Language AI HandbookMachine LearningSoftware EngineeringData, Analytics & AI

LLaVA: Visual Instruction Tuning

Feb 11, 2026·58 min read

LLaVA connects a frozen vision encoder to a language model through a projection layer. Covers visual instruction tuning, architecture, data, and evaluation.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Attention Visualization: Extracting and Interpreting Weights

Feb 11, 2026·51 min read

Extract attention weights from transformer models, visualize head patterns, measure head entropy, and understand the key caveats.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Caching for LLMs: Prompt, Semantic, and Invalidation

Feb 11, 2026·52 min read

Covers LLM caching strategies including prompt caching with KV reuse, semantic caching with embedding similarity, cache invalidation policies.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Model Routing: Selection, A/B Testing, Cascades & Strategies

Feb 10, 2026·53 min read

Route LLM requests intelligently across model tiers using rule-based, difficulty-based, and cascade strategies, plus A/B testing to validate routing decisions.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Auto-Scaling LLMs: Metrics, Policies, Production Strategies

Feb 10, 2026·54 min read

Auto-scale LLM serving infrastructure using GPU utilization, queue depth, and TTFT metrics. Topics include horizontal vs vertical scaling.

Open notebook
Data, Analytics & AIMachine LearningLanguage AI Handbook

Vision-Language Projection: Bridging Vision Encoders, LLMs

Feb 10, 2026·52 min read

Explains how projection layers bridge vision encoders and LLMs. Topics include linear transformation, MLP projection.

Open notebook
Economics & FinanceData, Analytics & AI

The February 2026 Selloff

Feb 8, 2026·27 min read

How $1 trillion vanished from software stocks in days. The mechanics of leverage, algorithms, and margin calls behind the Feb 2026 crash.

Read article
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Throughput Optimization: Batch Size, GPU Utilization

Feb 6, 2026·58 min read

Covers LLM throughput optimization with continuous batching, batch size tuning, parallelism strategies, and measuring tokens-per-second under real workloads.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Latency Optimization: TTFT, Batching, Speculative Decoding

Feb 6, 2026·52 min read

Covers LLM serving latency from TTFT to TPOT. Topics include KV cache management, continuous batching, speculative decoding, streaming responses.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Model Serving: Frameworks, Loading

Feb 6, 2026·57 min read

Serve language models in production: weight loading strategies, tensor parallelism, request batching, rate limiting.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Probing Layers in Transformer Models

Feb 5, 2026·55 min read

Explains how layer probing maps linguistic properties across BERT layers, from surface features in lower layers to semantic content in upper layers.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Code Evaluation: Functional Correctness and pass@k

Feb 5, 2026·53 min read

Measure code LLM quality using execution-based evaluation, the pass@k metric, and key benchmarks like HumanEval, APPS, and SWE-bench.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Code Execution: Sandboxed Feedback and Iterative Refinement

Feb 4, 2026·53 min read

Explains how LLM agents safely execute generated code using sandboxed environments, capture execution feedback.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Information Extraction: NER, Relations, Structured Output

Feb 4, 2026·53 min read

Transform raw text into structured data using named entity recognition, relation extraction, event extraction, and LLM-based JSON output generation.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Code Generation: Docstring-to-Code, Test-Driven, Strategies

Feb 3, 2026·52 min read

Covers code generation with LLMs, covering docstring-to-code, test-driven generation, sampling strategies, self-repair loops, and generation quality metrics.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Code Completion: Context, Ranking, Latency & UX

Feb 3, 2026·55 min read

Explains how LLM-powered code completion works: context assembly with fill-in-the-middle, candidate ranking, latency optimizations like speculative decoding.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Interpretability Goals: Debugging, Trust, and Safety

Feb 2, 2026·46 min read

Explains why interpretability matters for language models, covering debugging, trust, safety alignment, scientific understanding.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Code Understanding: Explanation, Bug Detection, Review

Feb 2, 2026·48 min read

Language models can explain code, detect bugs, support review, and run semantic code search. Covers dual encoders, contrastive training, and evaluation.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Red Teaming LLMs: Methodology, Taxonomies, and Automation

Jan 30, 2026·61 min read

Explains how red teams systematically probe language models for safety failures, covering attack taxonomies, ASR metrics, RL-based automated attack generation.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Prompt Injection: Attacks, RAG Risks, and Defenses

Jan 29, 2026·50 min read

Covers direct and indirect prompt injection attacks, how RAG pipelines create injection surfaces, and the layered defense strategies used to mitigate them.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Code LLM Training: Tokenization, FIM, Pretraining Objectives

Jan 29, 2026·56 min read

Explains how code language models are trained: code data curation, code-specific tokenization, fill-in-the-middle objectives.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Hyperparameter Selection: Search, Transfer, Default Recipes

Jan 29, 2026·58 min read

Covers systematic strategies for hyperparameter search, how muP enables transfer across model scales.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Activation Steering: Vectors and Representation Engineering

Jan 29, 2026·47 min read

Explains how steering vectors and activation addition let you modify language model behavior at inference time.

Open notebook
Data, Analytics & AIMachine LearningLanguage AI Handbook

Product Quantization: Vector Compression for ANN Search

Jan 28, 2026·62 min read

Explains how Product Quantization compresses embeddings up to 100x using learned codebooks and asymmetric distance computation for scalable vector search.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Training Stability: Loss Spikes, Gradient Norms & Debugging

Jan 28, 2026·60 min read

Detect and prevent training instability in deep learning. Topics include loss spikes, gradient norm monitoring, gradient clipping.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Gradient Accumulation: Memory-Efficient Large Batch Training

Jan 27, 2026·49 min read

Explains how gradient accumulation simulates large batch sizes on limited GPU memory by splitting batches into micro-batches.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Weight Decay: L2 Regularization, AdamW, Decoupled Training

Jan 27, 2026·56 min read

Explains how weight decay regularizes neural networks, why AdamW decouples weight decay from adaptive gradients.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Large Batch Training: Scaling Rules, Limits, and LAMB

Jan 27, 2026·49 min read

Explains how batch size affects gradient noise and generalization, apply the linear scaling rule for learning rates, identify the critical batch size.

Open notebook
Machine LearningData, Analytics & AILanguage AI Handbook

IVF Index: Clustering-Based Vector Search & Partitioning

Jan 27, 2026·66 min read

Covers IVF indexes for scalable vector search. Topics include clustering-based partitioning, nprobe tuning, and IVF-PQ compression for billion-scale retrieval.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Safety Risks in Language AI: Harms, Misuse, Threat Models

Jan 26, 2026·52 min read

Examines harmful content categories, misuse scenarios, unintended harms, and structured threat models for reasoning about safety in language AI systems.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

Cosine Learning Rate Schedule: Decay, Restarts, and Warmup

Jan 26, 2026·65 min read

Covers cosine learning rate schedule used in GPT and LLaMA training: the decay formula, warm restarts, key parameters, and comparison with linear decay.

Open notebook
Data, Analytics & AISoftware EngineeringMachine LearningLanguage AI Handbook

HNSW Index: Architecture for Fast Vector Search

Jan 26, 2026·67 min read

Covers Hierarchical Navigable Small World (HNSW) graphs for vector search. Topics include graph architecture, construction, and tuning for high-speed retrieval.

Open notebook
Community

Learn with the community

Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.

  • Keep your place across books and articles
  • Ask questions and take part in discussions
  • Read with fewer interruptions

Free to join · Takes less than a minute

Why

I created this space so readers can learn together, ask questions, and make sense of difficult ideas.

Michael Brenndoerfer
From readers1 / 12