QA: Extractive, Generative, and Open-Domain QA
Covers question answering systems from span extraction with BERT to retrieval-augmented generation, covering evaluation metrics and open-domain QA pipelines.
637 items · Page 4 of 14
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.

Covers question answering systems from span extraction with BERT to retrieval-augmented generation, covering evaluation metrics and open-domain QA pipelines.
Examines vector similarity search for RAG systems. Compare cosine, dot product, and Euclidean metrics, and implement exact vs. approximate search with FAISS.
Walk through the complete lifecycle of a quantitative trading strategy. Build a pairs trading system from scratch with rigorous backtesting and risk management.
Explains how learning rate schedules improve neural network training. Topics include step decay, exponential decay, inverse square root with warmup.
Covers ethical quantitative trading by learning to detect spoofing, navigate Reg NMS and MiFID II, implement kill switches, and ensure data privacy compliance.
Explains how embedding models convert text to vectors for RAG. Topics include bi-encoder architecture, pooling strategies, dimensionality trade-offs.
Covers optimal position sizing using the Kelly Criterion, risk budgeting, and volatility targeting.
Covers document chunking for RAG systems. Examines fixed-size, recursive, and semantic strategies to balance retrieval precision with context window limits.
Save and restore complete training state, choose checkpoint frequency, implement asynchronous I/O, and recover distributed LLM training runs from failures.
Covers contrastive learning for dense retrieval. Train models using InfoNCE loss, in-batch negatives, and hard negative mining strategies effectively.
Covers distributed training communication: gradient compression, topology-aware all-reduce, NCCL tuning.
Explains how mixed precision training uses FP16 and BF16 floating point formats to speed up LLM training and cut memory usage without sacrificing accuracy.
Build a robust quantitative research pipeline. From hypothesis formulation and backtesting to paper trading and live production deployment strategies.
Covers dense retrieval for semantic search. Examines bi-encoder architectures, embedding metrics, and contrastive learning to overcome keyword limitations.
How activation checkpointing trades compute for memory by discarding and recomputing activations.
Explains how FSDP shards model parameters, gradients, and optimizer states across GPUs to train billion-parameter models.
Explains how ZeRO eliminates memory redundancy in distributed training by partitioning optimizer states, gradients, and parameters.
Covers RAG system design by exploring retriever-generator interactions, timing strategies like iterative retrieval, and architectural variations like RETRO.
Examines quantitative trading-system architecture, including data pipelines, strategy engines, risk controls, and execution infrastructure.
Explains why LLMs need Retrieval-Augmented Generation. Explains how RAG bridges knowledge gaps, reduces hallucinations, and enables non-parametric memory.
Explains how pipeline parallelism splits deep models across devices, manages bubble overhead with micro-batching, and compares GPipe vs 1F1B schedules.
Covers extractive and abstractive summarization, from TextRank and MMR to BART and LLMs, with ROUGE and BERTScore evaluation techniques.
Explains how tensor parallelism splits weight matrices across GPUs using column and row strategies, enabling training of models too large for any single device.
Covers LLM inference serving architecture, token-aware load balancing, and auto-scaling. Optimize time-to-first-token and throughput for production systems.
Explains how continuous batching achieves 2-3x throughput gains in LLM inference through iteration-level scheduling, eliminating static batch inefficiencies.
Covers the mathematical framework for speculative decoding, including the exact acceptance criterion, rejection sampling logic.
Explains how activation patching locates where information flows in transformers through causal tracing, path patching, and component attribution experiments.
Explains how DDP trains large models across multiple GPUs using ring all-reduce gradient synchronization, gradient bucketing.
Speculative decoding uses a smaller draft model to propose tokens that a larger model verifies in parallel, reducing latency without changing output quality.
Explains how GPU memory breaks down into parameters, gradients, optimizer states, and activations. Estimate memory requirements and debug out-of-memory errors.
Covers GGUF format for storing quantized LLMs. Topics include file structure, quantization types, llama.cpp integration.
Explains how language models power creative writing, poetry generation, storytelling, and the ethical questions of authorship and originality in generative AI.
Covers GPU memory hierarchy, CUDA cores, Tensor Core throughput, and how to read GPU specs to optimize training workloads and debug performance bottlenecks.
Explains how Activation-aware Weight Quantization protects salient weights to compress LLMs. Topics include the algorithm, scaling factors.
Explains how adversarial prompts bypass LLM safety training, covering roleplay attacks, GCG gradient-based suffixes, PAIR automation.
Interpret sparse autoencoder features through activation patterns, automated naming, monosemanticity measurement, and causal feature circuit analysis.
Explains how GPTQ optimizes weight quantization using Hessian-based error compensation to compress LLMs to 4 bits while maintaining near-FP16 accuracy.
Explains how LLMs are trained on synthetic data, from Self-Instruct and Evol-Instruct to quality verification, diversity control, and knowledge distillation.
Covers INT4 quantization techniques for LLMs. Topics include group-wise quantization, NF4 format, double quantization.
Construct effective pretraining data recipes by setting domain proportions, applying quality weighting, running proxy experiments.
Detect and remove personal information from LLM training data with regex, named entity recognition, and hybrid privacy pipelines.
Explains how toxicity classifiers work, how to calibrate thresholds for pretraining data.
Filter low-quality text from web corpora using heuristic rules, perplexity scoring, and classifier-based methods with tunable thresholds.
Explains how weight quantization maps floating-point values to integers, reducing LLM memory by 4x. Topics include scale, zero-point.
Explains how MinHash compresses documents into compact signatures that estimate Jaccard similarity, enabling near-duplicate detection.
Explains how language identification works in NLP pipelines. Topics include n-gram models, FastText LID, code-switching, confidence thresholds.
Practical reporting guidelines, summary of key concepts, test selection parameters table, multiple comparison corrections table.
Extract clean text from HTML and PDFs for LLM training data. Topics include boilerplate removal, text density algorithms, Trafilatura.
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.
