KV Cache Compression: Eviction, Quantization & H2O Algorithm
Covers KV cache compression techniques including eviction strategies, attention sinks, the H2O algorithm, and INT8 quantization for efficient LLM inference.
637 items · Page 5 of 14
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.

Covers KV cache compression techniques including eviction strategies, attention sinks, the H2O algorithm, and INT8 quantization for efficient LLM inference.
Control false positives across multiple hypothesis tests with FWER, FDR, Bonferroni, Holm, and Benjamini-Hochberg procedures.
Explains how web crawlers discover and download internet content at scale, covering Common Crawl architecture, robots.txt compliance, politeness policies.
Benchmark scores can mislead when test data leaks, metrics saturate, or systems game proxies. Covers contamination, brittleness, Goodhart's Law, and mitigation.
Covers transaction cost analysis and market impact modeling. Estimate spread, slippage, and liquidity to build realistic backtests and execution strategies.
Cohen's d measures effect size separately from statistical significance. Covers practical significance and why small p-values can describe small effects.
Build rigorous NLP benchmarks: task definition, dataset collection, annotation guidelines, inter-annotator agreement, contamination detection.
Covers backtesting frameworks to validate trading strategies. Avoid look-ahead bias, measure risk-adjusted returns.
Choose sample size and minimum detectable effect with power analysis. Covers underpowered studies, statistical sensitivity, and design tradeoffs.
Explains how KV cache eliminates redundant attention computations in transformers. Topics include memory requirements, cache structure.
Covers event-driven trading strategies including merger arbitrage, earnings plays, and fixed income relative value.
Understanding false positives, false negatives, statistical power, and the tradeoff between error types. Balance Type I and Type II errors in study design.
Explains how multimodal AI powers document understanding, medical imaging diagnostics, video analysis, and complex reasoning tasks across real-world domains.
Covers iterative alignment for LLMs with online DPO, rolling references, Constitutional AI, and SPIN. Build self-improving models beyond single-shot training.
Examines cryptocurrency market structure, quantitative strategies for extreme volatility, and risk management in 24/7 decentralized trading.
Covers diffusion models from the DDPM objective to latent diffusion, classifier-free guidance, image editing techniques, ControlNet, and FID evaluation metrics.
One-way ANOVA, post-hoc tests, assumptions, and when to use ANOVA. Compare means across three or more groups while controlling Type I error rates.
Covers RLAIF and Constitutional AI for scalable model alignment. Use AI feedback, design constitutions, and train reward models effectively.
Extract trading signals from alternative data using NLP. Topics include sentiment analysis, text processing, and building news-based trading systems.
Covers four core image understanding tasks: visual question answering, image captioning, visual grounding, and scene understanding.
Why attention weights fail as explanations, how manipulation studies expose their limits, and what gradient-based methods offer as more faithful alternatives.
F-distribution, F-test for comparing variances, F-test in regression, and nested model comparison.
Multimodal models align text, images, and other inputs in shared embedding spaces. Covers contrastive loss, dual encoders, joint encoders, and tradeoffs.
Build safe autonomous agents using action constraints, sandboxing, real-time monitoring, prompt injection defenses, and human-in-the-loop intervention patterns.
Explains how multi-agent AI systems coordinate through topologies, communication protocols, role assignment.
Build ML-driven trading strategies covering return prediction, sentiment analysis, alternative data integration, and reinforcement learning for execution.
Covers t-tests including one-sample, two-sample (pooled and Welch), paired tests, assumptions, and decision framework.
Covers supervised ML algorithms for trading: linear models, random forests, gradient boosting.
Mathematical equivalence between confidence intervals and hypothesis tests, test assumptions (independence, normality, equal variances).
Covers z-tests including one-sample, two-sample, and proportion tests. Topics include when to use z-tests, how to calculate test statistics.
Explains how LLM agents plan, decompose tasks, execute multi-step goals, and recover from failures using architectures like ReWOO, Tree-of-Thought.
Foundation of hypothesis testing covering p-values, null and alternative hypotheses, one-sided vs two-sided tests.
Explains how market makers capture bid-ask spreads, manage inventory risk, and use the Avellaneda-Stoikov model for quote placement.
Covers volatility as an asset class. Topics include delta hedging, variance swaps, dispersion trading.
Build long-short factor portfolios using quintile rankings. Topics include value, momentum, quality, and volatility factors with exposure analysis.
Explains how PPO applies to language models. Topics include policy mapping, token action spaces, KL divergence penalties, and advantage estimation for RLHF.
Covers time-series and cross-sectional momentum strategies. Implement moving average crossovers, breakout systems, and CTA approaches with Python code.
Explains mean reversion through cointegration tests, pairs trading, factor-neutral portfolios, and regime risk management.
Covers policy gradient theory for language model alignment. Topics include REINFORCE algorithm, variance reduction with baselines, and foundations for PPO.
Covers quantitative trading fundamentals: alpha generation, strategy categories, backtesting workflows, and performance metrics for systematic investing.
Examines reward hacking in RLHF where language models exploit proxy objectives. Topics include distribution shift, over-optimization, and mitigation strategies.
Explains Credit Valuation Adjustment for derivatives pricing, including exposure profiles, default probability modeling, and the broader XVA framework.
Collect and process human preference data for RLHF. Topics include pairwise comparisons, annotator guidelines, quality metrics, and interface design.
Examines the AI alignment problem and HHH framework. Explains why training language models to be helpful, harmless, and honest presents fundamental challenges.
Covers credit risk measurement through Probability of Default, Loss Given Default, and Exposure at Default. Topics include loan pricing and portfolio analysis.
Evaluate instruction-tuned LLMs using benchmarks like Alpaca Eval and MT-Bench, human evaluation protocols, and LLM-as-Judge automatic methods.
Covers VaR calculation using parametric, historical, and Monte Carlo methods. Examines Expected Shortfall and stress testing for market risk management.
Covers market, credit, liquidity, operational, and model risk. Topics include Basel III capital requirements and risk management governance structures.
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.
