Hierarchical, Symbolic, Language, and Multi-Agent Planning
Explains how options, symbolic STRIPS planning, language grounding, and multi-agent belief models structure long-horizon planning and where abstractions fail.
682 items · Page 1 of 15
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.

Explains how options, symbolic STRIPS planning, language grounding, and multi-agent belief models structure long-horizon planning and where abstractions fail.
Explains how agents trade reward for information using dual control, information gain, curiosity, and safe exploration in world-model decision making.
Explains how actor-critic policies train inside learned world models, covering latent rollouts, lambda-returns, value expansion, and model exploitation.
Covers MCTS, UCT, PUCT, belief-space planning, and POMCP for decision-making under uncertainty with a hidden-state maze and tiger problem.
Explains how gradients flow through world-model rollouts, from adjoint backpropagation to action optimization, terminal values, and long-horizon stability.
Explains how sampling-based planning and model predictive control turn black-box world models into controllers, using random shooting, CEM.
Explains how world models are pretrained on diverse data, post-trained with preference objectives.
Explains how world models adapt over time through offline, online, and continual learning, covering replay, catastrophic forgetting, and safe drift guardrails.
Curation determines what a world model can learn. Examines deduplication, filtering, balancing, synthetic data, mixture weights, and contamination risks.
Explains how interaction data and passive observation shape world models, covering coverage, behavior policies, interventions, and data mixing strategies.
World models define and learn actions through inverse dynamics, controllability, and latent variables discovered in unlabeled video and aligned to controls.
Compare reconstruction, contrastive, masked, and multi-step objectives, including how each loss shapes rollout fidelity and control-relevant state.
Examines hierarchical, hybrid, and omnimodal world models, covering temporal abstraction, neural-symbolic hybrids, multimodal fusion.
Explains how JEPA predicts in latent space instead of pixels, why joint-embedding objectives collapse.
Explains how diffusion and flow world models represent multimodal futures via conditional denoising, trajectory diffusion, flow matching, guidance, and latency.
Explains how causal transformers predict world dynamics from tokenized observations and actions.
Compare attention-based transformers, structured state-space models, and hybrid architectures for world models.
Explains how recurrent state-space models combine deterministic memory with stochastic latents for filtering and imagination, and train them with the ELBO.
Explains how compositional world models reuse mechanisms, objects, and skills to generalize systematically to novel combinations.
Aleatoric and epistemic uncertainty shape world models, from multimodal transitions and ensembles to calibration and uncertainty propagation through rollouts.
Differentiable simulators, graph networks, neural ODEs, and operator learning add physical structure to world models for efficient data use and planning.
Covers structural causal models, the do-operator, counterfactuals, confounding, and invariance.
Embed physics into world models using Lagrangian and Hamiltonian architectures, contact and friction constraints, conservation laws.
Explains how coordinate frames, depth estimation, SLAM, occupancy grids, and neural fields give world models geometric state that stays stable under egomotion.
Explains how world models represent the agent itself: agent-environment boundaries, learned body schemas, capabilities, affordances.
Explains how object-centric and relational world models use slots, graphs, and message passing to predict object dynamics, track identity.
Explains how world models maintain state under partial observability using recurrent memory, gated LSTMs and GRUs, attention, episodic storage.
Explains how autoencoders, VAEs, and VQ-VAEs compress raw observations into compact latent states for world models.
Explains what a world model's state must preserve and discard: predictive versus control sufficiency, bisimulation, minimal abstractions.
Explains how sensors shape world models: calibration, synchronization, noise, missing and delayed observations.
Explains how learned world models power planning, policy optimization, and value learning in RL, plus model bias, rollout error.
Covers control as inference, optimality variables, soft Bellman updates, variational free energy, active inference, preferences, ambiguity, and epistemic value.
Covers classical planning and optimal control: state-space search with A*, dynamic programming, LQR.
Covers system identification through excitation, regression, prediction-error and subspace methods, neural dynamics, rollout tests, and model validation.
Bayesian filters turn noisy observations into belief states. Covers Kalman, extended Kalman, unscented Kalman, particle, and learned update methods.
Explains how MDPs, POMDPs, policies, values, and belief states formalize sequential decisions when dynamics are uncertain and observations hide the true state.
Covers math behind world models: conditional distributions, Markov transitions, stability, attractors, and information theory for sufficient states.
Trace the four intellectual lineages behind world models: neuroscience, cybernetics, AI planning, and foundation models.
Map the world-model design space across four axes: explicit versus implicit, state representation, deterministic versus stochastic dynamics.
Covers the four-criteria world-model qualification test covering state sufficiency, action conditioning, counterfactual validity.
Explains what a world model is through operational and philosophical definitions, its core components, and a hands-on pendulum example testing prediction.
Build a minimal world model in PyTorch with encoder, transition model, decoder, and reward head. Topics include filtering, imagination, CEM planning, and MPC.
Deploy language models responsibly through staged rollouts, tiered access control, content filtering, and production monitoring systems.
Covers long-form text generation with outline-based planning, hierarchical decomposition, entity tracking.
Covers test-time compute strategies: multiple sampling, iterative refinement, compute-optimal inference, and inference-time scaling laws for language models.
Explains how learning rate warmup stabilizes early training by gradually increasing the learning rate, with theory, linear warmup.
Explains how deduplication removes exact copies and near-duplicates from training corpora using SHA-256 hashing, Jaccard similarity over character shingles.
Reduce LLM hallucination using retrieval augmentation, self-consistency decoding, DPO training, and calibrated uncertainty expression.
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.
