Jul 27, 2026 PriorZero — Injecting LLM Priors into MuZero-Style World Models at the MCTS Root Jul 19, 2026 Reasoning Cache — Continual Improvement Over Long Horizons via Short-Horizon RL Jul 19, 2026 Context-Folding — Scaling Long-Horizon LLM Agents via Branch-and-Fold Jul 18, 2026 SOL — Self-Optimizing Language Models via Token-Level Efficiency Policies May 25, 2026 Polar — Agentic RL on Any Harness at Scale May 25, 2026 From RLHF to RULER: How the Reward Signal for RL Agents Evolved May 18, 2026 ECHO — Terminal Agents Learn World Models for Free May 18, 2026 SOAR — Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability May 18, 2026 Learning Fast and Slow: A New Recipe for Adapting LLMs Without Forgetting Sep 28, 2025 RL Study Notes — Lecture 2: Value Functions and Bellman Equations Sep 21, 2025 RL Study Notes — Lecture 1: MDP, Behavior Cloning, DAgger Sep 21, 2025 HICRA: Hierarchical Credit Assignment for LLM Reasoning