| Aug 09, 2026 | The Expressive Power of Transformers with Chain of Thought |
| Aug 07, 2026 | Landscape of Thoughts — Visualizing Where LLM Reasoning Actually Goes |
| Jul 27, 2026 | Magellan — Guided MCTS for Escaping the Gravity Wells of LLM Creativity |
| Jul 27, 2026 | PriorZero — Injecting LLM Priors into MuZero-Style World Models at the MCTS Root |
| Jul 27, 2026 | SuperThoughts — Reasoning Tokens in Superposition |
| Jul 27, 2026 | Huginn — Scaling Test-Time Compute via Recurrent Depth in Latent Space |
| Jul 27, 2026 | LoopFormer — Elastic-Depth Looped Transformers via Shortcut Modulation |
| Jul 27, 2026 | Think-at-Hard — Selective Latent Iteration for Looped Reasoning Transformers |
| Jul 19, 2026 | Process Reward Agents — Online Step-Wise Steering for Knowledge-Intensive Reasoning |
| Jul 19, 2026 | Tele-Lens — How Far Ahead Do LLMs Actually Plan in Chain-of-Thought? |
| Jul 19, 2026 | Reasoning Cache — Continual Improvement Over Long Horizons via Short-Horizon RL |
| Jul 19, 2026 | Context-Folding — Scaling Long-Horizon LLM Agents via Branch-and-Fold |
| Jul 19, 2026 | T3S — Training-Trajectory-Aware Token Selection for Continual Reasoning Distillation |
| Jul 18, 2026 | BG-MCTS — Budget-Guided Tree Search for Fixed Token Budgets in LLM Reasoning |
| Jul 18, 2026 | Blend-ASC — Optimal Self-Consistency via Power-Law Sample Efficiency |
| Jul 18, 2026 | SOL — Self-Optimizing Language Models via Token-Level Efficiency Policies |
| Jul 02, 2026 | WFM-TTS — Test-Time Scaling for World Foundation Models |
| Jul 02, 2026 | Monte Carlo Tree Diffusion — System 2 Planning with Diffusion Models |
| Jun 13, 2026 | Video Prediction Policy — Predictive Visual Representations as the Policy Backbone |
| Jun 13, 2026 | DreamGen — Scaling Robot Learning by Dreaming Trajectories |
| Jun 08, 2026 | LAPA — Latent Action Pretraining from Action-Label-Free Video |
| Jun 08, 2026 | Flow Matching — The Simulation-Free Recipe Under Modern Diffusion |
| Jun 08, 2026 | DreamZero — World Action Models as Zero-shot Robot Policies |
| May 25, 2026 | Polar — Agentic RL on Any Harness at Scale |
| May 25, 2026 | From RLHF to RULER: How the Reward Signal for RL Agents Evolved |
| May 21, 2026 | AlpaServe — Statistical Multiplexing with Model Parallelism for DL Serving |
| May 18, 2026 | LeWorldModel — A 15M-Parameter JEPA That Actually Trains End-to-End from Pixels |
| May 18, 2026 | FrontierSmith — Manufacturing Open-Ended Coding Problems to Train Better Code Agents |
| May 18, 2026 | ECHO — Terminal Agents Learn World Models for Free |
| May 18, 2026 | MEMENTO: Teaching Reasoning Models to Compress Their Own Thinking |
| May 18, 2026 | SOAR — Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability |
| May 18, 2026 | Learning Fast and Slow: A New Recipe for Adapting LLMs Without Forgetting |
| May 18, 2026 | Parcae: Scaling Laws for Stable Looped Language Models |
| May 18, 2026 | Meta-Harness: End-to-End Optimization of Model Harnesses |