2026

an archive of posts from this year

Aug 09, 2026 The Expressive Power of Transformers with Chain of Thought
Aug 07, 2026 Landscape of Thoughts — Visualizing Where LLM Reasoning Actually Goes
Jul 27, 2026 Magellan — Guided MCTS for Escaping the Gravity Wells of LLM Creativity
Jul 27, 2026 PriorZero — Injecting LLM Priors into MuZero-Style World Models at the MCTS Root
Jul 27, 2026 SuperThoughts — Reasoning Tokens in Superposition
Jul 27, 2026 Huginn — Scaling Test-Time Compute via Recurrent Depth in Latent Space
Jul 27, 2026 LoopFormer — Elastic-Depth Looped Transformers via Shortcut Modulation
Jul 27, 2026 Think-at-Hard — Selective Latent Iteration for Looped Reasoning Transformers
Jul 19, 2026 Process Reward Agents — Online Step-Wise Steering for Knowledge-Intensive Reasoning
Jul 19, 2026 Tele-Lens — How Far Ahead Do LLMs Actually Plan in Chain-of-Thought?
Jul 19, 2026 Reasoning Cache — Continual Improvement Over Long Horizons via Short-Horizon RL
Jul 19, 2026 Context-Folding — Scaling Long-Horizon LLM Agents via Branch-and-Fold
Jul 19, 2026 T3S — Training-Trajectory-Aware Token Selection for Continual Reasoning Distillation
Jul 18, 2026 BG-MCTS — Budget-Guided Tree Search for Fixed Token Budgets in LLM Reasoning
Jul 18, 2026 Blend-ASC — Optimal Self-Consistency via Power-Law Sample Efficiency
Jul 18, 2026 SOL — Self-Optimizing Language Models via Token-Level Efficiency Policies
Jul 02, 2026 WFM-TTS — Test-Time Scaling for World Foundation Models
Jul 02, 2026 Monte Carlo Tree Diffusion — System 2 Planning with Diffusion Models
Jun 13, 2026 Video Prediction Policy — Predictive Visual Representations as the Policy Backbone
Jun 13, 2026 DreamGen — Scaling Robot Learning by Dreaming Trajectories
Jun 08, 2026 LAPA — Latent Action Pretraining from Action-Label-Free Video
Jun 08, 2026 Flow Matching — The Simulation-Free Recipe Under Modern Diffusion
Jun 08, 2026 DreamZero — World Action Models as Zero-shot Robot Policies
May 25, 2026 Polar — Agentic RL on Any Harness at Scale
May 25, 2026 From RLHF to RULER: How the Reward Signal for RL Agents Evolved
May 21, 2026 AlpaServe — Statistical Multiplexing with Model Parallelism for DL Serving
May 18, 2026 LeWorldModel — A 15M-Parameter JEPA That Actually Trains End-to-End from Pixels
May 18, 2026 FrontierSmith — Manufacturing Open-Ended Coding Problems to Train Better Code Agents
May 18, 2026 ECHO — Terminal Agents Learn World Models for Free
May 18, 2026 MEMENTO: Teaching Reasoning Models to Compress Their Own Thinking
May 18, 2026 SOAR — Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
May 18, 2026 Learning Fast and Slow: A New Recipe for Adapting LLMs Without Forgetting
May 18, 2026 Parcae: Scaling Laws for Stable Looped Language Models
May 18, 2026 Meta-Harness: End-to-End Optimization of Model Harnesses