paper-review

an archive of posts in this category

Aug 09, 2026 The Expressive Power of Transformers with Chain of Thought
Aug 07, 2026 Landscape of Thoughts — Visualizing Where LLM Reasoning Actually Goes
Jul 27, 2026 Magellan — Guided MCTS for Escaping the Gravity Wells of LLM Creativity
Jul 27, 2026 PriorZero — Injecting LLM Priors into MuZero-Style World Models at the MCTS Root
Jul 27, 2026 SuperThoughts — Reasoning Tokens in Superposition
Jul 27, 2026 Huginn — Scaling Test-Time Compute via Recurrent Depth in Latent Space
Jul 27, 2026 LoopFormer — Elastic-Depth Looped Transformers via Shortcut Modulation
Jul 27, 2026 Think-at-Hard — Selective Latent Iteration for Looped Reasoning Transformers
Jul 19, 2026 Process Reward Agents — Online Step-Wise Steering for Knowledge-Intensive Reasoning
Jul 19, 2026 Tele-Lens — How Far Ahead Do LLMs Actually Plan in Chain-of-Thought?
Jul 19, 2026 Reasoning Cache — Continual Improvement Over Long Horizons via Short-Horizon RL
Jul 19, 2026 Context-Folding — Scaling Long-Horizon LLM Agents via Branch-and-Fold
Jul 19, 2026 T3S — Training-Trajectory-Aware Token Selection for Continual Reasoning Distillation
Jul 18, 2026 BG-MCTS — Budget-Guided Tree Search for Fixed Token Budgets in LLM Reasoning
Jul 18, 2026 Blend-ASC — Optimal Self-Consistency via Power-Law Sample Efficiency
Jul 18, 2026 SOL — Self-Optimizing Language Models via Token-Level Efficiency Policies
Jul 02, 2026 WFM-TTS — Test-Time Scaling for World Foundation Models
Jul 02, 2026 Monte Carlo Tree Diffusion — System 2 Planning with Diffusion Models
Jun 13, 2026 Video Prediction Policy — Predictive Visual Representations as the Policy Backbone
Jun 13, 2026 DreamGen — Scaling Robot Learning by Dreaming Trajectories
Jun 08, 2026 LAPA — Latent Action Pretraining from Action-Label-Free Video
Jun 08, 2026 Flow Matching — The Simulation-Free Recipe Under Modern Diffusion
Jun 08, 2026 DreamZero — World Action Models as Zero-shot Robot Policies
May 25, 2026 Polar — Agentic RL on Any Harness at Scale
May 21, 2026 AlpaServe — Statistical Multiplexing with Model Parallelism for DL Serving
May 18, 2026 LeWorldModel — A 15M-Parameter JEPA That Actually Trains End-to-End from Pixels
May 18, 2026 FrontierSmith — Manufacturing Open-Ended Coding Problems to Train Better Code Agents
May 18, 2026 SOAR — Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
May 18, 2026 Parcae: Scaling Laws for Stable Looped Language Models
May 18, 2026 Meta-Harness: End-to-End Optimization of Model Harnesses
Oct 27, 2025 Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Sep 28, 2025 LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
Sep 21, 2025 HICRA: Hierarchical Credit Assignment for LLM Reasoning
Sep 21, 2025 REFRAG: Rethinking RAG-Based Decoding
Dec 11, 2023 Sliced Mutual Information for Memorization and Generalization
Dec 11, 2023 When Memorizing Irrelevant Data Becomes Necessary
Dec 11, 2023 Chain-of-Thought Prompting: Key Papers and Variants
Oct 09, 2023 GPT-3: Language Models Are Few-Shot Learners
Oct 09, 2023 GPT-2: Language Models Are Unsupervised Multitask Learners
Oct 09, 2023 Auto-Regressive Next-Token Predictors Are Universal Learners
Sep 06, 2023 Memorization Without Overfitting in Large Language Models
Sep 03, 2023 Beyond Chain-of-Thought: Graph-of-Thought Reasoning in LLMs
Sep 02, 2023 Scaling Laws for Neural Language Models
Aug 30, 2023 Graph of Thought: Boosting Logical Reasoning in LLMs
Aug 29, 2023 Knowledge Graph Prompting Sparks Graph of Thoughts in LLMs
Aug 28, 2023 Graph of Thoughts: Solving Elaborate Problems with LLMs
Jul 27, 2023 Neural Tangent Kernel: Infinite-Width Networks as Kernel Methods
Jul 27, 2023 Mini-batch Optimization of Contrastive Loss
Jul 24, 2023 Frequency Effects on Syntactic Rule Learning in Transformers
Jul 23, 2023 Why Mask Reconstruction Pretraining Helps in Downstream Tasks
Jul 22, 2023 Emergent Abilities of Large Language Models
Jul 18, 2023 Contextual Representation Learning beyond Masked Language Modeling
Jul 17, 2023 AMOM: Adaptive Masking over Masking (AAAI 2023)
Jul 16, 2023 Mask More and Mask Later (ACL 2022)
Jul 15, 2023 Calibration, Entropy Rates, and Memory in Language Models
Jul 13, 2023 A Closer Look at How Fine-tuning Changes BERT
Jul 12, 2023 AdaGDA: Faster Adaptive Gradient Descent Ascent for Minimax Optimization
Jul 11, 2023 Masked Latent Semantic Modeling (MLSM)
Jul 09, 2023 Singular Value Representation: A Graph Perspective on Neural Networks
Jul 07, 2023 Blessing of Class Diversity in Pre-training
Jul 06, 2023 Deriving Language Models from Masked Language Models
Feb 01, 2023 Neural Collapse: Terminal Phase of Deep Network Training
Feb 26, 2022 BERT: Pre-training of Deep Bidirectional Transformers
Feb 26, 2022 Attention Is All You Need