- llm
- nlp
- reasoning
- rl
- rlhf
- ssm
- theory
- information-theory
- representation-learning
- optimization
- rag
- alignment
- paper-review
- survey
- notes
•
•
•
•
•
•
•
•
•
•
•
•
•
•
-
SSM Part I: Modeling Sequences with Structured State Spaces
Study notes on structured state space models (S4) — recurrent, convolutional, and continuous views unified.
-
RL Study Notes — Lecture 2: Value Functions and Bellman Equations
Value functions, Bellman equations, optimal policies, and a Gridworld example.
-
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
Combining autoregressive language modeling with JEPA-style joint embedding prediction to improve representations.
-
RL Study Notes — Lecture 1: MDP, Behavior Cloning, DAgger
Foundational concepts for reinforcement learning applied to language models — MDPs, value functions, and imitation learning.
-
HICRA: Hierarchical Credit Assignment for LLM Reasoning
Emergent two-phase hierarchy in RL training — fix procedural tokens early, then optimize strategic planning tokens.