- llm
- nlp
- reasoning
- rl
- rlhf
- ssm
- theory
- information-theory
- representation-learning
- optimization
- rag
- alignment
- paper-review
- survey
- notes
•
•
•
•
•
•
•
•
•
•
•
•
•
•
-
When Memorizing Irrelevant Data Becomes Necessary
Studying conditions under which memorizing irrelevant training examples is necessary for high accuracy.
-
Chain-of-Thought Prompting: Key Papers and Variants
A reading list of important Chain-of-Thought papers and variants — decomposition, plan-and-solve, faithfulness.
-
GPT-3: Language Models Are Few-Shot Learners
Review of GPT-3 — 175B parameters, in-context learning, and few-shot generalization without gradient updates.
-
GPT-2: Language Models Are Unsupervised Multitask Learners
Review of GPT-2 — scaling autoregressive language models reveals emergent zero-shot task performance.
-
Auto-Regressive Next-Token Predictors Are Universal Learners
Theoretical result showing that next-token prediction is a universal learning objective for sequence functions.