- llm
- nlp
- reasoning
- rl
- rlhf
- ssm
- theory
- information-theory
- representation-learning
- optimization
- rag
- alignment
- paper-review
- survey
- notes
•
•
•
•
•
•
•
•
•
•
•
•
•
•
-
Why Mask Reconstruction Pretraining Helps in Downstream Tasks
Theoretical analysis of why masked reconstruction (MAE-style) pretraining improves downstream performance.
-
Emergent Abilities of Large Language Models
Capabilities that appear discontinuously at scale — and the debate over whether emergence is a real phenomenon or measurement artifact.
-
Contextual Representation Learning beyond Masked Language Modeling
ACL 2022 — extending contextual representation learning beyond the standard MLM objective.
-
AMOM: Adaptive Masking over Masking (AAAI 2023)
Adaptive masking for Conditional Masked Language Models — improving seq2seq decoding with adaptive X- and Y-masking.
-
Mask More and Mask Later (ACL 2022)
Efficiency improvements to MLM pre-training by restructuring which information flows dominate.