- llm
- nlp
- reasoning
- rl
- rlhf
- ssm
- theory
- information-theory
- representation-learning
- optimization
- rag
- alignment
- paper-review
- survey
- notes
•
•
•
•
•
•
•
•
•
•
•
•
•
•
-
Calibration, Entropy Rates, and Memory in Language Models
Why perplexity fails to capture long-term properties — entropy rate calibration and mutual-information memory.
-
A Closer Look at How Fine-tuning Changes BERT
ACL 2022 — empirical analysis of what changes in BERT's internal representations during fine-tuning.
-
AdaGDA: Faster Adaptive Gradient Descent Ascent for Minimax Optimization
Adaptive gradient methods for minimax problems — faster convergence with adaptive step sizes.
-
Masked Latent Semantic Modeling (MLSM)
Sparse coding meets masked language modeling — masking in latent space via knowledge distillation.
-
Singular Value Representation: A Graph Perspective on Neural Networks
Recasting neural network analysis through singular value decomposition and graph-theoretic structure.