| Dec 11, 2023 | Sliced Mutual Information for Memorization and Generalization |
| Dec 11, 2023 | When Memorizing Irrelevant Data Becomes Necessary |
| Dec 11, 2023 | Chain-of-Thought Prompting: Key Papers and Variants |
| Oct 09, 2023 | GPT-3: Language Models Are Few-Shot Learners |
| Oct 09, 2023 | GPT-2: Language Models Are Unsupervised Multitask Learners |
| Oct 09, 2023 | Auto-Regressive Next-Token Predictors Are Universal Learners |
| Sep 06, 2023 | Memorization Without Overfitting in Large Language Models |
| Sep 03, 2023 | Beyond Chain-of-Thought: Graph-of-Thought Reasoning in LLMs |
| Sep 02, 2023 | Scaling Laws for Neural Language Models |
| Aug 30, 2023 | Graph of Thought: Boosting Logical Reasoning in LLMs |
| Aug 29, 2023 | Knowledge Graph Prompting Sparks Graph of Thoughts in LLMs |
| Aug 28, 2023 | Graph of Thoughts: Solving Elaborate Problems with LLMs |
| Aug 27, 2023 | Sparks of AGI: Early Experiments with GPT-4 |
| Jul 27, 2023 | Neural Tangent Kernel: Infinite-Width Networks as Kernel Methods |
| Jul 27, 2023 | Mini-batch Optimization of Contrastive Loss |
| Jul 24, 2023 | Frequency Effects on Syntactic Rule Learning in Transformers |
| Jul 23, 2023 | Why Mask Reconstruction Pretraining Helps in Downstream Tasks |
| Jul 22, 2023 | Emergent Abilities of Large Language Models |
| Jul 18, 2023 | Contextual Representation Learning beyond Masked Language Modeling |
| Jul 17, 2023 | AMOM: Adaptive Masking over Masking (AAAI 2023) |
| Jul 16, 2023 | Mask More and Mask Later (ACL 2022) |
| Jul 15, 2023 | Calibration, Entropy Rates, and Memory in Language Models |
| Jul 13, 2023 | A Closer Look at How Fine-tuning Changes BERT |
| Jul 12, 2023 | AdaGDA: Faster Adaptive Gradient Descent Ascent for Minimax Optimization |
| Jul 11, 2023 | Masked Latent Semantic Modeling (MLSM) |
| Jul 09, 2023 | Singular Value Representation: A Graph Perspective on Neural Networks |
| Jul 07, 2023 | Blessing of Class Diversity in Pre-training |
| Jul 06, 2023 | Deriving Language Models from Masked Language Models |
| Feb 01, 2023 | Neural Collapse: Terminal Phase of Deep Network Training |