2023

an archive of posts from this year

Dec 11, 2023 Sliced Mutual Information for Memorization and Generalization
Dec 11, 2023 When Memorizing Irrelevant Data Becomes Necessary
Dec 11, 2023 Chain-of-Thought Prompting: Key Papers and Variants
Oct 09, 2023 GPT-3: Language Models Are Few-Shot Learners
Oct 09, 2023 GPT-2: Language Models Are Unsupervised Multitask Learners
Oct 09, 2023 Auto-Regressive Next-Token Predictors Are Universal Learners
Sep 06, 2023 Memorization Without Overfitting in Large Language Models
Sep 03, 2023 Beyond Chain-of-Thought: Graph-of-Thought Reasoning in LLMs
Sep 02, 2023 Scaling Laws for Neural Language Models
Aug 30, 2023 Graph of Thought: Boosting Logical Reasoning in LLMs
Aug 29, 2023 Knowledge Graph Prompting Sparks Graph of Thoughts in LLMs
Aug 28, 2023 Graph of Thoughts: Solving Elaborate Problems with LLMs
Aug 27, 2023 Sparks of AGI: Early Experiments with GPT-4
Jul 27, 2023 Neural Tangent Kernel: Infinite-Width Networks as Kernel Methods
Jul 27, 2023 Mini-batch Optimization of Contrastive Loss
Jul 24, 2023 Frequency Effects on Syntactic Rule Learning in Transformers
Jul 23, 2023 Why Mask Reconstruction Pretraining Helps in Downstream Tasks
Jul 22, 2023 Emergent Abilities of Large Language Models
Jul 18, 2023 Contextual Representation Learning beyond Masked Language Modeling
Jul 17, 2023 AMOM: Adaptive Masking over Masking (AAAI 2023)
Jul 16, 2023 Mask More and Mask Later (ACL 2022)
Jul 15, 2023 Calibration, Entropy Rates, and Memory in Language Models
Jul 13, 2023 A Closer Look at How Fine-tuning Changes BERT
Jul 12, 2023 AdaGDA: Faster Adaptive Gradient Descent Ascent for Minimax Optimization
Jul 11, 2023 Masked Latent Semantic Modeling (MLSM)
Jul 09, 2023 Singular Value Representation: A Graph Perspective on Neural Networks
Jul 07, 2023 Blessing of Class Diversity in Pre-training
Jul 06, 2023 Deriving Language Models from Masked Language Models
Feb 01, 2023 Neural Collapse: Terminal Phase of Deep Network Training