[Survey] Recent approaches on Super-Alignment
A collection of recent approaches and papers about super-alignment and relevant topics.
Key Papers
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision — OpenAI
- Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning
- Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
RL-Based Approaches
- PPO: Proximal Policy Optimization Algorithms
- Deep Reinforcement Learning from Human Preferences
- Learning to Summarize from Human Feedback
- Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
- Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
Principles & Theory
- Understanding the Learning Dynamics of Alignment with Human Feedback
- On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
Learning Algorithms
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
- The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
Other Approaches
- Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
- Tuna: Instruction Tuning using Feedback from Large Language Models
- Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
- Rethinking Information Structures in RLHF: Reward Generalization from a Graph Theory Perspective
- Weak-to-Strong Jailbreaking on Large Language Models
Enjoy Reading This Article?
Here are some more articles you might like to read next: