Research
We study optimization, model efficiency, data, and technical alignment and safety.
Optimization
Understanding training dynamics and improving optimization.
The Marginal Value of Momentum for Small Learning Rate SGD ICLR 2024
Fine-Tuning Language Models with Just Forward Passes NeurIPS 2023 Oral
A Kernel-Based View of Language Model Fine-Tuning ICML 2023
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms NeurIPS 2022
On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs) NeurIPS 2021
End-to-end model efficiency
Producing the final model efficiently by optimizing the full pipeline, rather than each stage in isolation.
TailSFT: Filtered Fine-Tuning Improves Post-Training Performance Preprint · In submission
The Coverage Principle: How Pre-Training Enables Post-Training ICLR 2026 Oral
Overtrained Language Models are Harder to Fine-Tune ICML 2025
A Mathematical Exploration of Why Language Models Help Solve Downstream Tasks ICLR 2021
Progressive distillation induces an implicit curriculum ICLR 2025
Data
How training data shapes model behavior.
In Good GRACES: Principled Teacher Selection for Knowledge Distillation ICLR 2026
Metadata Conditioning Accelerates Language Model Pre-training ICML 2025
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws ICLR 2025
LESS: Selecting Influential Data for Targeted Instruction Tuning ICML 2024
Technical alignment & safety
Technical approaches to evaluating and improving how models operate in society.