Yekun Chai

london.jpg

I study language-model pretraining, with a focus on data, learning dynamics, and scaling behavior.

My research asks how model capabilities develop over the course of training, which effects persist across scales, and how reliably smaller runs predict larger ones.

I have contributed to ERNIE, ERNIE-Code, and StarCoder 2.

News

Aug 21, 2025 Four papers accepted to EMNLP 2025.
May 16, 2025 Curiosity-driven RLHF accepted to ACL 2025. [code]
Jan 23, 2025 MA-RLHF accepted to ICLR 2025. [paper] [code]

Selected Publications

  1. EMNLP Oral
    On Training Data Influence of GPT Models
    Yekun Chai, Qingyi Liu^Shuohuan WangYu SunQiwei Peng, and Hua Wu
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024
  1. ERNIE-Code: Beyond English-Centric Cross-lingual Pretraining for Programming Languages
    Yekun ChaiShuohuan Wang, Chao Pang, Yu SunHao Tian, and Hua Wu
    In Findings of the Association for Computational Linguistics: ACL 2023, Jul 2023
  1. Tokenization Falling Short: On Subword Robustness in Large Language Models
    Yekun Chai , Yewei Fang, Qiwei Peng, and Xuhong Li
    In Findings of the Association for Computational Linguistics: EMNLP 2024, Nov 2024
  1. Autoregressive Pre-Training on Pixels and Texts
    Yekun Chai, Qingyi Liu^, Jingwu Xiao^Shuohuan WangYu Sun, and Hua Wu
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024
  1. ICLR Spotlight
    Tool-Augmented Reward Modeling
    Lei Li*^Yekun Chai*†Shuohuan WangYu SunHao TianNingyu Zhang, and 1 more author
    In The Twelfth International Conference on Learning Representations(top 5%) , 2024