Yekun Chai

london.jpg

I study language-model pretraining, with a focus on data, learning dynamics, and scaling behavior.

My research asks how model capabilities develop over the course of training, which effects persist across scales, and how reliably smaller runs predict larger ones.

I have contributed to ERNIE 5.0, ERNIE-Code, and StarCoder 2.

News

Aug 21, 2025 Four papers accepted to EMNLP 2025.
May 16, 2025 Curiosity-driven RLHF accepted to ACL 2025. [code]
Jan 23, 2025 MA-RLHF accepted to ICLR 2025. [paper] [code]

Latest Posts

Selected Publications

  1. EMNLP Oral
    On Training Data Influence of GPT Models
    Yekun Chai, Qingyi Liu^Shuohuan WangYu SunQiwei Peng, and Hua Wu
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024
  1. StarCoder 2 and The Stack v2: The Next Generation
    Anton Lozhkov , Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, and 60 more authors
    Feb 2024
  1. ERNIE-Code: Beyond English-Centric Cross-lingual Pretraining for Programming Languages
    Yekun ChaiShuohuan Wang, Chao Pang, Yu SunHao Tian, and Hua Wu
    In Findings of the Association for Computational Linguistics: ACL 2023, Jul 2023
  1. Tokenization Falling Short: On Subword Robustness in Large Language Models
    Yekun Chai , Yewei Fang, Qiwei Peng, and Xuhong Li
    In Findings of the Association for Computational Linguistics: EMNLP 2024, Nov 2024