Yekun Chai

london.jpg

I study language-model pretraining, focusing on how data and scale shape learning dynamics.

My work seeks general principles for predicting how training decisions affect model capabilities across scales.

I have contributed to ERNIE 5.0, ERNIE-Code, and StarCoder 2.

News

Aug 21, 2025 Four papers accepted to EMNLP 2025.
May 16, 2025 Curiosity-driven RLHF accepted to ACL 2025. [code]
Jan 23, 2025 MA-RLHF accepted to ICLR 2025. [paper] [code]

Latest Posts

Selected Publications

  1. EMNLP Oral
    On Training Data Influence of GPT Models
    Yekun Chai, Qingyi Liu^Shuohuan WangYu SunQiwei Peng, and Hua Wu
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024
  1. StarCoder 2 and The Stack v2: The Next Generation
    Anton Lozhkov , Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, and 60 more authors
    Feb 2024
  1. ERNIE-Code: Beyond English-Centric Cross-lingual Pretraining for Programming Languages
    Yekun ChaiShuohuan Wang, Chao Pang, Yu SunHao Tian, and Hua Wu
    In Findings of the Association for Computational Linguistics: ACL 2023, Jul 2023
  1. Tokenization Falling Short: On Subword Robustness in Large Language Models
    Yekun Chai , Yewei Fang, Qiwei Peng, and Xuhong Li
    In Findings of the Association for Computational Linguistics: EMNLP 2024, Nov 2024