publications

Pretraining & Scaling

  1. ERNIE 5.0 Technical Report
    Baidu ERNIE
    arXiv, 2026
  2. ERNIE 4.5 Technical Report
    Baidu ERNIE
    2025
  3. Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code
    Taishi Nakamura, Mayank Mishra, Simone Tedeschi, Yekun Chai, Jason T. Stillerman, Felix Friedrich, and 35 more authors
    In Proceedings of the 31st International Conference on Computational Linguistics: Industry Track, Jan 2025
  4. StarCoder 2 and The Stack v2: The Next Generation
    Anton Lozhkov , Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, and 60 more authors
    Feb 2024
  5. Autoregressive Pre-Training on Pixels and Texts
    Yekun Chai, Qingyi Liu^, Jingwu Xiao^Shuohuan WangYu Sun, and Hua Wu
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024
  6. ERNIE-Code: Beyond English-Centric Cross-lingual Pretraining for Programming Languages
    Yekun ChaiShuohuan Wang, Chao Pang, Yu SunHao Tian, and Hua Wu
    In Findings of the Association for Computational Linguistics: ACL 2023, Jul 2023

Evaluation & Analysis

  1. CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
    Yilun Yang^, and Yekun Chai
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Nov 2025
  2. Understanding Subword Compositionality of Large Language Models
    Qiwei PengYekun Chai, and Anders Søgaard
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Nov 2025
  3. EvolKV: Evolutionary KV Cache Compression for LLM Inference
    Bohan Yu^, and Yekun Chai
    In Findings of the Association for Computational Linguistics: EMNLP 2025, Nov 2025
  4. EMNLP Oral
    On Training Data Influence of GPT Models
    Yekun Chai, Qingyi Liu^Shuohuan WangYu SunQiwei Peng, and Hua Wu
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024
  5. Tokenization Falling Short: On Subword Robustness in Large Language Models
    Yekun Chai , Yewei Fang, Qiwei Peng, and Xuhong Li
    In Findings of the Association for Computational Linguistics: EMNLP 2024, Nov 2024
  6. GiLOT: Interpreting Generative Language Models via Optimal Transport
    Xuhong Li*, Jiamin Chen*Yekun Chai*, and Haoyi Xiong
    In Forty-first International Conference on Machine Learning, Nov 2024
  7. HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
    Qiwei Peng*Yekun Chai*, and Xuhong Li
    In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), May 2024
  8. NeurIPS Datasets and Benchmarks
    M4: A Unified XAI Benchmark for Faithfulness Evaluation of Feature Attribution Methods across Metrics, Modalities and Models
    Xuhong LiMengnan Du, Jiamin Chen, Yekun ChaiHimabindu Lakkaraju, and Haoyi Xiong
    In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, May 2023

Post-training

  1. ACL
    Curiosity-Driven Reinforcement Learning from Human Feedback
    Haoran Sun*^Yekun Chai*†Shuohuan WangYu SunHua Wu , and Haifeng Wang
    In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2025
  2. MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
    Yekun Chai*†Haoran Sun*^Huang FangShuohuan WangYu Sun, and Hua Wu
    In The Thirteenth International Conference on Learning Representations, Jul 2025
  3. ICLR Spotlight
    Tool-Augmented Reward Modeling
    Lei Li*^Yekun Chai*†Shuohuan WangYu SunHao TianNingyu Zhang, and 1 more author
    In The Twelfth International Conference on Learning Representations(top 5%) , Jul 2024

Earlier Work

  1. ICASSP Oral
    Improved Training of Mixture-of-Experts Language GANs
    Yekun ChaiQiyue Yin , and Junge Zhang
    In 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023
  2. Clip-Tuning: Towards Derivative-free Prompt Learning with a Mixture of Rewards
    Yekun ChaiShuohuan WangYu SunHao TianHua Wu , and Haifeng Wang
    In Findings of the Association for Computational Linguistics: EMNLP 2022, Dec 2022
  3. Counter-Contrastive Learning for Language GANs
    Yekun ChaiHaidong ZhangQiyue Yin , and Junge Zhang
    In Findings of the Association for Computational Linguistics: EMNLP 2021, Nov 2021
  4. ACL
    Highway Transformer: Self-Gating Enhanced Self-Attentive Networks
    Yekun Chai, Shuo Jin, and Xinwen Hou
    In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020

* equal contribution    † corresponding author    ^ mentored