Sheng Jia

prof_pic.jpg

I’m an ML PhD candidate at the University of Toronto, advised by Jimmy Ba and Sheila McIlraith. My research is on large language models post-training.

I’m currently an Applied Scientist at Amazon in the Bay Area, working on RL scaling for LLM coding models. Previously, I interned on the post-training team at MiniMax (on M2.5 / M2.7), and worked as an Applied Scientist intern at Amazon.

Interests

  • Large Language Models
  • Reinforcement Learning
  • LLM reasoning
  • Agentic coding models

Education

PhD in Computer Science, Expected 2026
University of Toronto
MSc in Computer Science, 2021
University of Toronto
BASc in Engineering Science, 2019
University of Toronto

selected publications

  1. EMNLP 2026 (Main)
    Group Adaptive Clipping Policy Optimization
    Sheng Jia, Xiao Wang, Shiva Kasiviswanathan, and Rein Houthooft
    In Conference on Empirical Methods in Natural Language Processing (EMNLP), Main Conference, 2026 (15.4% acceptance)
    Camera-ready & code in preparation
  2. Tech Report
    The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
    MiniMax et al. (incl. Sheng Jia)
    2026
  3. ICLR 2026
    Training Large Language Models to Reason in Parallel with Global Forking Tokens
    Sheng Jia, Xiao Wang, and Shiva Kasiviswanathan
    In International Conference on Learning Representations (ICLR), 2026
  4. ICML 2021
    Efficient Statistical Tests: A Neural Tangent Kernel Approach
    Sheng Jia, Ehsan Nezhadarya, Yuhuai Wu, and Jimmy Ba
    In International Conference on Machine Learning (ICML), 2021
  5. ICLR 2019
    DOM-Q-NET: Grounded Reinforcement Learning on Structured Language
    Sheng Jia, Jamie Kiros, and Jimmy Ba
    In International Conference on Learning Representations (ICLR), 2019