Sheng Jia
I’m an ML PhD candidate at the University of Toronto, advised by Jimmy Ba and Sheila McIlraith. My research is on large language models post-training.
I’m currently an Applied Scientist at Amazon in the Bay Area, working on RL scaling for LLM coding models. Previously, I interned on the post-training team at MiniMax (on M2.5 / M2.7), and worked as an Applied Scientist intern at Amazon.
Interests
- Large Language Models
- Reinforcement Learning
- LLM reasoning
- Agentic coding models
Education
PhD in Computer Science, Expected 2026
University of Toronto
MSc in Computer Science, 2021
University of Toronto
BASc in Engineering Science, 2019
University of Toronto
selected publications
-
EMNLP 2026 (Main)Group Adaptive Clipping Policy OptimizationIn Conference on Empirical Methods in Natural Language Processing (EMNLP), Main Conference, 2026 (15.4% acceptance)Camera-ready & code in preparation