Leitian Tao

I am a Ph.D. student in the Computer Sciences Department at the University of Wisconsin–Madison, advised by Prof. Sharon Li. My research focuses on post-training for agentic systems and reasoning tasks.

I am currently a Research Intern at Microsoft Research, working with Dr. Baolin Peng and Dr. Jianfeng Gao on agentic reinforcement learning and coding-agent self-reflection for scaling.

Email  /  Google Scholar  /  Github  /  LinkedIn  /  CV

profile photo
Research

My research focuses on post-training for agentic systems and reasoning tasks, with an emphasis on agentic reinforcement learning, long-horizon credit assignment, self-improving coding agents, and learning beyond binary verifiable rewards. Representative papers are highlighted.

Experience
Microsoft Research — Research Intern
Feb 2026 – Present · Redmond, WA
Mentors: Dr. Baolin Peng and Dr. Jianfeng Gao. Research on agentic RL and coding-agent self-reflection for scaling.
Meta (Facebook AI Research) — Research Scientist Intern
May 2025 – Dec 2025 · Bellevue, WA
Mentors: Dr. Ping Yu and Dr. Jason Weston. Research on improving LLM reasoning beyond binary verifiable rewards.
Adobe Research — Research Scientist Intern
May 2024 – Dec 2024 · San Jose, CA
Mentors: Dr. Xiang Chen and Dr. Tong Yu. Developed CodeLutra, a preference-learning framework that improves code generation using success and failure signals.
Shanghai AI Lab — Research Intern
Oct 2022 – Mar 2023 · Shanghai, China
Mentor: Dr. Peng Gao. Studied fine-tuning methods for multimodal model alignment.
Preprints
Executable Verification Is All You Need for the Self-Improvement Agentic Loop
Leitian Tao, et al.
Ongoing
TRACE: Turn-Level Reward Assignment via Credit Estimation for Long-Horizon Agents
Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li
Preprint
The Era of Real-World Human Interaction paper The Era of Real-World Human Interaction: RL from User Conversations
Chuanyang Jin, Jing Xu, Bo Liu, Leitian Tao, Olga Golovneva, Tianmin Shu, Wenting Zhao, Xian Li, Jason Weston
Preprint, 2025
[arXiv]
Publications
Hybrid Reinforcement paper Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
Leitian Tao, Ilia Kulikov, Swarnadeep Saha, Tianlu Wang, Jing Xu, Yixuan Li, Jason Weston, Ping Yu
ICLR 2026
[arXiv]  
[BibTeX]
@inproceedings{tao2026hybrid,
  title = {Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense},
  author = {Tao, Leitian and Kulikov, Ilia and Saha, Swarnadeep and Wang, Tianlu and Xu, Jing and Li, Yixuan and Weston, Jason E. and Yu, Ping},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year = {2026}
}
RESTRAIN paper RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization
Zhaoning Yu, Will Su, Leitian Tao, Haozhu Wang, Aashu Singh, Hanchao Yu, Jianyu Wang, Hongyang Gao, Weizhe Yuan, Jason Weston, Ping Yu, Jing Xu
ICLR 2026
[arXiv]  
[BibTeX]
@inproceedings{yu2026restrain,
  title = {RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization},
  author = {Yu, Zhaoning and Su, Will and Tao, Leitian and Wang, Haozhu and Singh, Aashu and Yu, Hanchao and Wang, Jianyu and Gao, Hongyang and Yuan, Weizhe and Weston, Jason and Yu, Ping and Xu, Jing},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year = {2026}
}
Latent Space Synthesis paper Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
Leitian Tao, Xuefeng Du, Yixuan Li
NeurIPS 2025
[arXiv]
Your Weak LLM is Secretly a Strong Teacher for Alignment
Leitian Tao, Yixuan Li
ICLR 2025
CodeLutra: Boosting LLM Code Generation via Preference-Guided Refinement
Leitian Tao, Xiang Chen, Tong Yu, Tung Mai, Ryan Rossi, Yixuan Li, Saayan Mitra
TMLR 2025
Non-Parametric Outlier Synthesis
Leitian Tao, Xuefeng Du, Xiaojin Zhu, Yixuan Li
ICLR 2023
Predicate Correlation Learning for Scene Graph Generation
Leitian Tao, Li Mi, Nannan Li, Xianhang Cheng, Yaosi Hu, Zhenzhong Chen
IEEE TIP
Activate and Reject: Safe Domain Generalization under Category Shift
Chaoqi Chen, Luyao Tang, Leitian Tao, Hong-Yu Zhou, Yue Huang, Xiaoguang Han, Yizhou Yu
ICCV 2023
Position: Challenges and Future Directions of Data-Centric AI Alignment
Min-Hsuan Yeh, Jeffrey Wang, Xuefeng Du, Seongheon Park, Leitian Tao, Shawn Im, Yixuan Li
ICML 2025

Template inspired by Jon Barron.
Last updated: July 2026