Zixing “Elwood” Lei · PhD Student, Shanghai Jiao Tong University · Google Scholar ↗

Agents that improve, explore, and act.

Supervised by Prof. Siheng Chen, I focus on Recursive Self-Improvement, Reinforcement Learning, and Embodied AI. I believe that using powerful agents to reshape today’s scientific and engineering R&D pipelines is a critical path toward a superintelligent future.

Explore by direction

Three directions. One path toward autonomous intelligence.

RSI

Long-horizon autonomy

RSI systems and context management support agentic, long-horizon autonomous research.

RL

Learning through exploration

RL environments and algorithms let agents explore autonomously and acquire efficient, learnable signals.

Embodied AI

Verifiable physical intelligence

Embodied AI grounds agents in the physical world, where intelligence can be tested and verified.

Latest research

Embodied AI
2026

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Y. Wang, P. Lin, X.-H. Chen, et al.

A scalable pipeline that turns egocentric human manipulation videos into 18,561 hours of robot training data across 15 morphologies, improving out-of-distribution generalization for vision-language-action models.

Read paper ↗
Embodied AI
2026

Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

J. Zhang*, G. Zhou*, H. Yin*, Y. Huang*, Z. Lei*, Q. Peng*, H. Yuan, J. Zhang, X. Guo, X. Chen, A. Yang, F. Huang, Z. Yang, J. Lin, D. Liu, J. Zhou, Z. Yu, J. Fan, Z. Liang, P. Lin, Y. Wang, H. Li, A. Chen, K. Yan, X. Xu, J. Li, L. Hu, M. Zhang, S. Li, W. Xiao, S. Bai, X. Ren, C. Lv, C. Wu, X.-H. Chen. * Equal contribution

A scalable navigation model with a reconfigurable observation strategy, allowing a high-level planner to switch task modes and context strategies during long-horizon execution.

Read paper ↗
Embodied AI
2026

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

H. Yuan*, Z. Liang*, A. Chen*, Y. Wang*, H. Li*, P. Lin*, Y. Huang*, Z. Lei*, T. Zhang*, J. Zhang, J. Zhang, J. Fan, G. Zhou, Q. Peng, C. Lv, X. Chen, A. Yang, F. Huang, J. Lin, D. Liu, J. Zhou, C. Wu, X.-H. Chen. * Equal contribution

A generalizable vision-language-action foundation model that aligns heterogeneous manipulation data across representation, motion, and behavior to enable large-scale training and cross-embodiment transfer.

Read paper ↗
Embodied AI
2026

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

J. Zhang, X. Chen, A. Chen, et al.

A language-conditioned video world model that predicts physically grounded futures across manipulation, autonomous driving, navigation, and human-to-robot transfer.

Read paper ↗
Embodied AI
2026

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Q. Wang, M. Li, J. Guan, et al.

A unified embodied foundation model for manipulation, navigation, and trajectory prediction, extending vision-language reasoning into continuous action generation across robot embodiments.

Read paper ↗
Embodied AI
2026

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models

Z. Lei, C. Liu, Y. Xiong, et al.

VLAs-as-Tools separates temporal reasoning from local physical execution, enabling high-level agents to plan, select specialized VLA tools, monitor progress, and recover over long horizons.

Read paper ↗