Yunfei Xie | 谢云飞

I'm a CS PhD student in the Department of Computer Science at Rice University, advised by Prof. Chen Wei. My research focuses on robot learning from human video and post-training for foundation models.

I completed my bachelor's degree in the School of Artificial Intelligence and Automation at Huazhong University of Science & Technology. Previously, I was a research intern at Adobe Research, advised by Dr. Branislav Kveton, at VLAA, UC Santa Cruz, advised by Prof. Yuyin Zhou and Prof. Cihang Xie, and at CCVL, Johns Hopkins University, advised by Prof. Alan Yuille.

I am open for collaborations in research.

Email  /  Google Scholar  /  Github  /  Twitter  /  CV

profile photo Photo credit: my partner Xuan

News

Selected Publications

For the full publication list, see Google Scholar.

RoboTok human manipulation videos and reconstructed hand trajectories RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
Howard H. Qian, Yiting Chen, Yunfei Xie, Kejia Ren, Podshara Chanrungmaneekul, Gaotian Wang, Bowen Wen, Chen Wei, Kaiyu Hang
Preprint (under review), 2026
paper / website

TL;DR: We retrieve manipulation-relevant human demonstrations from 100,000 web video clips using 3D hand-motion trajectories, improving dexterous-hand RL policies in Isaac Gym.

Sensorimotor Alignment connects visual observations, demonstrated motion, and policy success Sensorimotor Alignment: A Training-Free Predictor of Robot Policy Success
Yunfei Xie, Howard H. Qian, Shiyi Lan, Kaiyu Hang, Chen Wei
Preprint (under review), 2026
website

TL;DR: We compare model representations with recorded motion trajectories to assess robot-policy and backbone capability using a training-free offline score.

Visual Game Learning method overview Play to Generalize: Learning to Reason Through Game Play
Yunfei Xie, Yinsong Ma, Shiyi Lan, Alan Yuille, Junfei Xiao, Chen Wei
ICLR 2026
website / paper

TL;DR: We show that RL post-training on simple arcade games improves multimodal LLMs’ out-of-domain math and spatial reasoning without worked solutions or diagrams during that post-training stage.

Training time comparison for GRPO and µ-GRPO How Off-Policy Can GRPO Be? µ-GRPO for Efficient LLM Reinforcement Learning
Minghao Tian, Yunfei Xie, Chen Wei
Preprint (under review), 2026
website / paper

TL;DR: We reuse stale rollouts in a few large generation and optimization stages, matching standard GRPO’s reasoning performance at lower training cost.

MedTrinity-25M multimodal data pipeline MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu, Xianhang Li, Hong-Yu Zhou, Sheng Liu, Lei Xing, James Zou, Cihang Xie, Yuyin Zhou
ICLR 2025. (400+ GitHub stars)
website / paper

TL;DR: We build an automated grounding, annotation, and filtering pipeline for multimodal medical pretraining, covering 25 million images across 10 modalities.

Hierarchical part and object segmentation From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
Yunfei Xie, Cihang Xie, Alan Yuille, Jieru Mei
ECCV 2024
paper

TL;DR: We use local and global aggregation in a hierarchical vision transformer to improve part and object segmentation.

Story-Iter long-story visualization examples Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
Jiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang, Mude Hui, Bingjie Xu, Zeyu Zheng, Zirui Wang, Cihang Xie, Yuyin Zhou
ICLR 2026. (900+ GitHub stars)
website / paper / code

TL;DR: We improve consistency in long-story visualization with a training-free method that refines story images using global reference context from the previous iteration.

MEMO context optimization framework MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games
Yunfei Xie, Kevin Wang, Bobby Cheng, Jianzhu Yao, Zhizhou Sha, Alexander Duffy, Yihan Xi, Hongyuan Mei, Cheston Tan, Chen Wei, Pramod Viswanath, Zhangyang Wang
ICML 2026
paper / website / code

TL;DR: We improve multi-agent LLM game performance and stability through self-play, persistent memory, and tournament-style prompt evolution.

Education

Ph.D. in Computer Science, Rice University
Sep. 2025 - 2029 (expected), Houston, TX
Advisor: Prof. Chen Wei
B.Eng. in Artificial Intelligence, Huazhong University of Science & Technology
Sep. 2021 - Jun. 2025, Wuhan, China

Experience

Adobe
Research

May 2026 - Aug. 2026: Research Intern, Adobe Research
Advisor: Dr. Branislav Kveton
RL for Flow Matching: Developed a reinforcement learning algorithm for post-training flow-matching generative models, applied to Adobe’s Firefly foundation model.

Rice University

Feb. 2025 - Present: Research Assistant, Rice University
Advisor: Prof. Chen Wei; robot learning in collaboration with Prof. Kaiyu Hang’s RobotPI Lab
Focus: Human-video learning and offline robot-model evaluation (RoboTok, S²), RL post-training (Play to Generalize, µ-GRPO), and self-improving agents (MEMO).

UC Santa Cruz

Dec. 2023 - Feb. 2025: Research Intern, VLAA Lab, University of California, Santa Cruz
Advisors: Prof. Yuyin Zhou and Prof. Cihang Xie
Focus: Multimodal data synthesis and annotation (MedTrinity-25M), and long-story visualization (Story-Iter).

Johns Hopkins University

Jul. 2023 - Dec. 2023: Research Intern, CCVL Lab, Johns Hopkins University
Advisor: Prof. Alan Yuille
Focus: Hierarchical vision transformers for part and object segmentation (From Pixels to Objects).

Honors and Awards

  • Aug. 2025: Lambda Research Grant ($10,000)
  • 2022: Science and Technology Scholarship, Huazhong University of Science and Technology (top 2% in school)

Invited Talks

Dec. 2025: “Play to Generalize: Learning to Reason Through Game Play,” Open AGI Symposium, NeurIPS 2025.

Professional Skills

ML Frameworks: PyTorch, verl, OpenRLHF, vLLM, DeepSpeed, FSDP, Slurm
Robotics: ROS, MuJoCo, NVIDIA Isaac Gym, SAPIEN
Languages & Tools: Python, Shell, Git, Docker, Linux

Reviewer

2026: ICLR, ICML, CVPR, ECCV
2025: ICCV, ICLR, CVPR, ICML, TMM, NeurIPS, TPAMI
2024: CVPR, ICML, IEEE ISBI


This website template was borrowed from Jon Barron.
Last updated on September 6, 2026.