News
Selected Publications
For the full publication list, see Google Scholar.
|
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
Howard H. Qian, Yiting Chen, Yunfei Xie, Kejia Ren, Podshara Chanrungmaneekul, Gaotian Wang, Bowen Wen, Chen Wei, Kaiyu Hang
Preprint (under review), 2026
paper / website
TL;DR: We retrieve manipulation-relevant human demonstrations from 100,000 web video clips using 3D hand-motion trajectories, improving dexterous-hand RL policies in Isaac Gym.
|
|
Sensorimotor Alignment: A Training-Free Predictor of Robot Policy Success
Yunfei Xie, Howard H. Qian, Shiyi Lan, Kaiyu Hang, Chen Wei
Preprint (under review), 2026
website
TL;DR: We compare model representations with recorded motion trajectories to assess robot-policy and backbone capability using a training-free offline score.
|
|
Play to Generalize: Learning to Reason Through Game Play
Yunfei Xie, Yinsong Ma, Shiyi Lan, Alan Yuille, Junfei Xiao, Chen Wei
ICLR 2026
website / paper
TL;DR: We show that RL post-training on simple arcade games improves multimodal LLMs’ out-of-domain math and spatial reasoning without worked solutions or diagrams during that post-training stage.
|
|
How Off-Policy Can GRPO Be? µ-GRPO for Efficient LLM Reinforcement Learning
Minghao Tian, Yunfei Xie, Chen Wei
Preprint (under review), 2026
website / paper
TL;DR: We reuse stale rollouts in a few large generation and optimization stages, matching standard GRPO’s reasoning performance at lower training cost.
|
|
MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu, Xianhang Li, Hong-Yu Zhou, Sheng Liu, Lei Xing, James Zou, Cihang Xie, Yuyin Zhou
ICLR 2025. (400+ GitHub stars)
website / paper
TL;DR: We build an automated grounding, annotation, and filtering pipeline for multimodal medical pretraining, covering 25 million images across 10 modalities.
|
|
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
Yunfei Xie, Cihang Xie, Alan Yuille, Jieru Mei
ECCV 2024
paper
TL;DR: We use local and global aggregation in a hierarchical vision transformer to improve part and object segmentation.
|
|
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
Jiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang, Mude Hui, Bingjie Xu, Zeyu Zheng, Zirui Wang, Cihang Xie, Yuyin Zhou
ICLR 2026. (900+ GitHub stars)
website / paper / code
TL;DR: We improve consistency in long-story visualization with a training-free method that refines story images using global reference context from the previous iteration.
|
|
MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games
Yunfei Xie, Kevin Wang, Bobby Cheng, Jianzhu Yao, Zhizhou Sha, Alexander Duffy, Yihan Xi, Hongyuan Mei, Cheston Tan, Chen Wei, Pramod Viswanath, Zhangyang Wang
ICML 2026
paper / website / code
TL;DR: We improve multi-agent LLM game performance and stability through self-play, persistent memory, and tournament-style prompt evolution.
|
Adobe Research |
May 2026 - Aug. 2026: Research Intern, Adobe Research
Advisor: Dr. Branislav Kveton
RL for Flow Matching: Developed a reinforcement learning algorithm for post-training flow-matching generative models, applied to Adobe’s Firefly foundation model. |
|
Feb. 2025 - Present: Research Assistant, Rice University
Advisor: Prof. Chen Wei; robot learning in collaboration with Prof. Kaiyu Hang’s RobotPI Lab
Focus: Human-video learning and offline robot-model evaluation (RoboTok, S²), RL post-training (Play to Generalize, µ-GRPO), and self-improving agents (MEMO). |
|
Dec. 2023 - Feb. 2025: Research Intern, VLAA Lab, University of California, Santa Cruz
Advisors: Prof. Yuyin Zhou and Prof. Cihang Xie
Focus: Multimodal data synthesis and annotation (MedTrinity-25M), and long-story visualization (Story-Iter). |
|
Jul. 2023 - Dec. 2023: Research Intern, CCVL Lab, Johns Hopkins University
Advisor: Prof. Alan Yuille
Focus: Hierarchical vision transformers for part and object segmentation (From Pixels to Objects). |
Honors and Awards
- Aug. 2025: Lambda Research Grant ($10,000)
- 2022: Science and Technology Scholarship, Huazhong University of Science and Technology (top 2% in school)
Invited Talks
Dec. 2025: “Play to Generalize: Learning to Reason Through Game Play,” Open AGI Symposium, NeurIPS 2025.
Professional Skills
ML Frameworks: PyTorch, verl, OpenRLHF, vLLM, DeepSpeed, FSDP, Slurm
Robotics: ROS, MuJoCo, NVIDIA Isaac Gym, SAPIEN
Languages & Tools: Python, Shell, Git, Docker, Linux
|
2026: ICLR, ICML, CVPR, ECCV
2025: ICCV, ICLR, CVPR, ICML, TMM, NeurIPS, TPAMI
2024: CVPR, ICML, IEEE ISBI
|
This website template was borrowed from Jon Barron.
Last updated on September 6, 2026.
|
|