Hey, I am happy you are here

Thanks for stopping by!

I am a Ph.D. student at the Institute of Automation, Chinese Academy of Sciences. My research interests include prompt learning, domain generalization, and representation learning for vision-language-action models.

My current projects explore generalizable action representations for embodied intelligence, including covariant action modeling, semantic-action manifold alignment, and human-to-robot transfer for robotic manipulation. Learn more about my research interests in publications.

Get in touch by sending me a note.

About me

Belief. Intelligence emerges when generalizable multimodal representations become stable.

My current work centers on VLA training and representation learning for embodied intelligence. I am interested in how vision-language-action models acquire compact, transferable action spaces that support robust manipulation and generalization.

I am a Ph.D. student at the Institute of Automation, Chinese Academy of Sciences, supervised by Changsheng Xu. I am broadly open to research collaborations, visiting opportunities, and industry research roles around embodied AI, robot learning, and multimodal foundation models.

2024 - Present

Research Intern

BAAI · Embodied vision-language-action model pretraining

2023 - Present

Ph.D. Student

CASIA · Pattern Recognition and Artificial Intelligence

News

May 2026
Our GAM and LAST have been accepted by ICML. (First Author)
Feb 2026
Our paper H2R-BM has been accepted by Pattern Recognition. (Corresponding Author)
Jul 2025
Our paper EgoPrompt has been accepted by ACM Multimedia. (First Author)
Mar 2025
Our paper on multiple local prompts distillation has been accepted by IEEE Transactions on Multimedia. (First Author)
Feb 2025
Our paper on bi-modality individual-aware prompt tuning has been accepted by IEEE TPAMI.
Oct 2024
I joined BAAI as a research intern working on embodied vision-language-action model pretraining.
Sep 2023
I started my Ph.D. study at the Institute of Automation, Chinese Academy of Sciences.

Lately ...

Teaser diagram for General Covariant Action Modeling

General Covariant Action Modeling

Constructing generalized manifolds for action learning through spatio-temporal decoupling.

See project ->
Diagram for Lie-algebraic Action Space Tokenizer

Lie-algebraic Action Space Tokenizer

Bridging vision-language and action manifolds via Gromov-Wasserstein alignment.

See project ->
Teaser diagram for Human to Robot for Bimanual Manipulation

Human to Robot for Bimanual Manipulation

Leveraging human videos to improve generalization in robotic bimanual manipulation.

See project ->