Computer Vision / Human Motion / Multimodal Generation

Yichen Peng

I am a research assistant professor at the Koike Lab, Institute of Science Tokyo (ex. Tokyo Institute of Technology), working with Prof. Hideki Koike, and Prof. Erwin Wu. I also serve as visiting faculty at Sony CSL and Keio University Graduate School of Science and Technology. Currently, I am visiting KLab, advised by Prof. Kris Kitani, Carnegie Mellon University. My research focuses on computer vision, human motion understanding, and multimodal generative models, with applications in skill training and human motion generation.

I received my Master and Ph.D. in Computer Science from the Japan Advanced Institute of Science and Technology (JAIST), supervised by Prof. Kazunori Miyata. (Funded by JST-SPRING) I worked closely with Prof. Haoran Xie, and Prof. Tsukasa Fukasato in Waseda University, on sketch-based 2D/3D Generation & Interface.

Yichen Peng

Research Interests

My research focuses on computer vision, human motion understanding (Pose Estimation/Motion Capture), and multimodal generative models, with applications in skill training (Piano/Skiing/Golf) and human motion generation (Speech/Conversation Gesture). I am also interested in sketch-based 2D/3D Generation & Interface.

News

Publications (Selected)

InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation

InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation

*Yichen Peng, *Jyun-Ting Song, *Chen-Chieh Liao, Kris Kitani, Hideki Koike, Erwin Wu

The 19th European Conference on Computer Vision (ECCV2026), 2026

Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion

Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion

Chen-Chieh Liao, Yichen Peng, Yiyi Cai, Yui Ono, Hiroki Hanaoka, Erwin Wu, Hideki Koike, Shuichi Kurabayashi

The 19th European Conference on Computer Vision (ECCV2026), 2026

DyaDiT: A Multi-Modal Diffusion Transformer for Socially-Aware Dyadic Gesture Generation

DyaDiT: A Multi-Modal Diffusion Transformer for Socially-Aware Dyadic Gesture Generation

Yichen Peng, Jyun-Ting Song, Siyeol Jung, Ruofan Liu, Haiyang Liu, Xuangeng Chu, Ruicong Liu, Erwin Wu, Hideki Koike, Kris Kitani

Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR2026), 2026

UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking

UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking

Xuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu, Yichen Peng, Bo Zheng

Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR2026), 2026

SoleCoach: Sole Pressure and IMU-based MLLMs for Skill Coaching

SoleCoach: Sole Pressure and IMU-based MLLMs for Skill Coaching

Toshihiro Hirano, Hitoshi Yoshihara, Yichen Peng, Chen-Chieh Liao, Erwin Wu, Hideki Koike

Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI) (Honorable Mention), 2026

From Pose to Muscle: Multimodal Learning for Piano Hand Muscle Electromyography

From Pose to Muscle: Multimodal Learning for Piano Hand Muscle Electromyography

Ruofan Liu, Yichen Peng, Takanori Oku, Chen-Chieh Liao, Erwin Wu, Shinichi Furuya, Hideki Koike

Conference on Neural Information Processing Systems (NeurIPS 2025), 2025

PiaMuscle: Improving Piano Skill Acquisition by Cost-effectively Estimating and Visualizing Activities of Miniature Hand Muscles

PiaMuscle: Improving Piano Skill Acquisition by Cost-effectively Estimating and Visualizing Activities of Miniature Hand Muscles

Ruofan Liu, Yichen Peng, Takanori Oku, Chen-Chieh Liao, Erwin Wu, Shinichi Furuya, Hideki Koike

Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI), 2025

LineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Model

LineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Model

Xi Wang, Hongzhen Li, Heng Fang, Yichen Peng, Haoran Xie, Xi Yang, Chuntao Li

Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR2025), 2025

Dual-modal 3d human pose estimation using insole foot pressure sensors

Dual-modal 3d human pose estimation using insole foot pressure sensors

Erwin Wu, Yichen Peng, Rawal Khirodkar, Hideki Koike, Kris Kitani

2024 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR 2024 workshop), 2024

Emage: Towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling

Emage: Towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling

Haiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng, Mingyang Su, You Zhou, Xuefei Zhe, Naoya Iwamoto, Bo Zheng, Michael J Black

Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR2024), 2024

Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis

Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis

Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng,, Zhengqing Li, You Zhou, Elif Bozkurt, Bo Zheng

European conference on computer vision (ECCV2022), 2024

DiffFaceSketch: Sketch-guided latent diffusion model for high-fidelity face image synthesis

DiffFaceSketch: Sketch-guided latent diffusion model for high-fidelity face image synthesis

Yichen Peng, Chunqi Zhao, Haoran Xie, Tsukasa Fukusato, Kazunori Miyata

IEEE Access 2023, 2023

Dualmotion: Global-to-local casual motion design for character animations

Dualmotion: Global-to-local casual motion design for character animations

Yichen Peng, Chunqi Zhao, Haoran Xie, Tsukasa Fukusato, Kazunori Miyata, Takeo Igarashi

IEICE TRANSACTIONS on Information and Systems (IEICE 2023), 2023