Yongzhong Wang

Embodied Intelligence & Computer Vision

Undergraduate @ SUSTech · Research Assistant @ VIP Lab

Updates

  • September 2026 We submitted our paper Navi-Agent to ICRA 2027!
  • June 2026 Our paper RADAR gets accepted by IROS 2026!
  • May 2026 Welcome to my homepage!

Publications

RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset
Yongzhong Wang*, Keyu Zhu*, Yong Zhong*, Liqiong Wang, Jinyu Yang, Feng Zheng (*Equal contribution)
IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026 (Accepted)

Projects

MOKA: VLM-based Keypoint Planning in Simulation

Oct 2025 – Dec 2025  ·  My RLBench Port  ·  Original MOKA (RSS 2024)

  • Sim-to-Real Reverse Engineering: Independently ported the visual prompting framework MOKA (RSS 2024) from its original hardware-exclusive, ROS-dependent codebase into the RLBench physics simulator, establishing a rigorously reproducible baseline for VLM-driven manipulation.
  • VLM-Simulation Bridge: Engineered a seamless pipeline connecting RLBench camera renders with the GPT-4V API, enabling the generation of Set-of-Mark visual prompts and the automated extraction of semantic "grasp" and "target" keypoints based on natural-language instructions.
  • Benchmarking & Analysis: Successfully deployed this adapted framework to rigorously evaluate multi-step tasks, exposing the inherent limitations of 2D-pixel heuristics in complex contact dynamics, which directly motivated the architecture of our IROS 2026 submission (RADAR).
Annotated keypoints Planned trajectory

Move the crumpled paper into the trash can

Grasp crumpled paper at P3, target trash can at Q1, motion: downward

Failure case: 2D-to-3D coordinate projection error caused the gripper to miss the crumpled paper entirely.

Phone on base annotated Phone on base motion

Move the phone onto its base

Failure case: grasp keypoints P5 incorrectly land on the base instead of the phone, highlighting the limitation of 2D visual prompting in ambiguous object boundaries.

Teleoperation-Free Dexterous Manipulation: Video-to-Sim Retargeting

Mar 2026 – Present  ·  Codebase (In Progress)

  • Custom Digital Twin Integration: Independently engineered a custom simulation environment in ManiSkill by flawlessly merging the disparate URDF models of an RM65 robotic arm and a Revo2 dexterous hand, establishing a robust physical baseline for complex, contact-rich manipulation.
  • End-to-End Kinematic Pipeline: Used a seamless retargeting bridge to translate raw human videos into simulated robot actions. Implemented rigorous SE(3) coordinate transformations (Camera-to-Base mapping) and utilized the Pinocchio library for 6D Inverse Kinematics (IK) to solve the highly non-linear mapping from human MANO 45D joint spaces to the physical constraints of the Revo2 hand.

Experience

  • Jun 2026 – Present Research Assistant, SJTU — Zero-shot vision-language navigation (VLN-CE) with monocular RGB and coordinate-free topological spatial memory
  • Jul 2025 – Sep 2025 Algorithm Intern, Spatialtemporal AI — VLM-based imitation learning retrieval system
  • Sep 2024 – Present Research Assistant, VIP Lab @ SUSTech — Embodied AI closed-loop data generation with VLMs and robot arms
  • Sep 2023 – Present Undergraduate, B.Eng. in Computer Science and Technology, SUSTech

Blog