CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
By Kechen Liu · Paper · cs.RO
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP