Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals socially learn from each other. Beyond its biological parallels, this modality provides a practical solution for data scaling in sample-inefficient and data-starved domains like robotics. Recent work has demonstrated promising results in learning manipulation skills from human videos, yet progress in this area remains difficult to assess. Existing methods vary widely in assumptions, hardware choices, and environment setups making it difficult to draw meaningful comparisons and identify advances in the field. To address these challenges, we introduce RoboReel: a unified benchmark for evaluating models that learn policies from human videos. RoboReel consists of bundled real-world human demonstration videos, simulated robot trajectories, and evaluation environments on ten manipulation tasks. We develop four test suites to evaluate the models’ performance on multiple axes, including the robustness to visual distractors and the ability to complete long-horizon tasks. Our benchmark covers learning-from-observation models from different categories, and studies the effectiveness of multiple representation choices in our benchmark evaluation that covers over seven state-of-the-art algorithms (including our VLA based variants) in the field of LfO. Finally, we present an analysis of the different types of algorithms showing that long-horizon tasks and tasks with low tolerances are still challenging for current models.
@misc{gu2026monkeyseemonkeydo,title={Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation},author={Gu, Weiwei and Gupta, Anmol and Sah, Anant and Varghese, Ryan and Vanam, Lalitha Shreya and Adireddi, Prabhath and Karkus, Peter and Gopalan, Nakul},year={2026},url={https://arxiv.org/abs/2609.08209},}
PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations
Anmol Gupta, Weiwei Gu, Omkar Patil, and 2 more authors
Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. Existing approaches often rely on prior knowledge about the objects, such as the number or type of joints. Some of these approaches also fail to recover occluded joints that are only revealed during interaction. Others require large numbers of multi-view images for every object, which is impractical in real-world settings. Furthermore, prior works neglect the order of manipulations, which is essential for many multi-DoF objects where one joint must be operated before another, such as a dishwasher. We introduce PokeNet, an end-to-end framework that estimates articulation models from a single human demonstration without prior object knowledge. Given a sequence of point cloud observations of a human manipulating an unknown object, PokeNet predicts joint parameters, infers manipulation order, and tracks joint states over time. PokeNet outperforms existing state-of-the-art methods, improving joint axis and state estimation accuracy by an average of over 27% across diverse objects, including novel and unseen categories. We demonstrate these gains in both simulation and real-world environments.
@misc{gupta2026pokenetlearningkinematicmodels,title={PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations},author={Gupta, Anmol and Gu, Weiwei and Patil, Omkar and Lee, Jun Ki and Gopalan, Nakul},year={2026},url={https://arxiv.org/abs/2602.02741},}
2025
Continual Robot Skill and Task Learning via Dialogue
Weiwei Gu, N Suresh K Kondepudi, Anmol Gupta, and 2 more authors
In ICRA 2025 Workshop on Foundation Models and Neuro-Symbolic AI for Robotics, 2025
@inproceedings{gu2025continual,title={Continual Robot Skill and Task Learning via Dialogue},author={Gu, Weiwei and Kondepudi, N Suresh K and Gupta, Anmol and Huang, Lixiao and Gopalan, Nakul},booktitle={ICRA 2025 Workshop on Foundation Models and Neuro-Symbolic AI for Robotics},year={2025},url={https://openreview.net/forum?id=BPK2OaKVdy},}
Learning Sequential Kinematic Models from Demonstrations for Multi-Jointed Articulated Objects
Anmol Gupta, Weiwei Gu, Omkar Patil, and 2 more authors
@article{gupta2025learning,title={Learning Sequential Kinematic Models from Demonstrations for Multi-Jointed Articulated Objects},author={Gupta, Anmol and Gu, Weiwei and Patil, Omkar and Lee, Jun Ki and Gopalan, Nakul},journal={arXiv preprint arXiv:2505.06363},year={2025},url={https://arxiv.org/abs/2505.06363},}
2023
Learning AI-System Capabilities under Stochasticity
Pulkit Verma, Rushang Karia, Gaurav Vipat, and 2 more authors
In NeurIPS 2023 Workshop on Generalization in Planning, 2023
@inproceedings{verma2023learning,title={Learning AI-System Capabilities under Stochasticity},author={Verma, Pulkit and Karia, Rushang and Vipat, Gaurav and Gupta, Anmol and Srivastava, Siddharth},booktitle={NeurIPS 2023 Workshop on Generalization in Planning},year={2023},url={https://openreview.net/pdf?id=boub8VqmZu},}