Motion Capture with Millimeter-Wave Tags
. Motion Capture with Millimeter-Wave Tags. SenSys, 2026.
Let it Cook: Learning to Wait in Sequential Decision Making
. Let it Cook: Learning to Wait in Sequential Decision Making. Reinforcement Learning Conference (RLC), 2026.
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
. TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance. ICML, 2026.
πŸ† ICML 2026 Spotlight
Expanding Spatial and Temporal Context for Robotic Imitation Learning Policies With Scene Graphs
. Expanding Spatial and Temporal Context for Robotic Imitation Learning Policies With Scene Graphs. CVPR, 2026.
VLMgineer: Vision-Language Models as Robotic Toolsmiths
. VLMgineer: Vision-Language Models as Robotic Toolsmiths. ICLR, 2026.
Correspondence-Driven Trajectory Warping for Data-Efficient Imitation and Autonomous Play
. Correspondence-Driven Trajectory Warping for Data-Efficient Imitation and Autonomous Play. ICLR, 2026.
TiPToP: A Modular Open-Vocabulary Planning System for Robotic Manipulation
. TiPToP: A Modular Open-Vocabulary Planning System for Robotic Manipulation. arXiv preprint arXiv:2603.09971, 2026.
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
. OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies. 2026.
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
. Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots. (Under review), 2026.
Real-World Reinforcement Learning of Interactive Perception Behaviors
. Real-World Reinforcement Learning of Interactive Perception Behaviors. NeurIPS, 2025.
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
. RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies. CORL, 2025.
πŸ† Oral Presentation
RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
. RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models. CORL, 2025.
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
. ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos. ICRA, 2025.
πŸ† Best Paper, Workshop on 3D Vision Language Models for Robotic Manipulation at CVPR 2025
Leveraging Symmetry to Accelerate Learning of Trajectory Tracking Controllers for Free-Flying Robotic Systems
. Leveraging Symmetry to Accelerate Learning of Trajectory Tracking Controllers for Free-Flying Robotic Systems. ICRA, 2025.
πŸ† Best Paper Award in the Neuroscience and Interpretability track, NeurReps Workshop at NeurIPS 2024
Vision Language Models are In-Context Value Learners
. Vision Language Models are In-Context Value Learners. ICLR, 2025.
The Value of Sensory Information to a Robot
. The Value of Sensory Information to a Robot. ICLR, 2025.
REGENT: A Retrieval-Augmented Generalist Agent That Can Act In-Context in New Environments
. REGENT: A Retrieval-Augmented Generalist Agent That Can Act In-Context in New Environments. ICLR, 2025.
πŸ† Oral Presentation, 1.82% accept rate
Learning to Achieve Goals with Belief State Transformers
. Learning to Achieve Goals with Belief State Transformers. ICLR, 2025.
Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model
. Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model. ICLR, 2025.
Illustrated Landmark Graphs for Long-Horizons Policy Learning
. Illustrated Landmark Graphs for Long-Horizons Policy Learning. TMLR, 2025.
Points2Reward: Robotic Manipulation Rewards from Just One Video
. Points2Reward: Robotic Manipulation Rewards from Just One Video. (Under review), 2025.
Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
. Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels. arXiv preprint arXiv:2508.17437, 2025.
Task-Oriented Hierarchical Object Decomposition for Visuomotor Control
. Task-Oriented Hierarchical Object Decomposition for Visuomotor Control. CORL, 2024.
Environment Curriculum Generation via Large Language Models
. Environment Curriculum Generation via Large Language Models. CORL, 2024.
πŸ† Oral Presentation
TLControl: Trajectory and Language Control for Human Motion Synthesis
. TLControl: Trajectory and Language Control for Human Motion Synthesis. ECCV, 2024.
Learning a Meta-Controller for Dynamic Grasping
. Learning a Meta-Controller for Dynamic Grasping. CASE, 2024.
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. RSS, 2024.
DrEureka: Language Model Guided Sim-To-Real Transfer
. DrEureka: Language Model Guided Sim-To-Real Transfer. RSS, 2024.
Universal Visual Decomposer: Long-Horizon Manipulation Made Easy
. Universal Visual Decomposer: Long-Horizon Manipulation Made Easy. ICRA, 2024.
πŸ† Best Paper, CORL 2023 Workshop on Learning Effective Abstractions for Planning; Best Paper Finalist at ICRA 2024
Long-HOT: A Modular Hierarchical Approach for Long-Horizon Object Transport
. Long-HOT: A Modular Hierarchical Approach for Long-Horizon Object Transport. ICRA, 2024.
ZeroFlow: Fast Zero Label Scene Flow via Distillation
. ZeroFlow: Fast Zero Label Scene Flow via Distillation. ICLR, 2024.
Privileged Sensing Scaffolds Reinforcement Learning
. Privileged Sensing Scaffolds Reinforcement Learning. ICLR, 2024.
πŸ† Spotlight Presentation, 5% accept rate
Memory-Consistent Neural Networks for Imitation Learning
. Memory-Consistent Neural Networks for Imitation Learning. ICLR, 2024.
Eureka: Human-Level Reward Design via Coding Large Language Models
. Eureka: Human-Level Reward Design via Coding Large Language Models. ICLR, 2024.
Can Transformers Capture Spatial Relations between Objects?
. Can Transformers Capture Spatial Relations between Objects?. ICLR, 2024.
Training self-learning circuits for power-efficient solutions
. Training self-learning circuits for power-efficient solutions. Applied Physics Letters (APL) Machine Learning, 2024.
Vision-Based Contact Localization Without Touch or Force Sensing
. Vision-Based Contact Localization Without Touch or Force Sensing. CORL, 2023.
Prospective Learning: Principled Extrapolation to the Future
. Prospective Learning: Principled Extrapolation to the Future. Proceedings of The 2nd Conference on Lifelong Learning Agents, 2023.
LIV: Language-Image Representations and Rewards for Robotic Control
. LIV: Language-Image Representations and Rewards for Robotic Control. ICML, 2023.
Learning Policy-Aware Models for Model-Based Reinforcement Learning via Transition Occupancy Matching
. Learning Policy-Aware Models for Model-Based Reinforcement Learning via Transition Occupancy Matching. L4DC, 2023.
VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
. VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training. ICLR, 2023.
πŸ† Spotlight Presentation, 5% accept rate
Planning Goals for Exploration
. Planning Goals for Exploration. ICLR, 2023.
πŸ† Spotlight Presentation, 5% accept rate; Best Paper Award, CoRL 2022 Roboadapt Workshop
Training Robots to Evaluate Robots: Example-Based Interactive Reward Functions for Policy Learning
. Training Robots to Evaluate Robots: Example-Based Interactive Reward Functions for Policy Learning. CORL, 2022.
πŸ† Oral Presentation; Best Paper Award at CORL 2022
How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $ f $-Advantage Regression
. How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $ f $-Advantage Regression. NeurIPS, 2022.
πŸ† Nominated for Outstanding Paper Award at NeurIPS 2022
Discovering Deformable Keypoint Pyramids
. Discovering Deformable Keypoint Pyramids. ECCV, 2022.
SMODICE: Versatile Offline Imitation Learning via State Occupancy Matching
. SMODICE: Versatile Offline Imitation Learning via State Occupancy Matching. ICML, 2022.
Fighting Fire with Fire: Avoiding DNN Shortcuts through Priming
. Fighting Fire with Fire: Avoiding DNN Shortcuts through Priming. ICML, 2022.
Know Thyself: Transferable Visuomotor Control Through Robot-Awareness
. Know Thyself: Transferable Visuomotor Control Through Robot-Awareness. ICLR, 2022.
Conservative and Adaptive Penalty for Model-Based Safe Reinforcement Learning
. Conservative and Adaptive Penalty for Model-Based Safe Reinforcement Learning. AAAI, 2022.
Prospective Learning: Back to the Future
. Prospective Learning: Back to the Future. 2022.
Object Representations Guided By Optical Flow
. Object Representations Guided By Optical Flow. NeurIPS 4th Robot Learning Workshop: Self-Supervised and Lifelong Learning, 2021.
Conservative Offline Distributional Reinforcement Learning
. Conservative Offline Distributional Reinforcement Learning. NeurIPS, 2021.
Likelihood-Based Diverse Sampling for Trajectory Forecasting
. Likelihood-Based Diverse Sampling for Trajectory Forecasting. ICCV, 2021.
Embracing the Reconstruction Uncertainty in 3D Human Pose Estimation
. Embracing the Reconstruction Uncertainty in 3D Human Pose Estimation. ICCV, 2021.
Keyframe-focused visual imitation learning
. Keyframe-focused visual imitation learning. ICML, 2021.
How Are Learned Perception-Based Controllers Impacted by the Limits of Robust Control?
. How Are Learned Perception-Based Controllers Impacted by the Limits of Robust Control?. L4DC, 2021.
An exploration of embodied visual exploration
. An exploration of embodied visual exploration. IJCV, 2021.
SMiRL: Surprise Minimizing RL in Dynamic Environments
. SMiRL: Surprise Minimizing RL in Dynamic Environments. ICLR, 2021.
πŸ† Oral Presentation
Long-horizon visual planning with goal-conditioned hierarchical predictors
. Long-horizon visual planning with goal-conditioned hierarchical predictors. NeurIPS, 2020.
Fighting Copycat Agents in Behavioral Cloning from Observation Histories
. Fighting Copycat Agents in Behavioral Cloning from Observation Histories. NeurIPS, 2020.
Model-Based Inverse Reinforcement Learning from Visual Demonstrations
. Model-Based Inverse Reinforcement Learning from Visual Demonstrations. CORL, 2020.
Cautious adaptation for reinforcement learning in safety-critical settings
. Cautious adaptation for reinforcement learning in safety-critical settings. ICML, 2020.
MAVRIC: Morphology-Agnostic Visual Robotic Control
. MAVRIC: Morphology-Agnostic Visual Robotic Control. ICRA and IEEE RA-L, 2020.
Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation
. Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation. ICRA and IEEE RA-L, 2020.
Causal Confusion in Imitation Learning
. Causal Confusion in Imitation Learning. NeurIPS, 2019.
REPLAB: A reproducible low-cost arm benchmark for robotic learning
. REPLAB: A reproducible low-cost arm benchmark for robotic learning. ICRA, 2019.
Manipulation by feel: Touch-based control with deep predictive models
. Manipulation by feel: Touch-based control with deep predictive models. ICRA, 2019.
Emergence of exploratory look-around behaviors through active observation completion
. Emergence of exploratory look-around behaviors through active observation completion. Science Robotics, 2019.
Time-agnostic prediction: Predicting predictable video frames
. Time-agnostic prediction: Predicting predictable video frames. ICLR, 2019.
More Than a Feeling: Learning to Grasp and Regrasp using Vision and Touch
. More Than a Feeling: Learning to Grasp and Regrasp using Vision and Touch. IROS and IEEE RA-L, 2018.
πŸ† Best Paper Award Runner-Up, IEEE Robotics and Automation-Letters (RA-L) 2018
Shapecodes: self-supervised feature learning by lifting views to viewgrids
. Shapecodes: self-supervised feature learning by lifting views to viewgrids. ECCV, 2018.
End-to-end policy learning for active visual categorization
. End-to-end policy learning for active visual categorization. IEEE TPAMI, 2018.
Techniques for rectification of camera arrays
. Techniques for rectification of camera arrays. 2018.
Learning Image Representations Tied to Egomotion from Unlabeled Video
. Learning Image Representations Tied to Egomotion from Unlabeled Video. IJCV Special Issue of Best Papers from ICCV 2015, 2017.
Techniques for improved focusing of camera arrays
. Techniques for improved focusing of camera arrays. 2017.
Divide, share, and conquer: Multi-task attribute learning with selective sharing
. Divide, share, and conquer: Multi-task attribute learning with selective sharing. Visual attributes, 2017.
Pano2Vid: Automatic cinematography for watching 360-degree videos
. Pano2Vid: Automatic cinematography for watching 360-degree videos. ACCV, 2016.
πŸ† Oral Presentation; Best Application Paper Award at ACCV 2016
Object-Centric Representation Learning from Unlabeled Videos
. Object-Centric Representation Learning from Unlabeled Videos. ACCV, 2016.
Look-ahead before you leap: end-to-end active recognition by forecasting the effect of motion
. Look-ahead before you leap: end-to-end active recognition by forecasting the effect of motion. ECCV, 2016.
πŸ† Oral Presentation
Slow and steady feature analysis: higher order temporal coherence in video
. Slow and steady feature analysis: higher order temporal coherence in video. CVPR, 2016.
πŸ† Spotlight Presentation
Learning image representations tied to ego-motion
. Learning image representations tied to ego-motion. ICCV, 2015.
πŸ† Oral Presentation
Zero-shot recognition with unreliable attributes
. Zero-shot recognition with unreliable attributes. NeurIPS, 2014.
Decorrelating semantic visual attributes by resisting the urge to share
. Decorrelating semantic visual attributes by resisting the urge to share. CVPR, 2014.
πŸ† Oral Presentation
Objective quality assessment of multiply distorted images
. Objective quality assessment of multiply distorted images. ASILOMAR Signals, Systems and Computers, 2012.