Reinforcement Learning

VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

April 2023

Training Robots to Evaluate Robots: Example-Based Interactive Reward Functions for Policy Learning
Training Robots to Evaluate Robots: Example-Based Interactive Reward Functions for Policy Learning

December 2022

How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $ f $-Advantage Regression
How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $ f $-Advantage Regression

December 2022

SMODICE: Versatile Offline Imitation Learning via State Occupancy Matching
SMODICE: Versatile Offline Imitation Learning via State Occupancy Matching

July 2022

Know Thyself: Transferable Visuomotor Control Through Robot-Awareness
Know Thyself: Transferable Visuomotor Control Through Robot-Awareness

April 2022

Conservative and Adaptive Penalty for Model-Based Safe Reinforcement Learning
Conservative and Adaptive Penalty for Model-Based Safe Reinforcement Learning

February 2022

Prospective Learning: Back to the Future
Prospective Learning: Back to the Future

January 2022

Conservative Offline Distributional Reinforcement Learning
Conservative Offline Distributional Reinforcement Learning

December 2021

Keyframe-focused visual imitation learning
Keyframe-focused visual imitation learning

Identifying and upsampling important frames from demonstration data can significantly boost imitation learning from histories, and scales easily to complex settings such as autonomous driving from vision.

July 2021

How Are Learned Perception-Based Controllers Impacted by the Limits of Robust Control?
How Are Learned Perception-Based Controllers Impacted by the Limits of Robust Control?

We show empirically that the sample complexity and asymptotic performance of learned non-linear controllers in partially observable settings continues to follow theoretical limits based on the difficulty of state estimation

June 2021

An exploration of embodied visual exploration
An exploration of embodied visual exploration

May 2021

SMiRL: Surprise Minimizing RL in Dynamic Environments
SMiRL: Surprise Minimizing RL in Dynamic Environments

We formulate homeostasis as an intrinsic motivation objective and show interesting emergent behavior from minimizing Bayesian surprise with RL across many environments.

April 2021

Long-horizon visual planning with goal-conditioned hierarchical predictors
Long-horizon visual planning with goal-conditioned hierarchical predictors

To plan towards long-term goals through visual prediction, we propose a model based on two key ideas: (i) predict in a goal-conditioned way to restrict planning only to useful sequences, and (ii) recursively decompose the goal-conditioned prediction task into an increasingly fine series of subgoals.

December 2020

Fighting Copycat Agents in Behavioral Cloning from Observation Histories
Fighting Copycat Agents in Behavioral Cloning from Observation Histories

December 2020

Cautious adaptation for reinforcement learning in safety-critical settings
Cautious adaptation for reinforcement learning in safety-critical settings

How to train RL agents safely? We propose to pretrain a model-based agent in a mix of sandbox environments, then plan pessimistically when finetuning in the target environment.

July 2020

Causal Confusion in Imitation Learning
Causal Confusion in Imitation Learning

"Causal confusion", where spurious correlates are mistaken to be causes of expert actions, is commonly prevalent in imitation learning, leading to counterintuitive results where additional information can lead to worse task performance. How might one address this?

December 2019

REPLAB: A reproducible low-cost arm benchmark for robotic learning
REPLAB: A reproducible low-cost arm benchmark for robotic learning

We propose a low-cost compact easily replicable hardware stack for manipulation tasks, that can be assembled within a few hours. We also provide implementations of robot learning algorithms for grasping (supervised learning) and reaching (reinforcement learning). Contributions invited!

May 2019

Emergence of exploratory look-around behaviors through active observation completion
Emergence of exploratory look-around behaviors through active observation completion

May 2019

End-to-end policy learning for active visual categorization
End-to-end policy learning for active visual categorization

Active visual perception with realistic and complex imagery can be formulated as an end-to-end reinforcement learning problem, the solution to which benefits from additionally exploiting the auxiliary task of action-conditioned future prediction.

July 2018

Learning to look around: Intelligently exploring unseen environments for unknown tasks
Learning to look around: Intelligently exploring unseen environments for unknown tasks

Task-agnostic visual exploration policies may be trained through a proxy "observation completion" task that requires an agent to "paint" unobserved views given a small set of observed views.

June 2018

Embodied learning for visual recognition
Embodied learning for visual recognition

January 2017

Look-ahead before you leap: end-to-end active recognition by forecasting the effect of motion
Look-ahead before you leap: end-to-end active recognition by forecasting the effect of motion

Active visual perception with realistic and complex imagery can be formulated as an end-to-end reinforcement learning problem, the solution to which benefits from additionally exploiting the auxiliary task of action-conditioned future prediction.

September 2016