Imitation Learning

Keyframe-focused visual imitation learning
Keyframe-focused visual imitation learning

Identifying and upsampling important frames from demonstration data can significantly boost imitation learning from histories, and scales easily to complex settings such as autonomous driving from vision.

July 2021

SMiRL: Surprise Minimizing RL in Dynamic Environments
SMiRL: Surprise Minimizing RL in Dynamic Environments

We formulate homeostasis as an intrinsic motivation objective and show interesting emergent behavior from minimizing Bayesian surprise with RL across many environments.

April 2021

Long-horizon visual planning with goal-conditioned hierarchical predictors
Long-horizon visual planning with goal-conditioned hierarchical predictors

To plan towards long-term goals through visual prediction, we propose a model based on two key ideas: (i) predict in a goal-conditioned way to restrict planning only to useful sequences, and (ii) recursively decompose the goal-conditioned prediction task into an increasingly fine series of subgoals.

December 2020

Fighting Copycat Agents in Behavioral Cloning from Observation Histories
Fighting Copycat Agents in Behavioral Cloning from Observation Histories

December 2020

Model-Based Inverse Reinforcement Learning from Visual Demonstrations
Model-Based Inverse Reinforcement Learning from Visual Demonstrations

We learn reward functions in unsupervised object keypoint space, to allow us to follow third-person demonstrations with model-based RL.

November 2020

Causal Confusion in Imitation Learning
Causal Confusion in Imitation Learning

"Causal confusion", where spurious correlates are mistaken to be causes of expert actions, is commonly prevalent in imitation learning, leading to counterintuitive results where additional information can lead to worse task performance. How might one address this?

December 2019

Time-agnostic prediction: Predicting predictable video frames
Time-agnostic prediction: Predicting predictable video frames

In visual prediction tasks, letting your predictive model choose which times to predict does two things: (i) improves prediction quality, and (ii) leads to semantically coherent "bottleneck state" predictions, which are useful for planning.

April 2019