Brown University

Action-driven Learning of Structured Representations for Sequential Decision Making

Description

Abstract:
Generally intelligent agents must learn and adapt by interacting with a complex world. In order to be generally capable of performing diverse tasks in their lifetime they must perceive the world through rich, high-dimensional sensors and have access to adaptable, fine controls. This, however, makes the learning problem intractable. In order to learn and act efficiently, they must use abstractions of state and time: they have to focus only on the relevant information and reason at the right time scale. Traditionally we have provided the problem formulation to our agents; implicitly giving them access to privileged knowledge about the abstractions and structure of the world. However, agents must be able to learn about these by themselves. In this thesis, we will focus on the problem of learning state representations directly from high-dimensional observations and show that agents' actions are the common thread, providing rich learning signal across two axes: abstraction and factorization. First, we focus on state abstraction: the agent must learn a representation that contains only the relevant information for planning. Specifically, we propose an algorithm to learn minimal continuous representations that are sufficient for planning with skills and show empirically that the learned model can be reused effectively to plan for different tasks. Second, we explore learning disentangled representations by discovering underlying factors of variation from raw observations: we introduce a contrastive algorithm that leverages the agent's actions to discover the independently controllable factors directly from pixels. That is, we leverage the agent’s interventions in the dynamics of the world to uncover a signal that disentangles the controllable factors without any prior knowledge. Next, we propose an approach to balance multiple sparsity conditions---action-effect sparsity and temporal-dependency sparsity---to recover the Dynamics Bayesian Network (DBN) by showing that the disentangled representation is the Pareto-optimal solution of a cooperative game between multiple constraints over a shared encoder. Finally, we generalize this idea to mechanism shifts that arise naturally in dynamical systems in which contacts can create new relations in the DBN, providing additional structural signal for disentanglement.
Notes:
Thesis (Ph. D.)--Brown University, 2026

Citation

Rodriguez Sanchez, Rafael Alberto, "Action-driven Learning of Structured Representations for Sequential Decision Making" (2026). Computer Science Theses and Dissertations. Brown Digital Repository. Brown University Library. https://repository.library.brown.edu/studio/item/bdr:ymzvfpjr/

Relations

Collection: