An increasing number of works combining neural networks and differential equations for spatio-temporal forecasting have been proposed for the last few years. Some of them show substantial improvements for the prediction of dynamical systems or videos compared to standard RNNs by defining the dynamics using learned ODEs. In this talk, we first present our recent approach for adapting such works for stochastic data. We introduce a novel dynamic model for stochastic video prediction which, unlike prior image-autoregressive models, decouples frame synthesis and dynamics. The dynamics of the model are governed in a latent space by a residual update rule, which is motivated by discretization schemes of differential equations. This endows our method with several desirable properties, such as temporal efficiency and latent space interpretability. Then, we will present a second method, more specifically focused on learning disentangled spatial and temporal representations of spatio-temporal phenomena, with the aim of more accurately predicting future tendencies from initial observations. We propose to model the evolution of partially observed spatiotemporal phenomena with unknown dynamics by taking inspiration from a formal method for the analytical resolution of PDEs: the functional separation of variables. We experimentally demonstrate the performance and broad applicability of our method against prior state-of-the-art models on physical and synthetic video datasets.
This video was produced by the SITE Research Center at New York University, as part of their talk series.
