Visual Statistical Learning (VSL) is typically investigated in isolation as either spatial or temporal learning, leaving open how these regularities interact when they co-occur in natural environments. Yet outside the laboratory, regularities unfold jointly across space and time and must be interpreted in context. Using a novel spatio-temporal paradigm in which spatially defined patterns dynamically moved in and out of view, with or without occluding elements, we examined how visual learning constructs structured internal descriptions from jointly available regularities.

We first replicated canonical spatial and temporal VSL effects within this integrated design. Crucially, learning reflected seamless integration of spatial and temporal statistics: behavior could not be explained by a simple additive combination of independent co-occurrence computations. Purely temporal regularities supported the construction of spatial structure, indicating cross-domain synthesis rather than parallel tracking. Moreover, motion-defined context and occlusion cues systematically shaped which regularities were learned from identical input, demonstrating that higher-level contextual biases are incorporated into the learning process in an equally integrated, non-linear manner.

Together, these findings reconceptualize VSL not as a passive recorder of isolated co-occurrences, but as a generative, interpretative process that selectively integrates spatio-temporal regularities with contextual biases to infer the latent structure of the environment.