A preprint reports a video model that learns which moments and tracked entities support driving-risk predictions from coarse video labels.