Distilling Vision-Language Models on Millions of Videos Paper β’ 2401.06129 β’ Published Jan 11, 2024 β’ 19