a multi-modal video caption dataset with richer annotation
-
Updated
Jun 5, 2024 - Python
a multi-modal video caption dataset with richer annotation
A repository of Video Language papers, code and datasets.
Pressure Testing Large Video-Language Models (LVLM): Doing multimodal retrieval from LVLM at any video lengths to measure accuracy
The official GitHub page for the survey paper "Self-Supervised learning for Videos: A survey"
VLG: General Video Recognition with Web Textual Knowledge (https://arxiv.org/abs/2212.01638)
ACM Multimedia 2023 (Oral) - RTQ: Rethinking Video-language Understanding Based on Image-text Model
Code for CVPR 2023 paper "SViTT: Temporal Learning of Sparse Video-Text Transformers"
[ICCV 2023] The official PyTorch implementation of the paper: "Localizing Moments in Long Video Via Multimodal Guidance"
Pytorch version of DeCEMBERT: Learning from Noisy Instructional Videos via Dense Captions and Entropy Minimization (NAACL 2021)
official repo of "VideoGUI: A Benchmark for GUI Automation from Instructional Videos"
A Video Chat Agent with Temporal Prior
An end-to-end masked contrastive video-and-language pre-training framework
Official implementation for paper Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos
A curated list of video-text datasets in a variety of languages. These datasets can be used for video captioning (video description) or video retrieval.
PyTorch code for "Perceiver-VL: Efficient Vision-and-Language Modeling with Iterative Latent Attention" (WACV 2023)
The Pytorch implementation for "Video-Text Pre-training with Learned Regions"
A Survey on video and language understanding.
[CVPR21] Visual Semantic Role Labeling for Video Understanding (https://arxiv.org/abs/2104.00990)
A new multi-shot video understanding benchmark Shot2Story with comprehensive video summaries and detailed shot-level captions.
Pytorch code for Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
Add a description, image, and links to the video-language topic page so that developers can more easily learn about it.
To associate your repository with the video-language topic, visit your repo's landing page and select "manage topics."