Related papers: VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommendation

VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommendation

URL: http://arxiv.org/abs/2602.04567v2
Date: Tue, 10 Feb 2026 08:09:13 GMT
Title: VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommendation
Authors: Aleksandr Poslavsky, Alexander D'yakonov, Yuriy Dorn, Andrey Zimovnov,
Abstract summary: We introduce the VK Large Short-Video dataset (VK-LSVD), the largest publicly available industrial dataset of its kind.<n>VK-LSVD offers an unprecedented scale of over 40 billion interactions from 10 million users and almost 20 million videos over six months, alongside rich features including content embeddings, diverse feedback signals, and contextual metadata.
Score: 74.72450521118019
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Short-video recommendation presents unique challenges, such as modeling rapid user interest shifts from implicit feedback, but progress is constrained by a lack of large-scale open datasets that reflect real-world platform dynamics. To bridge this gap, we introduce the VK Large Short-Video Dataset (VK-LSVD), the largest publicly available industrial dataset of its kind. VK-LSVD offers an unprecedented scale of over 40 billion interactions from 10 million users and almost 20 million videos over six months, alongside rich features including content embeddings, diverse feedback signals, and contextual metadata. Our analysis supports the dataset's quality and diversity. The dataset's immediate impact is confirmed by its central role in the live VK RecSys Challenge 2025. VK-LSVD provides a vital, open dataset to use in building realistic benchmarks to accelerate research in sequential recommendation, cold-start scenarios, and next-generation recommender systems.

Related papers

Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models [31.566051946153802]
Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios.<n>We introduce Impromptu VLA: over 80,000 meticulously curated video clips, distilled from over 2M source clips sourced from 8 open-source large-scale datasets.<n>This dataset is built upon our novel taxonomy of four challenging unstructured categories and features rich, planning-oriented question-answering annotations and action trajectories.
arXiv Detail & Related papers (2025-05-29T17:59:46Z)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents [57.59830804627066]
We introduce MONDAY, a large-scale dataset of 313K annotated frames from 20K instructional videos capturing real-world mobile OS navigation.<n>Models that include MONDAY in their pre-training phases demonstrate robust cross-platform generalization capabilities.<n>We present an automated framework that leverages publicly available video content to create comprehensive task datasets.
arXiv Detail & Related papers (2025-05-19T02:39:03Z)
Short-video Propagation Influence Rating: A New Real-world Dataset and A New Large Graph Model [66.8976337702408]
Cross-platform Short-Video dataset includes 117,720 videos, 381,926 samples, and 535 topics across 5 biggest Chinese platforms.<n>Large Graph Model (LGM) named NetGPT can bridge heterogeneous graph-structured data with the powerful reasoning ability and knowledge of Large Language Models (LLMs)<n>Our NetGPT can comprehend and analyze the short-video propagation graph, enabling it to predict the long-term propagation influence of short-videos.
arXiv Detail & Related papers (2025-03-31T05:53:15Z)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation [16.80010133425332]
We introduce Presto, a novel video diffusion model designed to generate 15-second videos with long-range coherence and rich content.<n>Presto achieves 78.5% on the VBench Semantic Score and 100% splits on the Dynamic Degree, outperforming existing state-of-the-art video generation methods.
arXiv Detail & Related papers (2024-12-02T09:32:36Z)
LLaVA-Video: Video Instruction Tuning With Synthetic Data [84.64519990333406]
We create a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K.<n>This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA.<n>By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM.
arXiv Detail & Related papers (2024-10-03T17:36:49Z)
CinePile: A Long Video Question Answering Dataset and Benchmark [55.30860239555001]
We present a novel dataset and benchmark, CinePile, specifically designed for authentic long-form video understanding. Our comprehensive dataset comprises 305,000 multiple-choice questions (MCQs), covering various visual and multimodal aspects. We fine-tuned open-source Video-LLMs on the training split and evaluated both open-source and proprietary video-centric LLMs on the test split of our dataset.
arXiv Detail & Related papers (2024-05-14T17:59:02Z)
Scaling Up Video Summarization Pretraining with Large Language Models [73.74662411006426]
We introduce an automated and scalable pipeline for generating a large-scale video summarization dataset. We analyze the limitations of existing approaches and propose a new video summarization model that effectively addresses them. Our work also presents a new benchmark dataset that contains 1200 long videos each with high-quality summaries annotated by professionals.
arXiv Detail & Related papers (2024-04-04T11:59:06Z)

This list is automatically generated from the titles and abstracts of the papers in this site.