Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training

July 05, 2024 · Declared Dead · 🏛 IEEE transactions on circuits and systems for video technology (Print)

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Mehryar Abbasi, Hadi Hadizadeh, Parvaneh Saeedi arXiv ID 2407.04258 Category cs.MM: Multimedia Cross-listed cs.AI, cs.CV, cs.LG Citations 1 Venue IEEE transactions on circuits and systems for video technology (Print) Last Checked 3 months ago

Abstract

This paper presents a novel approach for unsupervised video summarization using reinforcement learning (RL), addressing limitations like unstable adversarial training and reliance on heuristic-based reward functions. The method operates on the principle that reconstruction fidelity serves as a proxy for informativeness, correlating summary quality with reconstruction ability. The summarizer model assigns importance scores to frames to generate the final summary. For training, RL is coupled with a unique reward generation pipeline that incentivizes improved reconstructions. This pipeline uses a generator model to reconstruct the full video from the selected summary frames; the similarity between the original and reconstructed video provides the reward signal. The generator itself is pre-trained self-supervisedly to reconstruct randomly masked frames. This two-stage training process enhances stability compared to adversarial architectures. Experimental results show strong alignment with human judgments and promising F-scores, validating the reconstruction objective.