With the growing demand for short videos and personalized content, automated Video Log (Vlog) generation has become a key direction in multimodal content creation. Existing methods mostly rely on predefined scripts, lacking dynamism and personal expression. Therefore, there is an urgent need for an automated Vlog generation approach that enables effective multimodal collaboration and high personalization. To this end, we propose PersonaVlog, an automated multimodal stylized Vlog generation framework that can produce personalized Vlogs featuring videos, background music, and inner monologue speech based on a given theme and reference image. Specifically, we propose a multi-agent collaboration framework based on Multimodal Large Language Models (MLLMs). This framework can efficiently generate high-quality prompts required for multimodal content creation based on user input, thereby improving the efficiency and creativity of content creation. In addition, we incorporate a feedback and rollback mechanism that leverages MLLMs to evaluate and provide feedback on generated results, thereby achieving iterative optimization of multimodal contents. We also propose ThemeVlogEval, a theme-based automated benchmarking framework that provides standardized metrics and datasets for fair evaluation. Comprehensive experiments demonstrate the significant potential and advantages of our framework compared with several baselines, highlighting the effectiveness and great potential of our method for generating automated Vlogs.
PersonaVlog is based on a multimodal Multi-Agent Collaborative Framework (MACF) that can automatically generate complete and interesting stories, storyboards, video descriptions, character inner monologues, and background music descriptions based on input themes, character reference images, and styles. Subsequently, a Feedback and Rollback Mechanism (FRM) is used to iteratively optimize the multimodal generation results, ultimately achieving the automated generation of narrative, coherent, and diverse Video Log (Vlog) content.
Turn on the sound to enjoy the music and inner monologues.