StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin
Read on arXiv →Key claim
StoryTeller enhances narrative coherence in audio descriptions.
In plain English
Long-form audio descriptions need to convey more than just visible actions; they must maintain the story's context for blind and low-vision audiences. Current video-language models struggle with this, often treating scenes in isolation and missing important narrative connections. StoryTeller addresses this by using a narrative memory to keep track of story-relevant information across scenes, allowing for coherent and contextually rich descriptions. Builders might care because this method does not require extensive training or additional resources, making it accessible for various applications.
Introduces a novel framework for coherent long-form audio descriptions without training.
Demonstrates improvements over strong baselines with multiple evaluation methods.
Deep reliability assessment
The methodology supports the claim that StoryTeller improves narrative coherence and factual grounding in long-form audio descriptions without task-specific training. However, the reliance on public movie metadata and the lack of testing under different narrative complexities might overclaim its general applicability.
Reproducibility
yes, the paper provides a GitHub repository for the StoryAD-QA benchmark dataset and evaluation code.
Key figure
Figure 1 illustrates the StoryTeller framework, highlighting its use of a persistent narrative state with an identity graph and salience-weighted memory to maintain narrative coherence across scenes.
