← Back to feed
2026-07-13multimodalvisioncode

StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin

PDF preview for StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Read on arXiv →

Key claim

StoryTeller enhances narrative coherence in audio descriptions.

In plain English

Long-form audio descriptions need to convey more than just visible actions; they must maintain the story's context for blind and low-vision audiences. Current video-language models struggle with this, often treating scenes in isolation and missing important narrative connections. StoryTeller addresses this by using a narrative memory to keep track of story-relevant information across scenes, allowing for coherent and contextually rich descriptions. Builders might care because this method does not require extensive training or additional resources, making it accessible for various applications.

Novelty
8.0/10

Introduces a novel framework for coherent long-form audio descriptions without training.

Reliability
7.5/10

Demonstrates improvements over strong baselines with multiple evaluation methods.

Deep reliability assessment

The methodology supports the claim that StoryTeller improves narrative coherence and factual grounding in long-form audio descriptions without task-specific training. However, the reliance on public movie metadata and the lack of testing under different narrative complexities might overclaim its general applicability.

Reproducibility

yes, the paper provides a GitHub repository for the StoryAD-QA benchmark dataset and evaluation code.

Key figure

Figure 1 illustrates the StoryTeller framework, highlighting its use of a persistent narrative state with an identity graph and salience-weighted memory to maintain narrative coherence across scenes.

Benchmark results

StoryAD-QAaccuracy: 0.972vs AutoAD-Zero+0.061SOTA
GitHub1 repo
SEE-AI-Lab/ECCV2026_StoryTeller_StoryAD_QAOfficial