Seeing, Reading, Hearing, Interpreting: Computational Multimodal Framing Analysis in Short-Form News Videos
DOI:
https://doi.org/10.5117/CCR2026.4.4.HUKeywords:
computational methods, multimodal framing, multimodal large language model, tiktok, video-as-dataAbstract
Short-form news videos convey meaning through intertwined visual, auditory, and textual cues, yet most computational approaches analyze them primarily through textual features (e.g., transcripts) or isolated visual elements. Audio elements (e.g., music, speech, ambient sound), video-length visual dynamics, and their interactions remain largely underexamined. This study evaluates how effectively single-modality (visual, audio, text) and multimodal large language model (LLM)-assisted approaches identify frames in short-form news videos. Using 2 , 957 TikTok videos about the Russia-Ukraine war from three mainstream outlets (CNN, CGTN Europe, The Kyiv Independent), we compare LLM-assisted automated frame detection against manually coded ground truth across six frame categories: Conflict, Human Interest, Responsibility, Morality, Economic Consequence, and Diplomatic/Political Leader. Building on prior semiotic frameworks of image-text relations, we examine how modality interactions (complementary, reinforcing, or independent) shape overall multimodal framing outcome. We further implement a human-in-the-loop validation protocol to clarify the practical limits of automated inference. Benchmark results show that multimodal frameworks outperform single-modality approaches, though the modest improvement over text-only detection raises questions about cost-benefit tradeoffs for large-scale applications. Human judgment remains essential for disambiguating abstract frames, underscoring that methodological choices must account for both the multimodal nature of video and the practical limits of computational inference.Downloads
Published
2026-08-03
Issue
Section
Special Issue: Short Video
How to Cite
Hu, H. L., & Fu, K.- wa. (2026). Seeing, Reading, Hearing, Interpreting: Computational Multimodal Framing Analysis in Short-Form News Videos. Computational Communication Research, 8(4). https://doi.org/10.5117/CCR2026.4.4.HU


