Score: 1

Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing

Published: September 17, 2025 | arXiv ID: 2509.13646v2

By: Kexue Fu , Jingfei Huang , Long Ling and more

Potential Business Impact:

Helps writers create stories with pictures and words.

Business Areas:

Visual Search Internet Services

Humans think visually-we remember in images, dream in pictures, and use visual metaphors to communicate. Yet, most creative writing tools remain text-centric, limiting how authors plan and translate ideas. We present Vistoria, a system for synchronized text-image co-editing in fictional story writing that treats visuals and text as coequal narrative materials. A formative Wizard-of-Oz co-design study with 10 story writers revealed how sketches, images, and annotations serve as essential instruments for ideation and organization. Drawing on theories of Instrumental Interaction and Structural Mapping, Vistoria introduces multimodal operations-lasso, collage, filters, and perspective shifts that enable seamless narrative exploration across modalities. A controlled study with 12 participants shows that co-editing enhances expressiveness, immersion, and collaboration, enabling writers to explore divergent directions, embrace serendipitous randomness, and trace evolving storylines. While multimodality increased cognitive demand, participants reported stronger senses of authorship and agency. These findings demonstrate how multimodal co-editing expands creative potential by balancing abstraction and concreteness in narrative development.

Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing

Human-Computer Interaction

Helps writers create stories with pictures and words.

17 Sep 2025 1

88%

ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models

CV and Pattern Recognition

Makes stories with pictures that make sense.

13 Jun 2025 1

88%

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Computation and Language

Computers write stories from pictures.

27 Apr 2025 0

View PDF Login to Bookmark

Country of Origin

🇭🇰 🇺🇸 United States, Hong Kong

Page Count

34 pages

Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing

Helps writers create stories with pictures and words.

Technical Abstract

Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing

ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?