Score: 1

MusicAIR: A Multimodal AI Music Generation Framework Powered by an Algorithm-Driven Core

Published: November 21, 2025 | arXiv ID: 2511.17323v1

By: Callie C. Liao, Duoduo Liao, Ellie L. Zhang

BigTech Affiliations: Stanford University

Potential Business Impact:

Makes songs from just words and pictures.

Business Areas:
Internet Radio Media and Entertainment, Music and Audio

Recent advances in generative AI have made music generation a prominent research focus. However, many neural-based models rely on large datasets, raising concerns about copyright infringement and high-performance costs. In contrast, we propose MusicAIR, an innovative multimodal AI music generation framework powered by a novel algorithm-driven symbolic music core, effectively mitigating copyright infringement risks. The music core algorithms connect critical lyrical and rhythmic information to automatically derive musical features, creating a complete, coherent melodic score solely from the lyrics. The MusicAIR framework facilitates music generation from lyrics, text, and images. The generated score adheres to established principles of music theory, lyrical structure, and rhythmic conventions. We developed Generate AI Music (GenAIM), a web tool using MusicAIR for lyric-to-song, text-to-music, and image-to-music generation. In our experiments, we evaluated AI-generated music scores produced by the system using both standard music metrics and innovative analysis that compares these compositions with original works. The system achieves an average key confidence of 85%, outperforming human composers at 79%, and aligns closely with established music theory standards, demonstrating its ability to generate diverse, human-like compositions. As a co-pilot tool, GenAIM can serve as a reliable music composition assistant and a possible educational composition tutor while simultaneously lowering the entry barrier for all aspiring musicians, which is innovative and significantly contributes to AI for music generation.

Country of Origin
🇺🇸 United States

Page Count
10 pages

Category
Computer Science:
Sound