Score: 0

ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis

Published: January 7, 2026 | arXiv ID: 2601.03632v1

By: Haitao Li , Chunxiang Jin , Chenglin Li and more

Potential Business Impact:

Changes voice style without changing the speaker.

Business Areas:

Speech Recognition Data and Analytics, Software

Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result, synthesizing speech with a desired style often requires carefully selecting reference audio, which is impractical when only limited or mismatched references are available. While recent controllable TTS methods attempt to address this issue, they typically rely on absolute style targets and discrete textual prompts, and therefore do not support continuous and reference-relative style control. We propose ReStyle-TTS, a framework that enables continuous and reference-relative style control in zero-shot TTS. Our key insight is that effective style control requires first reducing the model's implicit dependence on reference style before introducing explicit control mechanisms. To this end, we introduce Decoupled Classifier-Free Guidance (DCFG), which independently controls text and reference guidance, reducing reliance on reference style while preserving text fidelity. On top of this, we apply style-specific LoRAs together with Orthogonal LoRA Fusion to enable continuous and disentangled multi-attribute control, and introduce a Timbre Consistency Optimization module to mitigate timbre drift caused by weakened reference guidance. Experiments show that ReStyle-TTS enables user-friendly, continuous, and relative control over pitch, energy, and multiple emotions while maintaining intelligibility and speaker timbre, and performs robustly in challenging mismatched reference-target style scenarios.

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

Sound

Makes computer voices sound like anyone you want.

8 Jan 2026 0

89%

Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement

Sound

Makes computer voices sound more real and emotional.

2 Oct 2025 0

88%

ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation

Sound

Makes computer voices sound happy or sad.

21 Oct 2025 2

View PDF Login to Bookmark

Country of Origin

🇨🇳 China

Page Count

13 pages

ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis

Changes voice style without changing the speaker.

Technical Abstract

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement

ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation