ICASSP AUDIO DEMO

InstCharVoice: Grounding Natural-Language Instructions for Character-Level Control in Text-to-Speech

Submitted to to ICASSP2027

Abstract — Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability.

Natural-Language Instruction → Keyword Grounding → Character-Level Acoustic Planning → Speech Generation
motivation
Motivation and overview of InstCharVoice. The framework bridges natural-language instructions and explicit character-level acoustic control through instruction grounding and acoustic planning.
llm
Model architecture of InstCharVoice, including (a) training sequence construction, (b) autoregressive modeling of character-level speech chunks, and (c) grounding-aware training objective.
Listening Recommendation: Please listen to the Reference audio first to familiarize yourself with the target speaker’s voice, and then compare how well each system follows the same synthesis text and instruction. The Reference audio corresponds to the reference text, while all other audio samples correspond to the target synthesis text. To prevent overlapping playback, starting a new audio sample will automatically pause the currently playing one.