Interpret the edit
Map a natural-language request to operations on nine perceptual voice attributes. Unmentioned attributes receive a keep operation.
Edit who speaks.
Direct how each segment sounds.
Change timbre through natural-language edits.
Direct each text segment with its own instruction.
First, EDICT edits the source voice. It then synthesizes speech in the edited timbre, following a local instruction for each text segment.
These six qualitative examples combine global timbre editing with local instruction control. EDICT first edits the source voice according to the timbre request. The editor-generated reference then anchors the edited timbre throughout synthesis, while each segment follows its own expression instruction.
Start with the source voice, read the edit, then compare the four systems on the same synthesis text.
Listen for how closely each output follows the edit instruction.
Instruction-conditioned systems receive the source reference and the same edit request. The target reference is provided only for listening comparison. Target-reference voice conversion is reported separately in the results.
Nine voice attributes: Perceived gender · Age impression · Pitch register · Brightness · Vocal weight · Resonance focus · Roughness · Breathiness · Nasality
One utterance, several instructions. Compare expression, voice consistency and the transitions between segments.
The colors trace instruction order across the text.
All four methods use the Qwen-VD backbone with the same text, global voice description and ordered local instructions. These examples are text-conditioned and use no reference audio.
Subjective listening scores and main experimental comparisons across timbre editing, local control and joint synthesis.
Points show the mean of item-level listener averages. Whiskers reproduce the reported symmetric display of 95% bootstrap confidence intervals from 10,000 audio-item resamples: the half-width is the larger distance from the mean to either interval endpoint. † Seed-VC uses target audio. Blue circles identify EDICT; the open diamond distinguishes Seed-VC.
All conditions are expanded below. Arrows show the preferred direction for each metric; blue rows identify EDICT.
EDICT improves edit accuracy, while Qwen-VD better preserves unedited attributes. Cache reconstruction improves local control and transitions, with increased recognition errors relative to Joint. Independent synthesis retains the strongest isolated segment accuracy.
Instruction-based text-to-speech (TTS) offers control over voice characteristics and speech expression through interfaces including voice cloning and text-based voice design. Voice cloning reproduces a reference voice, whereas text-based voice design creates a voice from a natural-language description. However, neither interface directly enables users to modify the timbre of a given reference and synthesize speech with the modified voice. Meanwhile, utterance-level expressive instructions leave changes across individual text segments underspecified. We introduce EDICT, a framework that unifies global timbre editing and local expressive control by using an edited acoustic reference to anchor voice identity across segments. To enable synthesis with an instruction-edited voice, EDICT combines reference audio with structured timbre edits to generate an edited reference in codec-token space. This representation serves as a shared voice anchor for a frozen TTS backbone, allowing segment-specific natural-language instructions to guide expression. To accommodate instruction changes while supporting acoustic continuity, EDICT rebuilds the KV cache at each segment boundary, refreshing instruction conditioning while retaining bounded acoustic context from previously generated speech. Evaluations on our proposed TimbreEdit-Bench and IntraTTS-Bench demonstrate improved timbre editing and a favorable balance between local instruction adherence, speaker consistency, and transition quality.
EDICT first edits the reference voice, then synthesizes speech in the edited timbre with segment-specific instructions.
Map a natural-language request to operations on nine perceptual voice attributes. Unmentioned attributes receive a keep operation.
The trained editor produces an edited codec reference, reused across all segments of the subsequent synthesis.
At each instruction switch, reconstruct the KV cache using the new instruction, shared reference, and initial and recent acoustic context.
Cached acoustic states can carry conditioning from an earlier instruction. EDICT discards the old cache and recomputes selected acoustic inputs under the new conditions. Initial context supplies evidence of the established voice; recent context supplies continuation cues. The generated segments remain in one codec stream.
The local control stage requires no additional backbone training.