ANONYMOUS RESEARCH DEMO · TEXT-TO-SPEECH

EDICT: Global Timbre Editing and
Local Instruction Control for TTS

Edit who speaks.
Direct how each segment sounds.

ONE VOICE, MANY EXPRESSIONS
01 / GLOBAL

A voice you can edit.

Reference voiceEdited voice

Change timbre through natural-language edits.

Shared voice anchor
02 / LOCAL

Expression that can change.

CalmExcitedRelieved

Direct each text segment with its own instruction.

01 / LISTEN TO THE COMBINATION

One voice. A change of expression.

First, EDICT edits the source voice. It then synthesizes speech in the edited timbre, following a local instruction for each text segment.

Joint control studio
Choose a sample
About these joint demos Editing + local synthesis

These six qualitative examples combine global timbre editing with local instruction control. EDICT first edits the source voice according to the timbre request. The editor-generated reference then anchors the edited timbre throughout synthesis, while each segment follows its own expression instruction.

02 / GLOBAL TIMBRE EDITING

Hear the voice change.

Start with the source voice, read the edit, then compare the four systems on the same synthesis text.

1 Source voice2 Edit instruction3 Compare outputs

Listen for how closely each output follows the edit instruction.

Comparison protocol Inputs + nine voice attributes

Instruction-conditioned systems receive the source reference and the same edit request. The target reference is provided only for listening comparison. Target-reference voice conversion is reported separately in the results.

Nine voice attributes: Perceived gender · Age impression · Pitch register · Brightness · Vocal weight · Resonance focus · Roughness · Breathiness · Nasality

03 / LOCAL INSTRUCTION CONTROL

Let the delivery evolve.

One utterance, several instructions. Compare expression, voice consistency and the transitions between segments.

1 Read the segment instructions2 Compare complete outputs

The colors trace instruction order across the text.

Comparison protocol Same backbone + same instructions

All four methods use the Qwen-VD backbone with the same text, global voice description and ordered local instructions. These examples are text-conditioned and use no reference audio.

04 / EXPERIMENTAL EVIDENCE

Behind the listening experience.

Subjective listening scores and main experimental comparisons across timbre editing, local control and joint synthesis.

1,955source–target pairsTimbreEdit-Bench · 200 held-out voice profiles
997synthesis requestsIntraTTS-Bench · Chinese + English
19listenersSubjective evaluation · five-point scales
HUMAN EVALUATION

How does it sound?

All six charts ↗
MEAN OPINION SCORE

Edit Accuracy

EDICT

EDICT Comparison systems Mean ± reported 95% CI display
Evaluation protocol and exact scores

Points show the mean of item-level listener averages. Whiskers reproduce the reported symmetric display of 95% bootstrap confidence intervals from 10,000 audio-item resamples: the half-width is the larger distance from the mean to either interval endpoint. † Seed-VC uses target audio. Blue circles identify EDICT; the open diamond distinguishes Seed-VC.

MAIN EXPERIMENTS

Main experimental comparisons

All conditions are expanded below. Arrows show the preferred direction for each metric; blue rows identify EDICT.

READING THE TRADE-OFFS

EDICT improves edit accuracy, while Qwen-VD better preserves unedited attributes. Cache reconstruction improves local control and transitions, with increased recognition errors relative to Joint. Independent synthesis retains the strongest isolated segment accuracy.

Abstract

Instruction-based text-to-speech (TTS) offers control over voice characteristics and speech expression through interfaces including voice cloning and text-based voice design. Voice cloning reproduces a reference voice, whereas text-based voice design creates a voice from a natural-language description. However, neither interface directly enables users to modify the timbre of a given reference and synthesize speech with the modified voice. Meanwhile, utterance-level expressive instructions leave changes across individual text segments underspecified. We introduce EDICT, a framework that unifies global timbre editing and local expressive control by using an edited acoustic reference to anchor voice identity across segments. To enable synthesis with an instruction-edited voice, EDICT combines reference audio with structured timbre edits to generate an edited reference in codec-token space. This representation serves as a shared voice anchor for a frozen TTS backbone, allowing segment-specific natural-language instructions to guide expression. To accommodate instruction changes while supporting acoustic continuity, EDICT rebuilds the KV cache at each segment boundary, refreshing instruction conditioning while retaining bounded acoustic context from previously generated speech. Evaluations on our proposed TimbreEdit-Bench and IntraTTS-Bench demonstrate improved timbre editing and a favorable balance between local instruction adherence, speaker consistency, and transition quality.

Method overview

EDICT first edits the reference voice, then synthesizes speech in the edited timbre with segment-specific instructions.

01

Interpret the edit

Map a natural-language request to operations on nine perceptual voice attributes. Unmentioned attributes receive a keep operation.

02

Edit the reference voice

The trained editor produces an edited codec reference, reused across all segments of the subsequent synthesis.

03

Refresh the context

At each instruction switch, reconstruct the KV cache using the new instruction, shared reference, and initial and recent acoustic context.

FIGURE 01 Global voice editing and local delivery control in one synthesis pipeline.
What changes at an instruction boundary?Explore the KV cache ↗

Cached acoustic states can carry conditioning from an earlier instruction. EDICT discards the old cache and recomputes selected acoustic inputs under the new conditions. Initial context supplies evidence of the established voice; recent context supplies continuation cues. The generated segments remain in one codec stream.

The local control stage requires no additional backbone training.