Speech bandwidth extension

CodecFlow: Speech Bandwidth Extension via Residual Flow Matching in Codec Latent Space

Bowen Zhang1,2Junchuan Zhao3Ian McLoughlin1Ye Wang3A. S. Madhukumar2

1Singapore Institute of Technology2Nanyang Technological University3National University of Singapore

Abstract

Speech bandwidth extension (BWE) recovers the high-frequency content missing from narrowband speech, and one low-band input admits many plausible full-band signals. Most methods work in the waveform or spectrogram domain. The continuous latent of a pretrained neural codec is an attractive alternative, since it is compact, aligned frame by frame between the narrowband input and its full-band target, and paired with a reusable decoder. Existing latent-domain methods, however, generate the whole target rather than only the part the low band does not determine, and they leave the codec's decoding path untouched. We present CodecFlow, a two-stage BWE system that addresses both points. Stage 1 trains a conditional flow-matching converter on the frozen encoder's latent to generate only the residual between the low-band and full-band latents, conditioned on the low-band latent and a frame-level voicing descriptor. Stage 2 adapts the codec decoder to reconstruct waveforms from these continuous predictions with the quantiser bypassed. Ablations show that the residual flow gives a better system than input-anchored or whole-latent flows once each has an adapted decoder. Decoder adaptation carries most of the broadband gain and quantiser removal most of the low-band gain. Against four published systems on three test sets, CodecFlow attains the lowest log-spectral distance (LSD) and high-band LSD at 16 and 44.1 kHz and the highest ViSQOL at 16 kHz. Low-band fidelity remains below that of waveform-domain methods.

Method

CodecFlow architecture: frozen DAC encoder and voicing descriptor condition a U-Conformer residual flow; an adapted DAC decoder reconstructs high-resolution speech with RVQ bypassed.View full size ↗
Overview of CodecFlow. A frozen DAC encoder maps band-limited speech to continuous latents 𝐳L, and a voicing descriptor provides per-frame scalars 𝐒 = [eng, log F0]; the two are fused into the condition. A U-Conformer conditioned on it and the flow time t ∼ 𝒰[0, 1] defines a velocity field whose ODE solution transports noise 𝐫0 to the residual 𝐫; the adapted DAC decoder then reconstructs HR speech from 𝐳H = 𝐳L + 𝐫 with the RVQ bypassed.
Architecture figure · PDF ↗

Results

LibriTTS test-clean8 → 16 kHz · Within-corpus
Full LibriTTS mel-spectrogram comparison of input, NU-Wave2, FlowHigh, Fre-Painter, AP-BWE, CodecFlow, and reference.View full size ↗
TIMIT8 → 16 kHz · Cross-corpus
Full TIMIT mel-spectrogram comparison of input, NU-Wave2, FlowHigh, Fre-Painter, AP-BWE, CodecFlow, and reference.View full size ↗
VCTK8 → 44.1 kHz · Held-out speakers
Full VCTK mel-spectrogram comparison of input, NU-Wave2, FlowHigh, Fre-Painter, AP-BWE, CodecFlow, and reference.View full size ↗

Baseline comparison

Two utterances per test set, shared across all systems. CodecFlow is highlighted in violet.

Loading audio comparisons…

Ablation study

Decoding-path ablation. RFC: residual flow converter; FLC: full-latent flow converter; RVQ: residual vector quantiser; D: decoder.

Loading ablation samples…

Full-size figureOpen original ↗