A Parametric Dual-Channel Audio Coding via Learned Time-Frequency Masking

Citation:

Wang Y, Qian Y, Huang Q, Qu T. A Parametric Dual-Channel Audio Coding via Learned Time-Frequency Masking, in the AES 160th Convention. Copenhagen, Denmark; 2026:10294.

Date Presented:

28-30 May

摘要:

While Neural Audio Codecs (NAC) have revolutionized monaural audio compression, achieving high-fidelity dual-channel coding at low bitrates remains a significant challenge. Existing approaches often rely on naive independent channel quantization, leading to phase incoherence, or entangled latent modeling, which sacrifices spatial precision for spectral energy. This paper proposes a novel dual-channel coding framework based on contentspatial disentanglement. Reframing spatial reconstruction as an informed source separation task, our architecturesynergizes a frozen, pre-trained DAC encoder for robust mono content preservation with a parameter-efficient side information encoder that predicts fine-grained time-frequency masks. To ensure precise spatial imaging, we introduce explicit physical constraints into the end-to-end training. Experimental results indicate that at low bitrates of 9 and 11 kbps, the proposed method outperforms state-of-the-art dual-mono neural baselines and industry standards in both objective spatial metrics and subjective MUSHRA evaluations.

访问链接