PDF
(6492KB)
Abstract
In distributed speech front-end frameworks, multi-channel enhancement typically relies on inter-channel time synchronization and is therefore susceptible to delay mismatch, which may lead to performance degradation. Based on this, this work regards independent single-channel enhancement as a complementary solution: when time synchronization information becomes unreliable, it can still maintain relatively stable performance. In single-channel research, noisy scenarios hinder accurate retrieval of non-structured, envelope-bound phase; without explicit constraints, amplitude-phase imbalance disrupts harmonics and continuity, making explicit modeling pivotal. MP-SENet (mask phase aware speech enhancement network) significantly advanced parallel amplitude-phase modeling for simultaneous optimization but struggles with residual noise and phase adaptation in low-SNR environments. To address this, an enhanced MP-SENet is proposed: FRI (filter recycle interguide) in TF-Transformer disentangles global-local speech-noise features; a lightweight SR post-processing module refines the amplitude–phase estimation, curbing noise and artifacts; parallel phase and metric discriminators enable joint optimization to enhance quality and intelligibility. It achieves state-of-the-art results on VoiceBank+DEMAND (PESQ: 3.67, CSIG: 4.96, STOI: 0.97) and robust generalization on real-world WHISPER_SET_1 at 0, –5, –10 dB, which makes it promising for distributed front-end systems.
Keywords
speech enhancement
/
explicit modeling
/
amplitude-phase
/
global-local
/
noise and artifacts
Cite this article
Download citation ▾
Yun Wu, Xinting Li, Zhimin Cao.
Beyond Baseline MP-SENet: A Synergistic Optimization Framework for High-Fidelity Monaural Speech Enhancement.
Journal of Beijing Institute of Technology, 2026, 35 (4) : 485-498 DOI:10.15918/j.jbit1004-0579.2025.052