Robust Speech Signal Improvement via Local Temporal Modeling Across Multi-Scale Resolutions
编号:95
访问权限:仅限参会人
更新:2026-10-04 23:39:21 浏览:7次
In-person
摘要
Real-world speech signals are often degraded by a complex combination of factors, including non-stationary noise, reverberation, and acoustic echo. While hierarchical encoder -decoder architectures are widely utilized for real-time communication due to their computational efficiency, existing models primarily rely on bottleneck transitions to capture long-term temporal dependencies, often overlooking the significance of fine-grained local features across different resolutions. In this paper, we propose a real-time speech restoration framework designed to explicitly extract local temporal features across multi-scale resolutions. By integrating localized temporal modeling within each hierarchical level, the proposed system effectively suppresses complex noise components while preserving critical speech nuances. Experimental results on the ICASSP 2024 Speech Signal Improvement blind test set demonstrate that our model achieves a superior balance between perceptual quality and execution speed, yielding an OVRL of 3.13 and a SIG of 3.56. Notably, the system outperforms several state-of-the-art baselines in overall signal restoration while maintaining a highly efficient real-time factor (RTF) of 0.33.
关键词
Speech Signal Improvement,Deep Learning,Multi-scale Resolutions
稿件作者
Thi Nhat Linh Nguyen
Hanoi University of Science and Technology
Minh Thuy Le
Hanoi University of Science and Technology
Kien Nguyen
Chiba University
Quoc Cuong Nguyen
Hanoi University of Science and Technology
发表评论