MSFEYNet: A Hybrid Convolution Transformer Network with Spectral Feature Extraction & MSFEB Modelling for Single-Channel Speech Enhancement

Main Article Content

Silpa Peethala
 Sunny Dayal Vanambathina

Abstract

This paper presents MSFEYNet, a novel dual-decoder U-Net architecture for single-channel speech enhancement that integrates convolutional and Transformer blocks to exploit local spectral information and long-range contextual dependencies jointly. The encoder extracts hierarchical multi-scale spectral representations, while two dedicated decoders simultaneously reconstruct the clean speech magnitude and phase spectra, enabling accurate speech recovery. To enhance feature representation, the proposed architecture incorporates three key components: the Triple Path Fused Attention (TPFA) module for efficient long-range axial spectral modelling, the Large Group Large Kernel Attention (LGLKA) module for capturing finegrained local axial structures, and an Enhanced Feed Forward Network (EFFN) with dynamic axial-aware feature reweighting to improve contextual feature learning. Furthermore, a Multi-Scale Feature Extraction Block (MSFEB) extracts both fine-grained local features and broad contextual information using convolutional filters with varying receptive fields, thereby enhancing multi-scale spectral-temporal feature representation. Extensive experiments on two datasets demonstrate that MSFEYNet achieves robust, well-balanced performance across a comprehensive set of objective evaluation metrics through the synergistic integration of the dual-decoder framework. The transformer-based attention mechanism and Multi-scale Feature Extraction Block (MSFEB) effectively capture local and global contextual dependencies, leading to significant improvements in speech intelligibility and perceptual quality. Experimental results further demonstrate that MSFEYNet consistently outperforms contemporary state-of-the-art speech enhancement methods, achieving substantial gains in STOI (Short-Time Objective Intelligibility) and PESQ (Perceptual Evaluation of Speech Quality) while maintaining excellent generalisation across diverse acoustic conditions and varying noise levels.

Downloads

Download data is not yet available.

Article Details

Section

Articles

How to Cite

[1]
Silpa Peethala and  Sunny Dayal Vanambathina , Trans., “MSFEYNet: A Hybrid Convolution Transformer Network with Spectral Feature Extraction & MSFEB Modelling for Single-Channel Speech Enhancement”, IJRTE, vol. 15, no. 3, pp. 1–11, Sep. 2026, doi: 10.35940/ijrte.C8389.15030926.
Share |

References

Ochieng, P. (2023). Deep neural network techniques for monaural speech enhancement and separation: State of the art analysis. Artificial Intelligence Review, 56, 3651–3703. DOI: 10.1007/s10462-023-10612-2

Li, Y., Sun, Y., Wang, W., & Naqvi, S. M. (2023). U-Shaped Transformer with Frequency-Band Aware Attention for Speech Enhancement. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31, 1511–1521. DOI: 10.1109/TASLP.2023.3265839.

Lu, Y.-X., Ai, Y., & Ling, Z.-H. (2023). MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra. In Proceedings of Interspeech 2023 (pp. 3834–3838). DOI: 10.21437/Interspeech.2023-1441

Kühne, N. L., Østergaard, J., Jensen, J., & Tan, Z.-H. (2025). xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement. In Proceedings of Interspeech 2025 (pp. 5148–5152). DOI: 10.21437/Interspeech.2025-108

Zhou, H., Zhou, Y., Cheng, Z., Zhao, Y., & Liu, Y. (2025). Improved Encoder–Decoder Architecture with Human-Like Perception Attention for Monaural Speech Enhancement. IEEE Signal Processing Letters, 32, 1670–1674. DOI: 10.1109/LSP.2025.3558690

Lin, Z.,Wang, J., Li, R., Shen, F., Xuan, X., 2025. PrimeK-Net: Multi-scale spectral learning via group prime-kernel convolutional neural networks for single-channel speech enhancement, in: Proc. ICASSP, pp. 1–5. DOI: 10.1109/ICASSP49660.2025.10890034

Liang, X., Zhang, Z., Wang, M., & Xu, R. (2024, April). Lightweight multi-axial transformer with frequency prompt for single-channel speech enhancement. In ICASSP 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 10511-10515). IEEE.

DOI: 10.1109/ICASSP48485.2024.10446787

Chao, R., et al., 2024. An investigation of incorporating Mamba for speech enhancement, in: Proc. IEEE Spoken Lang. Technol. Workshop (SLT), pp. 302–308. DOI: 10.1109/SLT61566.2024.10832332

Luo, T., Zhou, F., Bai, Z., 2024. MambaGAN: Mamba-based metric GAN for monaural speech enhancement, in: Proc. Int. Conf. Asian Lang. Process. (IALP), pp. 411–416. DOI: 10.1109/IALP63756.2024.10661187

Abdulatif, S., Cao, R., Yang, B., 2024. CMGAN: Conformer-based Metric-GAN for monaural speech enhancement. IEEE/ACM Trans. Audio Speech Lang. Process. 32, 2477–2493. DOI: 10.1109/TASLP.2024.3393718

Hu, Y., Yang, Q., Wei, W., Lin, L., He, L., Ou, Z., Yang, W., 2025. MNNet: Speech enhancement network via modelling the noise. IEEE Trans. Audio Speech Lang. Process. 33, 1208–1219. DOI: 10.1109/TASLPRO.2025.3546819

Lu, Y.X., Ai, Y., Ling, Z., 2025. Explicit estimation of magnitude and phase spectra in parallel for high-quality speech enhancement. Neural Networks 189, 107562. DOI: 10.1016/j.neunet.2025.107562

Chen, H., Zhang, J., Fu, Y., Zhou, X., Wang, R., Xu, Y., & Ke, D. (2025). TFDense-GAN: A Generative Adversarial Network for Single-Channel Speech Enhancement. EURASIP Journal on Advances in Signal Processing, 2025, Article 10. DOI: 10.1186/s13634-025-01210-1

Wang, H., Tian, B., 2025. ZipEnhancer: Dual-path down-up sampling-based Zipformer for monaural speech enhancement, in: Proc. ICASSP, pp. 1–5. DOI: 10.1109/ICASSP49660.2025.10888703

Wang, J., Lin, Z., Wang, T., Ge, M., Wang, L., Dang, J., 2025. Mamba-SEUNet: Mamba UNet for monaural speech enhancement, in: Proc.ICASSP, pp. 1–5. DOI: 10.1109/ICASSP49660.2025.10889525

Yu, W., Si, C., Zhou, P., Luo, M., Zhou, Y., Feng, J., Yan, S., Wang, X., 2024. MetaFormer baselines for vision. IEEE Trans. Pattern Anal. Mach. Intell. 46, 896–912. DOI: 10.1109/TPAMI.2023.3329173

Dang, F., Chen, H., & Zhang, P. (2022, May). DPT-FSNet: Dual-path transformer-based full-band and sub-band fusion network for speech enhancement. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6857-6861). IEEE. DOI: 10.1109/ICASSP43922.2022.9746171

Zhang, C., Wang, L., Qiao, Y., & Gao, W. A multi-scale feature extraction fusion model for human activity recognition. Scientific Reports, 12(1), 20620 (2022). DOI; 10.1038/s41598-022-24909-2

Chen, M., Zhang, Q., Wang, M., Zhang, X., Liu, H., Ambikairaiah, E., & Chen, D. (2024). Selective state space model for monaural speech enhancement. arXiv preprint arXiv:2411.06217. DOI: https://arxiv.org/abs/2411.06217.

Chao, R., Nasretdinov, R., Wang, Y. C. F., Jukić, A., Fu, S. W., & Tsao, Y. (2025). Universal speech enhancement with regression and generative Mamba. arXiv preprint arXiv:2505.21198. DOI: 10.48550/arXiv.2505.21198

Wang, J., Lin, Z., Wang, T., Ge, M., Wang, L., & Dang, J. (2025). Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2025) (pp. 1–5).

DOI: 10.1109/ICASSP49660.2025.10889525

Mathieu, F., Courtat, T., Richard, G., & Peeters, G. Learning interpretable filters in Wav-U-Net for speech enhancement. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2023).

DOI: 10.1109/ICASSP.49357.2023.10097153

J.,&Zhang, X. ICCRN: In-place cepstral convolutional recurrent neural network for monaural speech enhancement.In International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 1-5). IEEE (2023). DOI: 10.1109/ICASSP 49357.2023.10096918

Fang, H., Becker, D., Wermter, S., & Gerkmann, T. Integrating uncertainty into neural network-based speech enhancement. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31, 1587–1600. (2023). DOI: 10.1109/TASLP.2023.3274628

Zmolikova, K., Pedersen, M. S., & Jensen, J. Masked spectrogram prediction for unsupervised domain adaptation in speech enhancement. IEEE Open Journal of Signal Processing, 5, 274–283 (2023). DOI; 10.1109/OJSP.2023.3305948

Fiorio, L. V., Karanov, B., Defraene, B., David, J., Widdershoven, F., van Houtum, W., & Aarts, R. M. Spectral masking with explicit time-context windowing for neural network-based monaural speech enhancement. IEEE Access, 12, 154843–154852 (2024).

DOI: 10.1109/ACCESS.2024.3483443

Alohali, M. A., Saleem, N., Rhouma, D., Medani, M., Elmannai, H., & Bourouis, S.Temporally dynamic spiking transformer network for speech enhancement. IEEE Access, 12, 146513–146526 (2024). DOI: 10.1109/ACCESS.2024.3444596

Valentini-Botinhao, C., Wang, X., Takaki, S., & Yamagishi, J. (2016). Investigating RNN-based Speech Enhancement Methods for Noise-Robust Text-to-Speech. In Proceedings of the 9th ISCA Speech Synthesis Workshop (SSW 9) (pp. 159–165).

DOI: 10.21437/SSW.2016-19

Kadambi, M., Sharma, C. M., Mandal, A., & Ghosh, P. K. (2025, December). Bottleneck Transformer-Based Approach for Improved Automatic STOI Score Prediction. In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) (pp. 1-7). IEEE DOI: 10.1109/ASRU65441.2025.11434770

Xu, Z., Zhao, Z., & Fingscheidt, T. (2023). Coded speech quality measurement by a non-intrusive PESQ-DNN. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31, 3404-3417. DOI: 10.1109/TASLP.2023.3317574

Kadir, A. D. I. A., Al-Haiqi, A., & Din, N. M. (2021, December). A dataset and TinyML model for coarse age classification based on voice commands. In 2021 IEEE 15th Malaysia International Conference on Communication (MICC) (pp. 75-80). IEEE.

DOI: 10.1109/MICC53484.2021.9642091

Yu, G., Li, A., Zheng, C., Guo, Y., Wang, Y., & Wang, H. (2022, May). Dual-branch attention-in-attention transformer for single-channel speech enhancement. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 7847-7851). IEEE. DOI: 10.1109/ICASSP43922.2022.9746273

Ma, X., Dai, X., Bai, Y., Wang, Y., & Fu, Y. (2024). Rewrite the Stars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5694–5703). IEEE. DOI: 10.1109/CVPR52733.2024.00544.

Guo, M.-H., Lu, C.-Z., Liu, Z.-N., Cheng, M.-M., & Hu, S.-M. (2023). Visual attention network. Computational Visual Media, 9(4), 733–752. DOI: 10.1007/s41095-023-0364-2

Most read articles by the same author(s)

1 2 3 4 5 6 7 8 9 10 > >>