Szczegóły publikacji
Opis bibliograficzny
Combating reverberation in NTF-based speech separation using a sub-source weighted multichannel Wiener filter and linear prediction / Mieszko FRAŚ, Marcin WITKOWSKI, Konrad KOWALCZYK // W: INTERSPEECH 2021 [Dokument elektroniczny] : proceedings of the 22nd annual conference of the International Speech Communication Association : 30 August–3 September 2021, Brno, Czechia. — Brno : Brno University of Technology, 2021. — (Interspeech : Proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 1990-9772). — e-ISBN: 978-171383690-2. — S. 3895–3899. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-speech.org/archive/pdfs/interspeech_2021/fra... [2021-09-21]. — Bibliogr. s. 3899. Abstr. — W bazie Scopus zakres stron: 2403-2407
Autorzy (3)
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 136293 |
|---|---|
| Data dodania do BaDAP | 2021-09-27 |
| DOI | 10.21437/Interspeech.2021-1230 |
| Rok publikacji | 2021 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | Interspeech 2021 |
| Czasopismo/seria | Interspeech |
Abstract
Sound source separation (SS) from the microphone signals capturing speech in reverberant conditions is a formidable task. This paper addresses the problem of joint separation and dereverberation of speech using the multichannel Wiener filter (MWF) that is tailored to the sub-source modeling of each speech source with a full-rank mixing matrix. Specifically, the parameters of the proposed sub-source-weighted (SSW) spatial filter are estimated using the sub-source based expectation maximization (EM) algorithm with multiplicative updates (MU) and the localization prior distribution (LP) on the mixing matrix (SSEM-MU-LP). In addition, we strengthen dereverberation by incorporating a Generalized Weighted Prediction Error (GWPE) algorithm. The proposed method is evaluated using a large dataset of two-channel recordings of clean speech convolved with both real and synthesized impulse responses. The results of the experiments show the superior performance of the proposed method in reverberant conditions in comparison to using the standard NTF-based separation with the vanilla MWF in terms of signal-to-distortion ratio (improvement of 3–5.6 dB) and other commonly used sound separation metrics.