Szczegóły publikacji

Opis bibliograficzny

On ambisonic source separation with spatially informed non-negative tensor factorization / Mateusz GUZIK, Konrad KOWALCZYK // IEEE/ACM Transactions on Audio, Speech and Language Processing ; ISSN 2329-9290. — Tytuł poprz.: IEEE Transactions on Audio, Speech, and Language Processing ; ISSN: 1558-7916. — 2024 — vol. 32, s. 3238–3255. — Bibliogr. s. 3254–3255, Abstr. — Publikacja dostępna online od: 2024-05-10

Autorzy (2)

Słowa kluczowe

non negative tensor factorizationlocalization priorsource separationspherical harmonicsambisonics

Dane bibliometryczne

ID BaDAP154433
Data dodania do BaDAP2024-07-15
Tekst źródłowyURL
DOI10.1109/TASLP.2024.3399618
Rok publikacji2024
Typ publikacjiartykuł w czasopiśmie
Otwarty dostęptak
Czasopismo/seriaIEEE/ACM Transactions on Audio, Speech and Language Processing

Abstract

This article presents a Non-negative Tensor Factorization based method for sound source separation from Ambisonic microphone signals. The proposed method enables the use of prior knowledge about the Directions-of-Arrival (DOAs) of the sources, incorporated through a constraint on the Spatial Covariance Matrix (SCM) within a Maximum a Posteriori (MAP) framework. Specifically, this article presents a detailed derivation of four algorithms that are based on two types of cost functions, namely the squared Euclidean distance and the Itakura-Saito divergence, which are then combined with two prior probability distributions on the SCM, that is the Wishart and the Inverse Wishart. The experimental evaluation of the baseline Maximum Likelihood (ML) and the proposed MAP methods is primarily based on first-order Ambisonic recordings, using four different source signal datasets, three with musical pieces and one containing speech utterances. We consider underdetermined, determined, as well as over-determined scenarios by separating two, four and six sound sources, respectively. Furthermore, we evaluate the proposed algorithms for different spherical harmonic orders and at different reverberation time levels, as well as in non-ideal prior knowledge conditions, for increasingly more corrupted DOAs. Overall, in comparison with beamforming and a state-of-the-art separation technique, as well as the baseline ML methods, the proposed MAP approach offers superior separation performance in a variety of scenarios, as shown by the analysis of the experimental evaluation results, in terms of the standard objective separation measures, such as the SDR, ISR, SIR and SAR. IEEE

Publikacje, które mogą Cię zainteresować

fragment książki
#140050Data dodania: 6.5.2022
Wishart localization prior on spatial covariance matrix in ambisonic source separation using non-negative tensor factorization / Mateusz GUZIK, Konrad KOWALCZYK // W: ICASSP 2022 [Dokument elektroniczny] : 2022 IEEE International Conference on Acoustics, Speech, and Signal Processing : 7–13 May 2022, virtual, 22–27 May 2022, Singapore, satellite venue: Shenzhen, China : proceedings. — Wersja do Windows. — Dane tekstowe. — Piscataway : The Institute of Electrical and Electronics Engineers, cop. 2022. — ( Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing ; ISSN  1520-6149 ). — e-ISBN: 978-1-6654-0540-9. — S. 446–450. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 450, Abstr. — Publikacja dostępna online od: 2022-04-27
artykuł
#154436Data dodania: 15.7.2024
Reverberant source separation using NTF with delayed subsources and spatial priors / Mieszko FRAŚ, Konrad KOWALCZYK // IEEE/ACM Transactions on Audio, Speech and Language Processing ; ISSN 2329-9290. — Tytuł poprz.: IEEE Transactions on Audio, Speech, and Language Processing ; ISSN: 1558-7916. — 2024 — vol. 32, s. 1954–1967. — Bibliogr. s. 1966, Abstr. — Publikacja dostępna online od: 2024-03-06