Szczegóły publikacji
Opis bibliograficzny
Sign language recognition with Visual State Space Model / Bogdan KWOLEK // W: CVPR 2026 [Dokument elektroniczny] : IEEE/CVF Computer Vision and Pattern Recognition : 3-7 June 2026, Denver, [USA] : proceedings. — Wersja do Windows. — Dane tekstowe. — [Denver] : Computer Vision Foundation, [2026]. — S. 10520–10527. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 10526–10527, Abstr.
Autor
Dane bibliometryczne
| ID BaDAP | 168195 |
|---|---|
| Data dodania do BaDAP | 2026-07-09 |
| Tekst źródłowy | URL |
| Rok publikacji | 2026 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Creative Commons | |
| Konferencja | IEEE Conference on Computer Vision and Pattern Recognition 2026 |
Abstract
In this work, we propose a novel approach to isolated sign language recognition based on sequences of RGB images. Each dynamic gesture is represented using a Tree Structure Skeleton Image (TSSI), constructed from the signers skeletal joint configurations. The keypoints for each video are extracted in advance using the MediaPipe framework. Subsequently, a Visual State Space Model is trained on the resulting TSSI representations to perform gesture recognition. By transforming raw video frames into hierarchical skeletal structures, the model acquires a more structured understanding of gesture topology, which facilitates improved interpretation of complex sign patterns. Experimental results on the WLASL-100 dataset demonstrate that the proposed method achieves competitive performance relative to current state-of-the-art approaches.