Szczegóły publikacji
Opis bibliograficzny
Mask cross-attention transformer for robust exercise recognition under execution-speed domain shift / Kamil WARCHOŁ, Bogdan KWOLEK // W: Computational Science – ICCS 2026 workshops : 26th International Conference : Hamburg, Germany, June 29–July 1, 2026 : proceedings , Pt. 2 / eds. Maciej Paszynski, Amanda S. Barnard, Yongjie Jessica Zhang. — Cham : Springer, cop. 2026. — ( Lecture Notes in Computer Science ; ISSN 0302-9743 ; LNCS 16787 ). — ISBN: 978-3-032-29908-6; e-ISBN: 978-3-032-29909-3. — S. 443–457. — Bibliogr., Abstr. — Publikacja dostępna online od: 2026-06-28
Autorzy (2)
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 168919 |
|---|---|
| Data dodania do BaDAP | 2026-08-31 |
| DOI | 10.1007/978-3-032-29909-3_32 |
| Rok publikacji | 2026 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Wydawca | Springer |
| Konferencja | International Conference on Computational Science 2026 |
| Czasopismo/seria | Lecture Notes in Computer Science |
Abstract
Human action recognition has achieved remarkable progress on general-purpose benchmarks, yet the subproblem of short-time action recognition, where discriminative motion unfolds over brief and variable-duration intervals, remains underexplored. Video-based exercise recognition faces a critical challenge when deployment conditions differ from training environments. The execution speed variation introduces temporal domain shifts that degrade model performance. We propose the Mask Cross-Attention Transformer, a dual-stream architecture for short-time action recognition that conditions temporal reasoning on human-centric spatial priors through cross-attention between appearance features and per-frame human masks. By decoupling semantic motion patterns from execution tempo, the model achieves robust recognition across diverse scenarios. On the public MM-Fit benchmark, the model achieves test accuracy of 95.0% and macro-F1 of 88.8% on 11 exercise classes, surpassing recent multimodal approaches while using only grayscale video with automatically generated masks.