Szczegóły publikacji
Opis bibliograficzny
Hardware implementation of a bfloat16 exponential function for softmax computation / Radosław FEIGLEWICZ, Andrzej KOS // W: MIXDES 2026 [Dokument elektroniczny] : 2026 33rd international conference on Mixed Design of integrated circuits and systems : 25–26 June 2026, [Poznań, Poland] / ed. Wojciech Tylman. — Wersja do Windows. — Dane tekstowe. — Łódź : Lodz University of Technology ; IEEE, [2026]. — Dod. ISBN: 978-83-63578-29-9, 979-8-3195-1863-7. — e-ISBN: 978-83-63578-30-5. — S. 366–369. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 369, Abstr. — Publikacja dostępna online od: 2026-08-04
Autorzy (2)
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 169640 |
|---|---|
| Data dodania do BaDAP | 2026-09-28 |
| Tekst źródłowy | URL |
| DOI | 10.23919/MIXDES69535.2026.11617227 |
| Rok publikacji | 2026 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Wydawcy | Institute of Electrical and Electronics Engineers (IEEE), Politechnika Łódzka |
Abstract
Transformer-based artificial intelligence models are increasingly deployed in mobile robotic systems. Many applications require computations to be performed locally under an edge computing paradigm, which necessitates the development of customized AI accelerators tailored for robotics to increase throughput, reduce latency, and minimize energy consumption. This paper presents a hardware-efficient implementation of the exponential function based on piecewise linear (PWL) approximation to accelerate softmax computation within the attention mechanism during transformer inference. The proposed design targets resource-constrained edge devices used in robotic platforms. The implementation was evaluated using the MMLU-Pro benchmark, comparing the proposed custom solution operating in bfloat16 (brain float 16) precision with a float32 reference implementation. The results demonstrate that the proposed approach achieves inference accuracy comparable to the float32 baseline while reducing computational complexity. Furthermore, the design was synthesized using high-level synthesis (HLS) to estimate FPGA resource utilization, providing insight into the feasibility and efficiency of the proposed accelerator for edge robotic applications.