Szczegóły publikacji
Opis bibliograficzny
RL-Exponential-DAS: exponential decision schedules for dynamic algorithm selection / Władysław NIEĆ, Wojciech ACHTELIK, Hubert GUZOWSKI, Maciej SMOŁKA, Jacek MAŃDZIUK // W: Parallel Problem Solving from Nature – PPSN XIX : 19th international conference : PPSN 2026 : Trento, Italy, August 29–September 2, 2026 : proceedings , Pt. 5 / eds. Giovanni Iacca, [et al.]. — Cham : Springer Nature, cop. 2027. — ( Lecture Notes in Computer Science ; ISSN 0302-9743 ; vol. 16989 ). — ISBN: 978-3-032-36216-2; e-ISBN: 978-3-032-36217-9. — S. 152–168. — Bibliogr., Abstr. — Publikacja dostępna online od: 2026-08-25. — J. Mańdziuk - dod. afiliacja: Warsaw University of Technology
Autorzy (5)
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 170355 |
|---|---|
| Data dodania do BaDAP | 2026-10-06 |
| DOI | 10.1007/978-3-032-36217-9_10 |
| Rok publikacji | 2027 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Wydawca | Springer |
| Konferencja | Parallel Problem Solving from Nature 2026 |
| Czasopismo/seria | Lecture Notes in Computer Science |
Abstract
Dynamic Algorithm Selection (DAS) in meta-black-box optimization is fundamentally constrained when the underlying optimizer portfolio is structurally heterogeneous: switching between optimizers with incompatible internal states disrupts the adaptive mechanisms on which late-stage convergence critically depends. Prior reinforcement-learning-based DAS methods largely sidestep this difficulty by restricting selection to portfolios drawn from a single algorithmic family. We revisit DAS from a complementary standpoint and investigate whether the temporal placement of selection decisions, rather than their content alone, constitutes an exploitable axis of design. We introduce RL-Exponential-DAS, a PPO-based framework in which decision checkpoints follow an exponentially increasing schedule, inducing frequent early interventions and progressively longer uninterrupted optimization intervals as the search matures. The framework is coupled with a memory-handling mechanism that admits selection over a structurally heterogeneous portfolio comprising G3PCX, SPSO, and LMCMAES, whose internal states admit no common representation. On the noiseless BBOB benchmark, exponential scheduling yields statistically significant AOCC improvements over uniform scheduling in dimensions D∈{5,10}. A policy-level analysis reveals an interpretable exploration-to-exploitation structure that sharpens with dimensionality: early checkpoints are dominated by SPSO-driven exploration, while late checkpoints concentrate mass on LMCMAES.