Szczegóły publikacji
Opis bibliograficzny
A framework for large-scale synthetic graph dataset generation / Sajad Darabi, Piotr Bigaj, Dawid Majchrowski, Artur Kasymov, Paweł MORKISZ, Alex Fit-Florea // IEEE Transactions on Neural Networks and Learning Systems ; ISSN 2162-237X . — 2025 — vol. 36 iss. 8, s. 14258–14268. — Bibliogr. s. 14267, Abstr. — Publikacja dostępna online od: 2025-03-27. — P. Morkisz - dod. afiliacja: NVIDIA, Santa Clara, USA
Autorzy (6)
- Darabi Sajad
- Bigaj Piotr
- Majchrowski Dawid
- Kasymov Artur
- AGHMorkisz Paweł M.
- Fit-Florea Alex
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 162656 |
|---|---|
| Data dodania do BaDAP | 2025-09-19 |
| Tekst źródłowy | URL |
| DOI | 10.1109/TNNLS.2025.3540392 |
| Rok publikacji | 2025 |
| Typ publikacji | artykuł w czasopiśmie |
| Otwarty dostęp | |
| Czasopismo/seria | IEEE Transactions on Neural Networks and Learning Systems |
Abstract
Recently, there has been increasing interest in developing and deploying deep graph learning algorithms for various tasks, such as fraud detection and recommender systems. However, there is a limited number of publicly available graph-structured datasets, most of which are small compared with production-sized applications or limited in their application domain. In this work, we tackle this shortcoming by proposing a synthetic graph generation tool that enables scaling datasets to production-size graphs with trillions of edges and billions of nodes. The proposed method comprises a series of parametric models that can either be randomly initialized or fit to proprietary datasets. These models can then be released to researchers to study graph methods on the synthetic data, facilitating prototype development and novel applications. We demonstrate the generalizability of the framework across various datasets, mimicking their structural and feature distributions, as well as the ability to scale them to varying sizes, demonstrating their usefulness for benchmarking and model development. Code can be found on GitHub.