Szczegóły publikacji

Opis bibliograficzny

AI-driven image generation: algorithms, architectures, quality assessment, and applications — a structured narrative review / Mikołaj LESZCZUK, Yi Zhang, Mylène C. Q. Farias, Damon M. Chandler, Ruth Kalola // Electronics [Dokument elektroniczny]. — Czasopismo elektroniczne ; ISSN  2079-9292 . — 2026 — vol. 15 iss. 13 art. no. 2891, s. 1–69. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 63–69, Abstr. — Publikacja dostępna online od: 2026-07-01

Autorzy (5)

Słowa kluczowe

diffusion modelsAI driven image generationtext-to-image synthesisstructured narrative reviewmulti-modal learningfoundation modelsimage quality assessmentGenerative Adversarial Networks

Dane bibliometryczne

ID BaDAP169099
Data dodania do BaDAP2026-07-31
Tekst źródłowyURL
DOI10.3390/electronics15132891
Rok publikacji2026
Typ publikacjiprzegląd
Otwarty dostęptak
Creative Commons
Czasopismo/seriaElectronics

Abstract

Background: This paper presents a structured narrative review of recent advances in AI-driven image generation across four complementary perspectives: generative models and architectures, quality assessment and performance metrics, application domains, and multi-modal or cross-lingual extensions. The review aimed to identify dominant methodological trends, representative evaluation practices, and open research challenges in contemporary image generation research. Methods: A structured literature search was conducted in the Scopus database on 29 January 2026 using a predefined query focused on modern generative-image paradigms and excluding clearly out-of-scope domains. Eligible records addressed contemporary AI-driven image generation or closely related multi-modal generation settings within the temporal and topical scope of the search. Retrieved records were first assigned to four thematic branches and then screened with branch-specific relevance criteria for narrative synthesis. Results: The search returned 1524 records, and the final narrative synthesis included 117 publications: 34 on generative models, 23 on quality assessment, 29 on applications, and 31 on multi-modal and cross-lingual aspects. Across the reviewed literature, progress was shaped not only by visual fidelity, but also by controllability, semantic grounding, human-centred evaluation, multi-modal integration, and practical deployment constraints. Limitations: The review was limited to a single primary bibliographic source and to a qualitative narrative synthesis without meta-analysis. Conclusions: The review provides a structured reference point for researchers and practitioners working on AI-based image generation and its evaluation, while also highlighting benchmark, comparability, and multi-modal-transfer challenges that remain unresolved.