Comprehensive Evaluation Frameworks and Future Directions: Current State and Challenges of Text-to-Image Generation Models
DOI:
https://doi.org/10.54097/njeaxf87Keywords:
Text-to-Image Generation; Performance Evaluation; Human Preferences; Comprehensive Benchmarks; Generative AI.Abstract
Text-to-image (T2I) generation models have great innovations in generative AI. However, evaluating these models could still be regarded as a crucial challenge. Traditional metrics include Frechet Inception Distance (FID) and CLIPScore, and they concentrate on image quality and text-image alignments. However, they often fail to address fine-grained details, for example, spatial relationships and attribute binding. Recent frameworks, for example, GENEVAL, TIFA, HEIM, and HRS-Bench, have involved multidimensional evaluation methods. They have combined automated metrics and human preference-based approaches. Although these frameworks improve interpretability and comprehensiveness, they still face limitations. To be more specific, they are not good enough in scalability or standardization. Moreover, they lack emphasis on social and ethical considerations. These methods often do not have enough granularity. This paper will analyze the advantages and disadvantages of the current evaluation methods, and highlight the needs for modular, interpretable, and scalable systems. They integrate automated metrics with selective human feedback. By dealing with these challenges, T2I models can better cater to people’s expectations. In addition, they support various applications in art and data generation. This study would offer practical perspectives for promoting progress in the evaluation and development of next-generation T2I models.
Downloads
References
[1] Tang R, Liu L, Pandey A, Jiang Z, Yang G, Kumar K, et al. What the daam: Interpreting stable diffusion using cross attention. arXiv preprint arXiv:2210.04885. 2022.
[2] Hessel J, Holtzman A, Forbes M, Bras R L, & Choi Y. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718. 2021.
[3] Ghosh D, Hajishirzi H, Schmidt L. GENEVAL: An object-focused framework for evaluating text-to-image alignment. Proceedings of NeurIPS. 2023.
[4] Hu Y, Liu B, Kasai J, Wang Y, Ostendorf M, Krishna R, Smith N A. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 20406-20417). 2023.
[5] Wu H. Human Preference Score: Better aligning text-to-image models with human preference. Proceedings of ICCV. 2023.
[6] Xu J, Liu X, Wu Y, Tong Y, Li Q, Ding M, Tang J, Dong Y. ImageReward: Learning and evaluating human preferences for text-to-image generation. Proceedings of NeurIPS. 2023.
[7] Lee T, Yasunaga M, Meng C, Mai Y, Park JS, Gupta A, et al. Holistic Evaluation of Text-to-Image Models (HEIM). Proceedings of NeurIPS. 2023.
[8] Bakr E M, Sun P, Shen X, Khan F F, Li L E, Elhoseiny M. HRS-Bench: Holistic, reliable and scalable benchmark for text-to-image models. Proceedings of ICCV. 2023.
[9] Jayasumana S, Ramalingam S, Veit A, Glasner D, Chakrabarti A, Kumar S. Rethinking fid: Towards a better evaluation metric for image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9307-9315). 2024.
[10] Yarom M, Bitton Y, Changpinyo S, Aharoni R, Herzig J, Lang O, Ofek E, Szpektor I. What you see is what you read? Improving text-image alignment evaluation. Proceedings of NeurIPS. 2023.
[11] Petsiuk V, Siemenn A, Surbehera S, Chin Z, Tyser K, Hunter G, et al. Human evaluation of text-to-image models on a multi-task benchmark. NeurIPS 2022 Workshop on Human Evaluation of Generative Models. 2022.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Highlights in Science, Engineering and Technology

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







