Comprehensive Evaluation Frameworks and Future Directions: Current State and Challenges of Text-to-Image Generation Models

Authors

  • Linpeng Shang

DOI:

https://doi.org/10.54097/njeaxf87

Keywords:

Text-to-Image Generation; Performance Evaluation; Human Preferences; Comprehensive Benchmarks; Generative AI.

Abstract

Text-to-image (T2I) generation models have great innovations in generative AI. However, evaluating these models could still be regarded as a crucial challenge. Traditional metrics include Frechet Inception Distance (FID) and CLIPScore, and they concentrate on image quality and text-image alignments. However, they often fail to address fine-grained details, for example, spatial relationships and attribute binding. Recent frameworks, for example, GENEVAL, TIFA, HEIM, and HRS-Bench, have involved multidimensional evaluation methods. They have combined automated metrics and human preference-based approaches. Although these frameworks improve interpretability and comprehensiveness, they still face limitations. To be more specific, they are not good enough in scalability or standardization. Moreover, they lack emphasis on social and ethical considerations. These methods often do not have enough granularity. This paper will analyze the advantages and disadvantages of the current evaluation methods, and highlight the needs for modular, interpretable, and scalable systems. They integrate automated metrics with selective human feedback. By dealing with these challenges, T2I models can better cater to people’s expectations. In addition, they support various applications in art and data generation. This study would offer practical perspectives for promoting progress in the evaluation and development of next-generation T2I models.

Downloads

Download data is not yet available.

References

[1] Tang R, Liu L, Pandey A, Jiang Z, Yang G, Kumar K, et al. What the daam: Interpreting stable diffusion using cross attention. arXiv preprint arXiv:2210.04885. 2022.

[2] Hessel J, Holtzman A, Forbes M, Bras R L, & Choi Y. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718. 2021.

[3] Ghosh D, Hajishirzi H, Schmidt L. GENEVAL: An object-focused framework for evaluating text-to-image alignment. Proceedings of NeurIPS. 2023.

[4] Hu Y, Liu B, Kasai J, Wang Y, Ostendorf M, Krishna R, Smith N A. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 20406-20417). 2023.

[5] Wu H. Human Preference Score: Better aligning text-to-image models with human preference. Proceedings of ICCV. 2023.

[6] Xu J, Liu X, Wu Y, Tong Y, Li Q, Ding M, Tang J, Dong Y. ImageReward: Learning and evaluating human preferences for text-to-image generation. Proceedings of NeurIPS. 2023.

[7] Lee T, Yasunaga M, Meng C, Mai Y, Park JS, Gupta A, et al. Holistic Evaluation of Text-to-Image Models (HEIM). Proceedings of NeurIPS. 2023.

[8] Bakr E M, Sun P, Shen X, Khan F F, Li L E, Elhoseiny M. HRS-Bench: Holistic, reliable and scalable benchmark for text-to-image models. Proceedings of ICCV. 2023.

[9] Jayasumana S, Ramalingam S, Veit A, Glasner D, Chakrabarti A, Kumar S. Rethinking fid: Towards a better evaluation metric for image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9307-9315). 2024.

[10] Yarom M, Bitton Y, Changpinyo S, Aharoni R, Herzig J, Lang O, Ofek E, Szpektor I. What you see is what you read? Improving text-image alignment evaluation. Proceedings of NeurIPS. 2023.

[11] Petsiuk V, Siemenn A, Surbehera S, Chin Z, Tyser K, Hunter G, et al. Human evaluation of text-to-image models on a multi-task benchmark. NeurIPS 2022 Workshop on Human Evaluation of Generative Models. 2022.

Downloads

Published

11-05-2025

How to Cite

Shang, L. (2025). Comprehensive Evaluation Frameworks and Future Directions: Current State and Challenges of Text-to-Image Generation Models. Highlights in Science, Engineering and Technology, 138, 96-101. https://doi.org/10.54097/njeaxf87