Image Aesthetic Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance
By: Zhiyuan Hu , Zheng Sun , Yi Wei and more
Potential Business Impact:
Teaches computers to judge if pictures look good.
The performance of image generation has been significantly improved in recent years. However, the study of image screening is rare and its performance with Multimodal Large Language Models (MLLMs) is unsatisfactory due to the lack of data and the weak image aesthetic reasoning ability in MLLMs. In this work, we propose a complete solution to address these problems in terms of data and methodology. For data, we collect a comprehensive image screening dataset with over 128k samples, about 640k images. Each sample consists of an original image, four generated images. The dataset evaluates the image aesthetic reasoning ability under four aspects: appearance deformation, physical shadow, placement layout, and extension rationality. Regarding data annotation, we investigate multiple approaches, including purely manual, fully automated, and answer-driven annotations, to acquire high-quality chains of thought (CoT) data in the most cost-effective manner. Methodologically, we introduce a Hard Cases Mining (HCM) strategy with a Dynamic Proportional Accuracy (DPA) reward into the Group Relative Policy Optimization (GRPO) framework, called HCM-GRPO. This enhanced method demonstrates superior image aesthetic reasoning capabilities compared to the original GRPO. Our experimental results reveal that even state-of-the-art closed-source MLLMs, such as GPT4o and Qwen-VL-Max, exhibit performance akin to random guessing in image aesthetic reasoning. In contrast, by leveraging the HCM-GRPO, we are able to surpass the scores of both large-scale open-source and leading closed-source models with a much smaller model.
Similar Papers
Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
CV and Pattern Recognition
Helps pick the best online ads using AI.
REASONEDIT: Towards Reasoning-Enhanced Image Editing Models
CV and Pattern Recognition
Makes AI better at changing pictures with words.
Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
CV and Pattern Recognition
Helps AI "see" and "think" about pictures better.