TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation
By: Yuheng Feng , Jianhui Wang , Kun Li and more
Potential Business Impact:
Makes AI pictures match your exact ideas.
Although text-to-image generation technologies have made significant advancements, they still face challenges when dealing with ambiguous prompts and aligning outputs with user intent.Our proposed framework, TDRI (Two-Phase Dialogue Refinement and Co-Adaptation), addresses these issues by enhancing image generation through iterative user interaction. It consists of two phases: the Initial Generation Phase, which creates base images based on user prompts, and the Interactive Refinement Phase, which integrates user feedback through three key modules. The Dialogue-to-Prompt (D2P) module ensures that user feedback is effectively transformed into actionable prompts, which improves the alignment between user intent and model input. By evaluating generated outputs against user expectations, the Feedback-Reflection (FR) module identifies discrepancies and facilitates improvements. In an effort to ensure consistently high-quality results, the Adaptive Optimization (AO) module fine-tunes the generation process by balancing user preferences and maintaining prompt fidelity. Experimental results show that TDRI outperforms existing methods by achieving 33.6% human preference, compared to 6.2% for GPT-4 augmentation, and the highest CLIP and BLIP alignment scores (0.338 and 0.336, respectively). In iterative feedback tasks, user satisfaction increased to 88% after 8 rounds, with diminishing returns beyond 6 rounds. Furthermore, TDRI has been found to reduce the number of iterations and improve personalization in the creation of fashion products. TDRI exhibits a strong potential for a wide range of applications in the creative and industrial domains, as it streamlines the creative process and improves alignment with user preferences
Similar Papers
Test-time Prompt Refinement for Text-to-Image Models
Machine Learning (CS)
Fixes AI art mistakes by checking its own work.
DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
CV and Pattern Recognition
Finds exact pictures by talking with you.
Twin Co-Adaptive Dialogue for Progressive Image Generation
CV and Pattern Recognition
Makes computer pictures match your ideas better.