AI Research Proposes Multimodal Follow-Up Suggestions for Image-Editing Conversations
A recent study published on arXiv (2608.07565) explores the largely neglected topic of follow-up editing suggestions in conversational image-creation systems. Researchers gathered 100,000 authentic multi-turn conversation samples from the Qwen App, revealing that 80.1% of these exchanges are reliant on images, underscoring the necessity for multimodal recommendations. They introduce a three-phase framework: initially, compiling a human-reviewed list of suitable follow-up editing intents based on real online data and refining a multimodal policy through supervised fine-tuning (SFT); next, enhancing the policy using user click feedback via multi-objective reinforcement learning to align suggestions with user preferences; and finally, (the abstract is truncated, likely relating to evaluation or implementation). This research seeks to enhance how conversational assistants suggest follow-up edits that resonate with user preferences, provide varied options, and are feasible with the existing image. The paper can be accessed at https://arxiv.org/abs/2608.07565.
Key facts
- Paper arXiv:2608.07565 addresses follow-up edit suggestions in conversational image-creation systems.
- Collected 100,000 real multi-turn image-creation conversation samples from Qwen App.
- 80.1% of interactions are image-dependent, underscoring need for multimodal recommendation.
- Proposes a three-stage framework: SFT, multi-objective reinforcement learning, and more.
- Aims to make follow-up suggestions reflect user preferences, offer diverse directions, and be executable.
- Published on arXiv with abstract available at https://arxiv.org/abs/2608.07565.
Entities
Institutions
- Qwen App
- arXiv