CityRiSE: Enhancing LVLMs for Urban Socio-Economic Reasoning via Reinforcement Learning
A novel framework named CityRiSE has been introduced to enhance Large Vision-Language Models' (LVLMs) capability to analyze urban socio-economic conditions using visual information. This framework is outlined in an arXiv publication (arXiv:2510.22282v2) and utilizes reinforcement learning (RL) alongside a meticulously assembled multi-modal dataset and a design for verifiable rewards. CityRiSE directs LVLMs to concentrate on visually significant cues, facilitating structured and objective-driven reasoning for predicting socio-economic status. Experimental results indicate that CityRiSE, featuring advanced reasoning abilities, surpasses current methodologies. This study tackles the issue of urban socio-economic sensing, crucial for promoting global sustainable development objectives. The paper was released as a replace-cross update on arXiv.
Key facts
- CityRiSE is a novel framework for reasoning urban socio-economic status in LVLMs via reinforcement learning.
- The framework uses a curated multi-modal dataset and verifiable reward design.
- It guides LVLMs to focus on semantically meaningful visual cues.
- Experiments show significant improvement over existing methods.
- The paper is available on arXiv with ID 2510.22282.
- The research addresses urban socio-economic sensing for sustainable development goals.
- The announcement type is replace-cross.
- The approach enables structured and goal-oriented reasoning.
Entities
Institutions
- arXiv