ARTFEED — Contemporary Art Intelligence

OliveGemma: A Vision Language Model for Mediterranean Diet Recognition

ai-technology · 2026-08-06

A new vision language model named OliveGemma has been developed to identify and analyze Mediterranean and European dishes, tackling the complexities of detailed food recognition in image-based dietary evaluations. Utilizing the open-weight PaliGemma-2-3B framework, OliveGemma underwent fine-tuning via LoRA with a consolidated dataset of 17,340 images from three European research initiatives: MedGR, ODIN, and VIPPSTAR. This model integrates these datasets into 216 distinct dish categories and includes 102,642 instructional question-answer pairs related to dish identification, ingredients, visual evidence, and overall food comprehension. In a 3-fold cross-validation test, OliveGemma achieved a top-1 accuracy of 92.96% ± 0.91%, outperforming the leading CNN baseline (DensNet). The findings, published as arXiv preprint 2608.03428, suggest that vision language models could serve as effective alternatives to traditional food diaries, improving dietary assessments and nutritional research, especially in Mediterranean regions.

Key facts

  • OliveGemma is a vision language model for recognizing Mediterranean and European cuisine.
  • It is built on the PaliGemma-2-3B architecture and fine-tuned with LoRA.
  • Training data includes 17,340 images from MedGR, ODIN, and VIPPSTAR datasets.
  • The model covers 216 composed dish categories and 102,642 instruction-style QA items.
  • Achieves 92.96% ± 0.91% top-1 accuracy, exceeding CNN baselines.
  • Aims to provide scalable alternative to self-reported food diaries.
  • Released as arXiv preprint 2608.03428.
  • Supports dish recognition, ingredient identification, and class boundary discrimination.

Entities

Institutions

  • arXiv

Sources