ARTFEED — Contemporary Art Intelligence

Frontier LLMs vs. Natively Multimodal Embeddings: A Hard-Negative Text-to-Image Retrieval Comparison

ai-technology · 2026-08-13

A recent study published on arXiv (2608.11343) marks the inaugural direct assessment of native multimodal embeddings in contrast to LLM-based visual ranking for text-to-image retrieval. Released in 2026, the research analyzes Google's Gemini Embedding 2, the pioneering natively multimodal embedding model that integrates text, images, video, audio, and documents into a unified space. It compares this model with leading large language models (LLMs) like GPT-4.1 and Claude Sonnet 4.6, utilizing the Flickr30k dataset with challenging retrieval tasks. Results reveal that GPT-4.1 and Claude Sonnet 4.6 match Gemini Embedding 2's performance, while multimodal embeddings excel in situations where embeddings are precomputed, providing efficiency benefits. The findings underscore the intensifying rivalry in multimodal retrieval systems and LLMs' potential as zero-shot rankers.

Key facts

  • Study compares native multimodal embeddings with LLM-based visual ranking on Flickr30k.
  • Gemini Embedding 2 is Google's first natively multimodal embedding model.
  • GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2.
  • Multimodal embeddings are better suited when embeddings are precomputed.
  • Research is available on arXiv with ID 2608.11343.
  • The study addresses hard-negative text-to-image retrieval.
  • The release date of Gemini Embedding 2 is March 2026.
  • The study is the first direct comparison of its kind.

Entities

Institutions

  • Google
  • arXiv

Sources