ARTFEED — Contemporary Art Intelligence

Audio Foundation Models Recover Phylogenetic Signal from Marine Mammal and Bird Vocalizations

ai-technology · 2026-07-27

A new study published on arXiv (2607.22458) shows that powerful pretrained audio models can detect evolutionary relationships based on animal sounds without needing specialized training. The researchers analyzed four models—AST, CLAP, BEATs-bio, and BirdNET—using Mantel tests across two datasets. They focused on 32 marine mammal species, pulling from 1,754 recordings in the Watkins Marine Mammal Sound Database. The models identified strong phylogenetic signals in 26 cetaceans, with CLAP and BEATs-bio scoring r=0.82, and AST at r=0.74 (all p<0.001). In comparison, manually crafted MFCC features (105 dimensions) revealed no significant correlation (r=0.040, p=0.338), indicating that the learned models capture deeper biological insights beyond simple labels.

Key facts

  • Study probes four pretrained audio models (AST, CLAP, BEATs-bio, BirdNET) for phylogenetic signal in vocalizations.
  • Mantel tests used on marine mammal and bird datasets.
  • 32 marine mammal species (1,754 recordings) from Watkins Marine Mammal Sound Database.
  • Strong signal in 26 cetaceans: CLAP r=0.82, BEATs-bio r=0.82, AST r=0.74 (all p<0.001).
  • MFCC features (105d) show no correlation (r=0.040, p=0.338).
  • Gap persists after PCA projection to 105 dimensions.
  • Published on arXiv with ID 2607.22458.
  • Study suggests audio embeddings encode biological structure beyond training labels.

Entities

Institutions

  • arXiv
  • Watkins Marine Mammal Sound Database

Sources