ARTFEED — Contemporary Art Intelligence

LLM Validation of Pseudo-Science Varies by Model and Deployment

ai-technology · 2026-07-27

A study on arXiv (2607.22513) tested how four major LLM families—Claude, Grok, GPT, Gemini—evaluate ethnonationalist pseudo-science from Frank Salter's biosocial framework. Across four snapshots (October 2025–February 2026), Grok's Fast versions (powering X's default experience) assigned credibility scores of 70–75, two to five times higher than other models (15–40). This pattern was absent in control prompts on basic evolutionary consensus and refuted Lamarckian claims. Three additional findings: (1) a silent patch reversed Grok's behavior from chaotic to stably high validation overnight without documentation; (2) the same Grok model identifier showed different behavior across API and web interfaces; (3) GPT and Gemini exhibited inconsistent stances over time. The study highlights opaque epistemic mediation in LLM deployment configurations.

Key facts

  • Study tested Claude, Grok, GPT, Gemini on ethnonationalist pseudo-science from Frank Salter's biosocial framework
  • Four temporal snapshots from October 2025 to February 2026
  • Grok Fast versions scored 70-75 credibility, others 15-40
  • Control prompts showed comparable performance across models
  • Silent patch reversed Grok's behavior overnight without documentation
  • Same Grok model identifier behaved differently on API vs web
  • GPT and Gemini showed inconsistent stances over time
  • Study published on arXiv with ID 2607.22513

Entities

Institutions

  • arXiv
  • X

Sources