ARTFEED — Contemporary Art Intelligence

Semalith v1.4: 184M-Parameter Safety Classifier Outperforms Llama-Guard-3-8B on Prompt Injection

ai-technology · 2026-07-29

The release of Semalith v1.4 introduces a DeBERTa-v3-base classifier with 184 million parameters, capable of conducting three-axis safety classification—prompt injection, general harm, and compliance with financial services regulations—within a single forward pass. Its head features 22 classes, including BENIGN, nine types of prompt injection, general harm, and eleven BFSI labels. This model was trained using a 4-class auxiliary super-category head with jointly weighted loss on a corpus of 76,204 rows sourced from 49 public origins, employing SHA-1 deduplication against all held-out evaluation sets. Impressively, 21 out of 22 benchmarks exhibit zero contamination (maximum 0.22%). In comparisons with Llama-Guard-3-8B across 22 benchmarks, Semalith v1.4 outperforms in every prompt-injection assessment, filling a crucial gap in current open guardrails for financial services and agentic contexts.

Key facts

  • Semalith v1.4 is a 184M-parameter DeBERTa-v3-base classifier.
  • It performs simultaneous three-axis safety classification: prompt injection, general harm, and financial-services regulatory compliance.
  • The model has a 22-class head with BENIGN, nine prompt-injection sub-types, general-harm, and eleven BFSI labels.
  • Training used a 4-class auxiliary super-category head under jointly weighted loss.
  • Training corpus: 76,204 rows from 49 public sources with SHA-1 deduplication.
  • 21 of 22 benchmarks have zero contamination (max 0.22%).
  • Outperforms Llama-Guard-3-8B on all 22 held-out prompt-injection benchmarks.
  • Addresses a gap in open guardrails for financial-services and agentic settings.

Entities

Sources