VARM-Bench: New Benchmark for Verifiable Reasoning in Chinese Abusive Speech Moderation
Researchers have created a new tool called VARM-Bench to evaluate how effectively structured reasoning can be applied to moderating abusive language in Chinese. This development comes in response to the rising issue of harmful content on Chinese social media and points out the limitations of current benchmarks, which lack a solid framework for verifying moderation decisions. VARM-Bench features detailed rationales for six key decisions: identifying the target, the type of target, how explicit the target is, the author’s stance, harmfulness labeling, and detailed categorization. It uses a set process to assess various aspects like correctness and agreement, without depending on a language model for judgment. You can find more details in a paper on arXiv (identifier 2608.15600).
Key facts
- VARM-Bench is a new benchmark for Chinese abusive speech moderation.
- It provides field-anchored chain-of-thought rationales.
- Each instance includes six decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category.
- The benchmark uses a deterministic protocol without an LLM judge.
- It evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors.
- The paper is available on arXiv with identifier 2608.15600.
- Existing benchmarks lack unified representation for verifying moderation decisions.
- The benchmark aims to improve reliability of Chinese social-media text moderation.
Entities
Institutions
- arXiv