MermaidSeqBench: New Benchmark for Evaluating LLMs in Generating Mermaid Sequence Diagrams
MermaidSeqBench has been launched by researchers as a novel benchmark aimed at assessing large language models (LLMs) in their ability to create Mermaid sequence diagrams from natural language inputs. This initiative responds to the current absence of evaluation tools for this specific task, which has obstructed thorough comparisons of model performance. The benchmark includes 132 samples developed through a combination of human-verified flows, LLM enhancements, and rule-based expansions. An LLM-as-a-judge model is utilized to evaluate the quality of the generated diagrams based on criteria like syntax accuracy, error management, activation handling, and usability. Detailed in a paper on arXiv (arXiv:2511.14967), this advancement holds considerable importance for the software engineering field, as Mermaid diagrams are crucial for visualizing interactions, potentially improving documentation and design workflows. The benchmark seeks to establish a robust standard for future advancements in this domain.
Key facts
- MermaidSeqBench is a new benchmark for evaluating LLMs in generating Mermaid sequence diagrams.
- The benchmark consists of 132 samples.
- Samples were developed via hybrid methodology: human-verified flows, LLM-based augmentation, and rule-based expansion.
- Evaluation uses LLM-as-a-judge model.
- Metrics include syntax correctness, activation handling, error handling, and practical usability.
- The paper is available on arXiv with ID 2511.14967.
- The announcement type is replace-cross.
- The benchmark addresses the lack of existing benchmarks for NL-to-Mermaid diagram generation.
Entities
Institutions
- arXiv