EduPluginBench: Executable Assurance for AI-Generated Educational Plugins
A new standard called EduPluginBench has been created to ensure that AI-generated code plugins adhere to governance guidelines like least privilege and telemetry consent. This standard was explained in a paper posted on arXiv (arXiv:2608.00739v1) and uses a phased approach for admitting plugins in regulated software settings. The P0-P4 stages led to a 74.7 percentage point increase in catching release-blocking defects compared to the P0-P2 stages, with no rejections from 120 clean references. In a study with 600 unmodified outputs from two coding models, only 300 parsed, and none met the P0 or P0-P4 criteria. This research is available on arXiv.
Key facts
- EduPluginBench is an executable benchmark for AI-generated educational plugins.
- It checks compliance with least privilege, telemetry consent, provenance, privileged-write authority, lifecycle constraints, and bounded failure.
- The benchmark uses a staged admission method (P0-P4).
- Across 1,440 activation-checked first-order mutants from 30 specifications, P0-P4 increased release-blocking-defect recall by 74.7 percentage points over P0-P2.
- No rejection was observed among 120 clean references (95% Wilson upper bound 3.1%).
- A frozen transfer study of 600 unmodified generations from two coding models found 300/600 parsed, but none passed P0 or P0-P4 conformance.
- The paper is available on arXiv with identifier 2608.00739v1.
- The study includes an independently labelled Moodle study (details cut off).
Entities
Institutions
- arXiv
- Moodle