# Bench QA

> Quality-assurance automation for the diolog-swe-bench benchmark corpus: judge whether a task is difficult enough and fairly verified (bench-task-judge), re-run the whole-corpus review with any model (bench-corpus-review), and author new gated tasks including ui design tasks (bench-task-author).

## Facts
- Page: https://tashan.sh/capability/plugin-diologir-diolog-plugins-bench-qa
- tashan id: plugin:diologir/diolog-plugins/bench-qa
- Source: https://github.com/DiologIR/diolog-plugins
- Type: plugin
- Category: other
- tashan score: 38.0 / 100
- Adoption: 7.0
- Upkeep: 78.0
- Freshness: 91.0
- Evidence coverage: 84% of the inputs this score can use
- Health: active
- Instruction depth: not yet graded
- License: MIT
- Official: no

## Install

```sh
/plugin marketplace add DiologIR/diolog-plugins
/plugin install bench-qa@diolog-plugins
```

## Security audit
Not scanned. We audit npm-published capabilities; this one has no npm package we can resolve, or has not reached the queue. This is not a clean bill of health.

---
Measured 2026-09-13 by tashan (https://tashan.sh) from public evidence. Scorer s5.
