‹ The Index

Bench QA

plugin

Quality-assurance automation for the diolog-swe-bench benchmark corpus: judge whether a task is difficult enough and fairly verified (bench-task-judge), re-run the whole-corpus review with any model (bench-corpus-review), and author new gated tasks including ui design tasks (bench-task-author).

Works with: Claude Code (native)
native: this artifact type is that client's own format

Install (Claude Code):

/plugin marketplace add DiologIR/diolog-plugins
/plugin install bench-qa@diolog-plugins

Security audit

Not scanned yet. We audit npm-published capabilities for known advisories, install-time scripts and permission surface; this one has no npm package we can resolve, or has not reached the queue.

You searched for one. Check the rest of your stack:

npx tashan-cli doctor

Reads the config already on your machine and names what is dead, deprecated or running code at install time. No account, nothing uploaded.

tashan Pro$6/mo

Bench QA scores 38 today. Pro keeps the series, so you can see whether that is a project getting better or one on its way down.

Start a 7-day trial › Everything measured on this page stays free.

Show your score

Measured this well? Put the live badge in your README — it updates as the score does.

tashan badge for Bench QA
[![tashan](https://tashan.sh/badge/plugin-diologir-diolog-plugins-bench-qa.svg)](https://tashan.sh/capability/plugin-diologir-diolog-plugins-bench-qa.html)

source ↗  ·  plugin:diologir/diolog-plugins/bench-qa

Everything on this page is public evidence and free. What it cannot know is whether you run this — check your whole config, free, in the browser. tashan Pro adds the series behind each row and names a replacement for anything dying.

Already running this? Check your whole config — free, in your browser, nothing installed. Or npx tashan-cli doctor locally, which sends nothing at all.

Measured 2026-09-13  ·  scorer s5  ·  how  ·  something wrong here?