# Wf Sandbox Testing

> wf skill-eval harness pack: the sandbox-testing feature capability. Runs real headless wf: skill invocations hermetically in a fingerprinted container (runner/), judges the runner's structural outputs — terminal block, workspace file set, invoked contract-op set — statistically over N runs and emits a variance-aware report separating drift from regression (assert/), ships the SMOKE/STATISTICAL tier commands and the baseline-comparison primitive, the behavioral-regression corpus mined retrofit-f

## Facts
- Page: https://tashan.sh/capability/plugin-pavel-rp-wf-plugin-wf-sandbox-testing
- tashan id: plugin:pavel-rp/wf-plugin/wf-sandbox-testing
- Source: https://github.com/pavel-rp/wf-plugin
- Type: plugin
- Category: devtools
- tashan score: 35.0 / 100
- Adoption: 7.0
- Upkeep: 61.0
- Freshness: 92.0
- Evidence coverage: 84% of the inputs this score can use
- Health: active
- Instruction depth: not yet graded
- Official: no

## Install

```sh
/plugin marketplace add pavel-rp/wf-plugin
/plugin install wf-sandbox-testing@wf-marketplace
```

## Security audit
Not scanned. We audit npm-published capabilities; this one has no npm package we can resolve, or has not reached the queue. This is not a clean bill of health.

---
Measured 2026-09-13 by tashan (https://tashan.sh) from public evidence. Scorer s5.
