# Eval Harness

> | [WHAT] Bias-aware evaluation harness — orchestrate fixture runs across 19 agents + detect 4 bias types (sycophancy, anchoring, verbosity, pattern-bias, position) + cross-LLM judge methodology. [AUDIENCE] Maintainer (offline — sole runner; Evan agent reverted in v3.3); Patrick (retro consumer); Stan (cross-team calibration). [WHEN] Pre-major-release (v3→v4 prompt refactor); Phase 4 Triage (PEV loop, per-bd) ถ้า disputerate 20% per agent (per drift M3); ก่อน promote prompt change to default. [TRIGGER] /shode-house:eval-harness, "eval", "bias detection", "agent regression", "sycophancy test", "no-bias evaluation", "harness".

## Facts
- Page: https://tashan.sh/capability/skill-shode666-eval-harness
- tashan id: skill:shode666/eval-harness
- Source: https://github.com/shode666/claude-skill-shode-house
- Type: skill
- Category: other
- tashan score: 45.0 / 100
- Adoption: 14.0
- Upkeep: 80.0
- Freshness: 95.0
- Evidence coverage: 84% of the inputs this score can use
- Health: active
- Instruction depth: not yet graded
- Official: no

## Install

```sh
cp -r eval-harness ~/.claude/skills/
```

## Security audit
Not scanned. We audit npm-published capabilities; this one has no npm package we can resolve, or has not reached the queue. This is not a clean bill of health.

---
Measured 2026-08-14 by tashan (https://tashan.sh) from public evidence. Scorer s5.
