Evalview
Behavior regression testing for AI agents. EvalView detects when your agent's behavior drifts e.g changed tool calls, different outputs, or degraded quality even when traditional tests still pass. Snapshot baselines, check for regressions, and auto-heal flaky results, all from inside Claude Code.
Works with: Claude Code, Cursor, Claude Desktop, Codex CLI, Gemini CLI, Cline, Windsurf, VS Code
Category: Dev Tools & CI — see all ranked ›
Work: Model evaluation · Test automation
Who it is for: AI engineer · Software engineer
“Prints real check output with three verdicts; says it is not an observability tool”
- tashan score: 69.0
- Expertise: 79.0 (solid)
- Adoption: 1 repos
- Health: active
- GitHub stars: 124
- Contributors: 15
- License: Apache-2.0
Security audit
Not scanned yet. We audit npm-published capabilities for known advisories, install-time scripts and permission surface; this one has no npm package we can resolve, or has not reached the queue.
source ↗ · plugin:hidai25/eval-view/evalview
Already running this? npx tashan-cli doctor checks your whole config against the Index — how it works ›