‹ The Index

Eval

skill

Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.

Works with: Claude Code, Cursor, Codex CLI

Category: Dev Tools & CI — see all ranked ›

Work: Model evaluation

Who it is for: AI engineer

Security audit

Not scanned yet. We audit npm-published capabilities for known advisories, install-time scripts and permission surface; this one has no npm package we can resolve, or has not reached the queue.

source ↗  ·  skill:alirezarezvani/eval

Already running this? npx tashan-cli doctor checks your whole config against the Index — how it works ›