# What to use for model evaluation

> 'Evals' — scoring model output and catching regressions.

Source: https://tashan.sh/task/model-evaluation.html
Ranked by fit for the task, then how well it documents itself, then the tashan score
  (upkeep and freshness, gated by real adoption). Public evidence only — nothing paid can
  change a rank. Method: https://tashan.sh/methodology.html

## Ranked

| # | Capability | tashan score | Adoption evidence | Activity |
|---|---|---|---|---|
| 1 | [Deepeval](https://tashan.sh/capability/plugin-confident-ai-deepeval-deepeval.html) | 77 | 17k ★ | active |
| 2 | [Datarobot Agent Skills](https://tashan.sh/capability/plugin-datarobot-oss-datarobot-agent-skills-datarobot-agent-skills.html) | 62 | 23 ★ | active |
| 3 | [Evalview](https://tashan.sh/capability/plugin-hidai25-eval-view-evalview.html) | 62 | 124 ★ | active |
| 4 | [Fiftyone](https://tashan.sh/capability/plugin-voxel51-fiftyone-skills-fiftyone.html) | 61 | 37 ★ | active |
| 5 | [Mlflow](https://tashan.sh/capability/plugin-mlflow-skills-mlflow.html) | 61 | 61 ★ | active |
| 6 | [Probabl Skills](https://tashan.sh/capability/plugin-probabl-ai-skills-probabl-skills.html) | 57 | 74 ★ | active |
| 7 | [Iris](https://tashan.sh/capability/plugin-iris-eval-mcp-server-iris.html) | 45 | 8 ★ | active |
| 8 | [Nnsight](https://tashan.sh/capability/plugin-ndif-team-skills-nnsight.html) | 42 | 9 ★ | active |
| 9 | [Autoresearch AI Plugin](https://tashan.sh/capability/plugin-proyecto26-autoresearch-ai-plugin-autoresearch-ai-plugin.html) | 40 | 12 ★ | active |
| 10 | [HuggingFace Skills](https://tashan.sh/capability/plugin-huggingface-skills-huggingface-skills.html) | 77 | 11k ★ | active |
| 11 | [Promptfoo Evals](https://tashan.sh/capability/plugin-promptfoo-promptfoo-promptfoo-evals.html) | 76 | 24k ★ | active |
| 12 | [Agent Eval Harness · redhat-global-engineering](https://tashan.sh/capability/plugin-redhat-global-engineering-ge-public-skills-agent-eval-harness.html) | 48 | 5 ★ | active |
| 13 | [Bitfab](https://tashan.sh/capability/plugin-project-white-rabbit-bitfab-claude-plugin-bitfab.html) | 39 | 1 ★ | active |
| 14 | [Langsmith](https://tashan.sh/capability/pkg-langsmith-mcp-server.html) | 39 | 3k/wk | abandoned |
| 15 | [Claude Performance](https://tashan.sh/capability/plugin-adelaidasofia-claude-performance-claude-performance.html) | 38 | 1 ★ | active |
| 16 | [Everdict](https://tashan.sh/capability/plugin-everdict-everdict-everdict.html) | 34 | 1 ★ | active |
| 17 | [Setup](https://tashan.sh/capability/skill-alirezarezvani-setup.html) | not scored | 11 repos | active |
| 18 | [Skill Creator](https://tashan.sh/capability/skill-anthropics-skill-creator.html) | not scored | 7 repos | active |
| 19 | [Langfuse](https://tashan.sh/capability/plugin-langfuse-skills-langfuse.html) | 70 | 218 ★ | active |
| 20 | [Nexus Agents · williamzujkowski](https://tashan.sh/capability/plugin-williamzujkowski-nexus-agents-nexus-agents.html) | 53 | 16 ★ | active |
| 21 | [Auxiliar](https://tashan.sh/capability/pkg-auxiliar-mcp.html) | 54 | 158/wk | active |
| 22 | [My Pi](https://tashan.sh/capability/pkg-my-pi.html) | 62 | 216/wk | active |
| 23 | [Evals](https://tashan.sh/capability/pkg-cyanheads-evals-mcp-server.html) | 53 | 224/wk | active |
| 24 | [Mcpscope](https://tashan.sh/capability/pkg-mcpscope.html) | 52 | 300/wk | active |
| 25 | [Trustmodel](https://tashan.sh/capability/pkg-trustmodel-mcp-server.html) | 48 | 59/wk | active |
| 26 | [Orizu](https://tashan.sh/capability/plugin-orizuai-orizu-cli-orizu.html) | 40 | 0 ★ | active |
| 27 | [Prove](https://tashan.sh/capability/plugin-vassilissoum-prove-prove.html) | 31 | 1 ★ | active |
| 28 | [Plzebo](https://tashan.sh/capability/pkg-plzebo.html) | 51 | 326/wk | active |

## What these numbers are not

- The tashan score measures upkeep, freshness and adoption. It is **not** a security
  verdict and **not** a measure of whether the capability works well.
- `not scored` means too little public evidence to rank, never that something is bad.
- The security audit is separate and free per capability, on each page above.
