The Index · Model evaluation

Model evaluation

'Evals' — scoring model output and catching regressions. tashan measures 28 capabilities for this work. Ranked on upkeep, freshness and real adoption, every input public. How we measure ›

#CapabilitytashanEvidenceHealth
1
Deepeval primary plugin solid
77
17k ★active
2
Datarobot Agent Skills primary plugin solid
62
23 ★active
3
Evalview primary plugin solid
62
124 ★active
4
Fiftyone primary plugin solid
61
37 ★active
5
Mlflow primary plugin solid
61
61 ★active
6
Probabl Skills primary plugin solid
57
74 ★active
7
Iris primary plugin solid
45
8 ★active
8
Nnsight primary plugin solid
42
9 ★active
9
Autoresearch AI Plugin primary plugin solid
40
12 ★active
10
HuggingFace Skills primary plugin
77
11k ★active
11
Promptfoo Evals primary plugin
76
24k ★active
12
48
5 ★active
13
Bitfab primary plugin
39
1 ★active
14
Langsmith primary npm
39
3k/wkabandoned
15
Claude Performance primary plugin
38
1 ★active
16
Everdict primary plugin
34
1 ★active
·
Setup primary skill
not scored yet
11 reposactive
·
Skill Creator primary skill ✓ Anthropic
not scored yet
7 reposactive
19
Langfuse primary plugin thin
70
218 ★active
20
53
16 ★active
21
Auxiliar supporting npm solid
54
158/wkactive
22
My Pi supporting npm
62
216/wkactive
23
Evals supporting npm
53
224/wkactive
24
Mcpscope supporting npm
52
300/wkactive
25
Trustmodel supporting npm
48
59/wkactive
26
Orizu supporting plugin
40
0 ★active
27
Prove supporting plugin
31
1 ★active
28
Plzebo supporting npm thin
51
326/wkactive

Who does this work

O*NET has no process step for this yet — its software occupations were surveyed before agentic tooling existed. We list it because the corpus plainly shows people doing it, and we say so rather than forcing it onto an unrelated step.

Recently changed in Model evaluation

14 changes recorded here in the last 45 days, newest and most serious first.

This page cannot know what you run. tashan doctor reads your own config and names which of these you have — Pro adds the history behind each, and what to move to.
tashan Pro — $6/mo ›

Other work

Agent configurationAgent developmentAnimation and generated imageryApplication developmentAttributionAudio productionAuditBookkeepingBrowser automationCAD modellingCareer developmentCode reviewContent marketingContract reviewCopy editingCopywritingCustomer supportDashboards and reportingData pipelinesData qualityDatabase accessDebuggingDesign critiqueDocument productionExploratory data analysisFinancial modellingHardware designIncident responseInfrastructure and deploymentKnowledge managementLiterature reviewLogistics and fulfilmentML engineeringMarket analysisMessaging and emailObservabilityPRDs and specsPerformance optimisationProcess automationProduct strategyProject managementPrompt engineeringQuery optimisationRecruiting and hiringRegulatory complianceRetrieval systemsRisk assessmentSEOSales pipelineScientific researchSecurity reviewSoftware architectureStatistical modellingTechnical documentationTest automationTraining materialUser researchVersion controlVideo editingVisual designWeb developmentWeb researchWeb scraping

Occupational data from the O*NET 30.3 Database by the U.S. Department of Labor, Employment and Training Administration, used under CC BY 4.0. tashan consolidated its process steps into the terms practitioners use; O*NET does not endorse this site.

See the full Index ›

tashan Pro$6/mo

14 changes across Model evaluation in the last 45 days. A new advisory, an install script appearing, a maintainer leaving. All listed free above. Pro keeps the series behind each row, so a number today comes with a direction.

Start a 7-day trial › Everything measured on this page stays free.