‹ The Index

LLM Evals Engineer

skill

Author’s activation text, quoted as published

- Evaluate and production-harden LLM applications end to end. Use when the user says "evaluate my LLM app", "add evals", "is my RAG hallucinating", "RAG quality metrics", "LLM-as-judge", "add tracing for my chatbot", "guardrails", "prompt versioning", "rate-limit my LLM spend", "LLM budgets", or "make my LLM app production-ready". Covers tracing architecture (callback handler plus decorator instrumentation, trace-per-query spans, session/user propagation), the score data model, Ragas-style metrics (context precision/recall, context entities recall, noise sensitivity, faithfulness, response relevancy, tool call accuracy), LLM-as-judge pipelines, prompt versioning, chat-history trimming. …

Works with: Claude Code (native)  ·  Cursor, Codex CLI (manual)
native: this artifact type is that client's own format

Install (Claude Code):

cp -r llm-evals-engineer ~/.claude/skills/

Security audit

Not scanned yet. We audit npm-published capabilities for known advisories, install-time scripts and permission surface; this one has no npm package we can resolve, or has not reached the queue.

Its own instructions

Its SKILL.md says when to use it, shows worked examples and states a limitation.

Read from the capability’s own SKILL.md. This is not a grade and does not compare to the instruction-depth verdict on an MCP server — a skill has no tools to document, so that rubric does not apply to it.

You searched for one. Check the rest of your stack:

npx tashan-cli doctor

Reads the config already on your machine and names what is dead, deprecated or running code at install time. No account, nothing uploaded.

tashan Pro$6/mo

Pro adds the history to tashan doctor, so a run over your own config says which of yours gained an advisory, started running an install script, or lost its last maintainer — and what to move to.

Start a 7-day trial › Everything measured on this page stays free.

source ↗  ·  skill:Cloud-Byte-Consulting/llm-evals-engineer

Everything on this page is public evidence and free. What it cannot know is whether you run this — check your whole config, free, in the browser. tashan Pro adds the series behind each row and names a replacement for anything dying.

Already running this? Check your whole config — free, in your browser, nothing installed. Or npx tashan-cli doctor locally, which sends nothing at all.

Measured 2026-08-22  ·  scorer s5  ·  how  ·  something wrong here?