# LLM Evals Engineer

> - Evaluate and production-harden LLM applications end to end. Use when the user says "evaluate my LLM app", "add evals", "is my RAG hallucinating", "RAG quality metrics", "LLM-as-judge", "add tracing for my chatbot", "guardrails", "prompt versioning", "rate-limit my LLM spend", "LLM budgets", or "make my LLM app production-ready". Covers tracing architecture (callback handler plus decorator instrumentation, trace-per-query spans, session/user propagation), the score data model, Ragas-style metrics (context precision/recall, context entities recall, noise sensitivity, faithfulness, response relevancy, tool call accuracy), LLM-as-judge pipelines, prompt versioning, chat-history trimming, input guardrails, and gateway controls. Langfuse, Ragas, NeMo Guardrails, LiteLLM, and Redis are example capability fills. For trust auditing of autonomous agent actions and delegation gates, use agent-trust-auditor.

## Facts
- Page: https://tashan.sh/capability/skill-cloud-byte-consulting-llm-evals-engineer
- tashan id: skill:Cloud-Byte-Consulting/llm-evals-engineer
- Source: https://github.com/Cloud-Byte-Consulting/plugins
- Type: skill
- Category: other
- tashan score: not scored (catalogued only — too little public evidence)
- Adoption: 9.0
- Upkeep: 95.0
- Freshness: 90.0
- Evidence coverage: 84% of the inputs this score can use
- Health: active
- Instruction depth: not yet graded
- License: Apache-2.0
- Official: no

## Install

```sh
cp -r llm-evals-engineer ~/.claude/skills/
```

## Security audit
Not scanned. We audit npm-published capabilities; this one has no npm package we can resolve, or has not reached the queue. This is not a clean bill of health.

---
Measured 2026-08-22 by tashan (https://tashan.sh) from public evidence. Scorer s5.
