‹ The Index

Databricks Mlflow Evaluation

skill

MLflow 3 GenAI agent evaluation. Use when writing mlflow.genai.evaluate() code, creating @scorer functions, using built-in scorers (Guidelines, Correctness, Safety, RetrievalGroundedness), building eval datasets from traces, setting up trace ingestion and production monitoring, aligning judges with MemAlign from domain expert feedback, or running optimizeprompts() with GEPA for automated prompt improvement.

Works with: Claude Code (native)  ·  Cursor, Codex CLI (manual)
native: this artifact type is that client's own format

Category: Other — see all ranked ›

Install (Claude Code):

cp -r databricks-mlflow-evaluation ~/.claude/skills/

Security audit

Not scanned yet. We audit npm-published capabilities for known advisories, install-time scripts and permission surface; this one has no npm package we can resolve, or has not reached the queue.

source ↗  ·  skill:databricks/databricks-mlflow-evaluation

Already running this? npx tashan-cli doctor checks your whole config against the Index — how it works ›

Measured 2026-08-03  ·  scorer s5  ·  how  ·  something wrong here?