Measuring what a capability is actually worth

The AI ecosystem has reached its GitHub moment: millions of skills, thousands of MCP servers, endless prompts and agents.

Supply is infinite. What is scarce is trustworthy intelligence about which of them work, for whom, under what conditions, at what cost.

Every marketplace so far optimizes for the wrong thing: SEO, downloads, stars, listings. Anyone can build another directory tomorrow. Discovery has become a commodity. Intelligence has not.

Why this is not another directory

A directory’s job is to list. An instrument’s job is to be right. Those two conflict constantly — and every time they do, the result is on the page.

#1  Withheld

We delete rankings that mean nothing

722 skills published. None of them ranked. A skill’s only upkeep signal is its repository’s, shared by every skill in it.

One publisher had 263 skills and three distinct scores between them. A directory keeps those rows; ranked rows are the product.

#2  Absent

Unmeasured is never rendered as zero

1,015 rows carry no score and a written reason for it.

We record when a payment address was last checked, so “never paid” cannot be read as “we did not look”. A directory cannot afford an empty cell. We cannot afford a wrong one.

#3  Unflattering

Coverage falls when we do well

Every successful discovery lowers the measured-to-tracked ratio. We publish it and never target it.

No directory reports a number that punishes its own growth.

Everything above, you could recompute. That is the point. What you cannot recompute is what it said last month.

45 days of daily measurements, none of them backfillable. Someone starting tomorrow starts at zero and stays that far behind. What changed is built from it.

The tell: we publish an x402 price, so tashan sits on our own paid board at $0.00, never paid. Nobody exempted us.

A skill is not software. It is closer to a handbook, an SOP, a playbook. Expertise. Code has good measures. The expertise inside a capability has almost none. That is the opening.

So the rank here is not who shouts loudest. It is did it get adopted, was it kept, is it still maintained.

Confidence is an output of evidence, not a badge. Every input is named and checks back against the source it came from.

The name. 他山 (tashan) comes from the proverb 他山之石,可以攻玉a stone from another mountain can polish your jade.

Every tool in your AI’s toolbox was made on somebody else’s mountain. None of it is ours. We host nothing and sell nothing in it. The old line is about that kind of borrowing, and about the part people forget: not every stone is worth carrying home. Some polish the jade. Some scratch it.

Telling those apart, with evidence anyone can check, is the whole job.

What we're building

Narrow and honest first. Track the whole MCP field and score it on public signal: adoption, upkeep, freshness, and an LLM read of how well it documents itself, deep, solid or thin. Live on the Index today.

Then the things nobody can read off a listing: retention and churn from git history, controlled evals we run ourselves, compatibility across models, real cost. Capability by capability, the evidence base compounds into something no listing can copy.

Where this goes

The measured layer for the whole AI ecosystem. Skills, MCP servers, prompts, agents, models, workflows.

A verdict is only useful for as long as the work behind it holds up, so the plan is to keep widening what we can measure rather than what we can assert.

Principles

Mission. Help people and AI systems confidently choose the best capability for the job — by measuring what actually works, not just what exists.

See the Index ›