Measuring what a capability is actually worth
The AI ecosystem has reached its GitHub moment: millions of skills, thousands of MCP servers, endless prompts and agents.
Supply is infinite. What is scarce is trustworthy intelligence about which of them work, for whom, under what conditions, at what cost.
Every marketplace so far optimizes for the wrong thing: SEO, downloads, stars, listings. Anyone can build another directory tomorrow. Discovery has become a commodity. Intelligence has not.
Why this is not another directory
A directory’s job is to list. An instrument’s job is to be right. Those two conflict constantly — and every time they do, the result is on the page.
#1 Withheld
We delete rankings that mean nothing
722 skills published. None of them ranked. A skill’s only upkeep signal is its repository’s, shared by every skill in it.
One publisher had 263 skills and three distinct scores between them. A directory keeps those rows; ranked rows are the product.
#2 Absent
Unmeasured is never rendered as zero
1,015 rows carry no score and a written reason for it.
We record when a payment address was last checked, so “never paid” cannot be read as “we did not look”. A directory cannot afford an empty cell. We cannot afford a wrong one.
#3 Unflattering
Coverage falls when we do well
Every successful discovery lowers the measured-to-tracked ratio. We publish it and never target it.
No directory reports a number that punishes its own growth.
Everything above, you could recompute. That is the point. What you cannot recompute is what it said last month.
45 days of daily measurements, none of them backfillable. Someone starting tomorrow starts at zero and stays that far behind. What changed is built from it.
The tell: we publish an x402 price, so tashan sits on our own paid board at $0.00, never paid. Nobody exempted us.
A skill is not software. It is closer to a handbook, an SOP, a playbook. Expertise. Code has good measures. The expertise inside a capability has almost none. That is the opening.
So the rank here is not who shouts loudest. It is did it get adopted, was it kept, is it still maintained.
Confidence is an output of evidence, not a badge. Every input is named and checks back against the source it came from.
The name. 他山 (tashan) comes from the proverb 他山之石,可以攻玉 — a stone from another mountain can polish your jade.
Every tool in your AI’s toolbox was made on somebody else’s mountain. None of it is ours. We host nothing and sell nothing in it. The old line is about that kind of borrowing, and about the part people forget: not every stone is worth carrying home. Some polish the jade. Some scratch it.
Telling those apart, with evidence anyone can check, is the whole job.
What we're building
Narrow and honest first. Track the whole MCP field and score it on public signal: adoption, upkeep, freshness, and an LLM read of how well it documents itself, deep, solid or thin. Live on the Index today.
Then the things nobody can read off a listing: retention and churn from git history, controlled evals we run ourselves, compatibility across models, real cost. Capability by capability, the evidence base compounds into something no listing can copy.
Where this goes
The measured layer for the whole AI ecosystem. Skills, MCP servers, prompts, agents, models, workflows.
A verdict is only useful for as long as the work behind it holds up, so the plan is to keep widening what we can measure rather than what we can assert.
Principles
- Derived from public evidence. If we cannot derive it from a public source, we do not publish it.
- Transparent by default. Every signal is defined and traceable. See the Methodology.
- Independent judgement. Commercial relationships never influence task fit, measurements, findings, rankings or editorial recommendations.
- Humble about the unknown. We name what we cannot yet measure. That is the roadmap, not a secret.