# Text Corpus Analysis

> Skills for analyzing large text corpora — topic modeling (BERTopic with temporal evolution), NER, categorization into fixed taxonomies, bottom-up category derivation, multi-level taxonomy design, word frequency, synonym clustering for voice-note/STT corpora, parametric stats, and metadata↔content correlation. Three execution lanes (classical NLP, local LLM via Ollama, cloud LLM via OpenRouter) with explicit cost-awareness: mandatory pre-run estimates for 1k-doc LLM passes, two-pass cheap→premiu

## Facts
- Page: https://tashan.sh/capability/plugin-danielrosehill-claude-code-plugins-text-corpus-analysis
- tashan id: plugin:danielrosehill/claude-code-plugins/text-corpus-analysis
- Source: https://github.com/danielrosehill/Claude-Code-Plugins
- Type: plugin
- Category: docs
- tashan score: 41.0 / 100
- Adoption: 7.0
- Upkeep: 82.0
- Freshness: 100.0
- Evidence coverage: 84% of the inputs this score can use
- Health: active
- Instruction depth: not yet graded
- Official: no

## Install

```sh
/plugin marketplace add danielrosehill/Claude-Code-Plugins
/plugin install text-corpus-analysis@danielrosehill
```

## Security audit
Not scanned. We audit npm-published capabilities; this one has no npm package we can resolve, or has not reached the queue. This is not a clean bill of health.

---
Measured 2026-08-05 by tashan (https://tashan.sh) from public evidence. Scorer s5.
