# Document Extraction API

> - Two primitives of the Nutrient Data Extraction API. parse (/extraction/parse) returns the whole-document model — a structural JSON of typed elements with bounding boxes, or whole-document Markdown — for RAG ingestion, search indexing, content migration, or layout-aware understanding. extract (/extraction/extract) returns just the fields you define in a JSON Schema, each with a per-field citation grounding it to a page region. Route to extract for "pull the invoice number and total", "extract these fields", "map to my schema", or "with citations"; route to parse for "parse this document", "whole-document Markdown", "chunk for embeddings", or "extract every table/element" (no target schema). Triggers include parse this document, extract layout, RAG pipeline, schema extraction, field extraction, cited fields, invoice/form field extraction, document understanding.

## Facts
- Page: https://tashan.sh/capability/skill-pspdfkit-labs-document-extraction-api
- tashan id: skill:PSPDFKit-labs/document-extraction-api
- Source: https://github.com/PSPDFKit-labs/nutrient-skills
- Type: skill
- Category: other
- tashan score: not scored (catalogued only — too little public evidence)
- Adoption: 9.0
- Upkeep: 99.0
- Freshness: 97.0
- Evidence coverage: 84% of the inputs this score can use
- Health: active
- Instruction depth: not yet graded
- Official: no

## Install

```sh
cp -r document-extraction-api ~/.claude/skills/
```

## Security audit
Not scanned. We audit npm-published capabilities; this one has no npm package we can resolve, or has not reached the queue. This is not a clean bill of health.

---
Measured 2026-09-12 by tashan (https://tashan.sh) from public evidence. Scorer s5.
