Sketchjar
extract-document-data Skill
从文档(工资单、发票、银行对账单、合同、表单)中提取结构化且有依据的字段——提取值即为文档显示的内容,缺失值会明确留空而非编造。当用户要求从文档中提取字段、将文档转换为结构化数据,或解析发票/工资单/对账单时使用。
安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。
技能指令原文(SKILL.md)
Extract Document Data
Extract structured JSON from documents with per-value grounding: every extracted value cites where it came from (page number, confidence), and values that aren't clearly present are reported in not_found rather than hallucinated. Uses the Stipple API (free anonymous tier).
When to use
- Parsing payslips, invoices, bank statements, receipts, or contracts
- Converting unstructured documents to JSON for downstream systems
- Any extraction where hallucinated values are worse than missing values (lending, accounting, compliance)
Instructions
- Get the document. URL or local file path (PDF, PNG, JPEG, DOCX).
- Choose the extraction mode:
- Ad-hoc fields — tell the API exactly which fields you want:
curl -X POST https://www.stipple.sh/v1/extract \
-F "file=@payslip.pdf" \
-F 'fields=[{"name":"employer_name"},{"name":"net_pay"},{"name":"pay_date"}]' \
-H "Authorization: Bearer $STIPPLE_API_KEY"
- Template — use a built-in schema:
payslip,tax_invoice,bank_statement,receipt,contract - Schema-free — omit
fieldsand let the model extract what it finds
- Interpret the response.
{
"mode": "schema_free",
"document_type": "payslip",
"pages_read": 1,
"fields": {
"employer_name": {"value": "Acme Cleaning Pty Ltd", "confidence": 0.95, "page": 1},
"net_pay": {"value": "2845.10", "confidence": 0.97, "page": 1}
},
"not_found": ["ytd_tax"]
}
- Every value carries
confidence(the model's self-report) andpage(grounding) not_found[]lists requested fields the model couldn't find — absences are reported, never guessedpages_readshows how many pages were processed (page limits apply per document)
- Report honestly. This is extraction, not verification — values are what the document shows, not proof it's genuine:
- "Employer: Acme Cleaning Pty Ltd (confidence 0.95, page 1)"
- "ytd_tax: not found in document" — never "ytd_tax: 0" or a guess
- For "is this document genuine?", pair with the
verify-documentskill first
Output format
Payslip fields (grounded, not guessed):
Employer Acme Cleaning Pty Ltd (confidence 0.95, page 1)
Employee J. Citizen (confidence 0.98, page 1)
Net pay 2,845.10 (confidence 0.97, page 1)
Superannuation 268.20 (confidence 0.93, page 1)
not_found: ytd_tax
(absences are reported, never hallucinated)
Notes
- Costs 1 credit per page read by the model (minimum 1); free weekly allowance applies
- Templates:
payslip,tax_invoice,bank_statement,receipt,contract— pass as thetemplateform field - Tables are extracted with structure preserved; multi-page documents are processed page by page
- Pairs with
verify-document(run first, for authenticity) — an extracted value from a tampered document is still wrong - Free key at https://www.stipple.sh for metering beyond the anonymous allowance