Sketchjar

extract-document-data Skill

从文档(工资单、发票、银行对账单、合同、表单)中提取结构化且有依据的字段——提取值即为文档显示的内容,缺失值会明确留空而非编造。当用户要求从文档中提取字段、将文档转换为结构化数据,或解析发票/工资单/对账单时使用。

安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。

查看源码

技能指令原文(SKILL.md)

Extract Document Data

Extract structured JSON from documents with per-value grounding: every extracted value cites where it came from (page number, confidence), and values that aren't clearly present are reported in not_found rather than hallucinated. Uses the Stipple API (free anonymous tier).

When to use

  • Parsing payslips, invoices, bank statements, receipts, or contracts
  • Converting unstructured documents to JSON for downstream systems
  • Any extraction where hallucinated values are worse than missing values (lending, accounting, compliance)

Instructions

  1. Get the document. URL or local file path (PDF, PNG, JPEG, DOCX).
  1. Choose the extraction mode:
  • Ad-hoc fields — tell the API exactly which fields you want:
     curl -X POST https://www.stipple.sh/v1/extract \
       -F "file=@payslip.pdf" \
       -F 'fields=[{"name":"employer_name"},{"name":"net_pay"},{"name":"pay_date"}]' \
       -H "Authorization: Bearer $STIPPLE_API_KEY"
  • Template — use a built-in schema: payslip, tax_invoice, bank_statement, receipt, contract
  • Schema-free — omit fields and let the model extract what it finds
  1. Interpret the response.
   {
     "mode": "schema_free",
     "document_type": "payslip",
     "pages_read": 1,
     "fields": {
       "employer_name": {"value": "Acme Cleaning Pty Ltd", "confidence": 0.95, "page": 1},
       "net_pay": {"value": "2845.10", "confidence": 0.97, "page": 1}
     },
     "not_found": ["ytd_tax"]
   }
  • Every value carries confidence (the model's self-report) and page (grounding)
  • not_found[] lists requested fields the model couldn't find — absences are reported, never guessed
  • pages_read shows how many pages were processed (page limits apply per document)
  1. Report honestly. This is extraction, not verification — values are what the document shows, not proof it's genuine:
  • "Employer: Acme Cleaning Pty Ltd (confidence 0.95, page 1)"
  • "ytd_tax: not found in document" — never "ytd_tax: 0" or a guess
  • For "is this document genuine?", pair with the verify-document skill first

Output format

Payslip fields (grounded, not guessed):

  Employer          Acme Cleaning Pty Ltd  (confidence 0.95, page 1)
  Employee          J. Citizen             (confidence 0.98, page 1)
  Net pay           2,845.10               (confidence 0.97, page 1)
  Superannuation    268.20                 (confidence 0.93, page 1)

not_found: ytd_tax
(absences are reported, never hallucinated)

Notes

  • Costs 1 credit per page read by the model (minimum 1); free weekly allowance applies
  • Templates: payslip, tax_invoice, bank_statement, receipt, contract — pass as the template form field
  • Tables are extracted with structure preserved; multi-page documents are processed page by page
  • Pairs with verify-document (run first, for authenticity) — an extracted value from a tampered document is still wrong
  • Free key at https://www.stipple.sh for metering beyond the anonymous allowance