Skip to main content

document-extraction skills

1 skills. Source of truth: skills/document-extraction/.

kreuzberg​

Extract text, tables, metadata, and images from 91+ document formats (PDF, Office, images, HTML, email, archives, academic) using Kreuzberg. Use when writing code that calls Kreuzberg APIs in Python, Node.js/TypeScript, Rust, or CLI. Covers installation, extraction (sync/async), configuration (OCR, chunking, output format), batch processing, error handling, and plugins.

v1.0.0 · document-extraction pdf ocr text-extraction · source