Parse a PDF to text and Markdown
Parse a PDF to text and Markdown: send a PDF url or base64 file, get clean Markdown with paragraphs rebuilt from the text layer, per-page character offsets for citing pages, metadata (title, author, dates, page count) and the page numbers that are scanned images with no text layer. Up to 50 pages per call. Text layer only, no OCR: an image-only PDF returns an uncharged 422.
Not tested: has real-world effectsour last check, 2026-10-11
0 of 0checks answered this week
n/amedian answer time
$0.005listed price per call
n/aprice it asked us
Paid test badge: not yet. The checks above are free: we call the tool without paying and read the payment request it sends back. The Verified badge needs paid calls whose answers match the promised output, and nobody can buy a badge.
Endpoint
POST https://netintel.dev/pdf/parse
| Category | Image and media |
|---|---|
| Provider host | netintel.dev |
| Networks | eip155:8453, solana:5eykt4UsFv8P8NJdTREpY1vzqKqZKvdp |
| Payment schemes | exact |
| Self-reported calls, 30 days | 2 from 2 payers (the provider's figure, not ours) |
Our checks, last 30 days
We never call tools that send, buy, move money or file anything, not even without paying.
Example input (from the provider)
{
"body": {
"url": "https://netintel.dev/samples/pdf-parse-sample.pdf"
},
"bodyType": "json",
"method": "POST",
"type": "http"
}
Promised output schema (from the provider)
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"properties": {
"input": {
"additionalProperties": false,
"properties": {
"body": {
"properties": {
"file_base64": {
"description": "The PDF as base64 or a data: URL (up to ~700 KB). Send url OR file_base64.",
"type": "string"
},
"max_chars": {
"description": "Cap on the markdown length (default 100000)",
"type": "number"
},
"pages": {
"description": "Pages to parse, 1-based, e.g. \"1-3,7\" (default: the first 50)",
"type": "string"
},
"url": {
"description": "Public URL of the PDF (up to 10 MB)",
"type": "string"
}
},
"type": "object"
},
"bodyType": {
"enum": [
"json",
"form-data",
"text"
],
"type": "string"
},
"method": {
"enum": [
"POST"
],
"type": "string"
},
"type": {
"const": "http",
"type": "string"
}
},
"required": [
"type",
"method",
"bodyType",
"body"
],
"type": "object"
},
"output": {
"properties": {
"example": {
"properties": {
"char_count": {
"type": "number"
},
"findings": {
"items": {
"type": "string"
},
"type": "array"
},
"markdown": {
"description": "All parsed pages, each after a <!-- page N --> marker",
"type": "string"
},
"metadata": {
"type": "object"
},
"page_count": {
"type": "number"
},
"pages": {
"description": "Per page: page, word_count, has_text, char_start, char_end (offsets into markdown)",
"type": "array"
},
"pages_parsed": {
"type": "number"
},
"scanned_pages": {
"items": {
"type": "number"
},
"type": "array"
},
"source": {
"type": "object"
},
"truncated": {
"type": "boolean"
},
"word_count": {
"type": "number"
}
},
"type": "object"
},
"type": {
"type": "string"
}
},
"required": [
"type"
],
"type": "object"
}
},
"required": [
"input"
],
"type": "object"
}