apyhub
Back
▣ DATA EXTRACTION · FILE MANIPULATION

Extract Text from PDF API

Hosted on ApyHub

What it does

The PDF to Text API takes a PDF and returns its text. Send a remote url or upload a file of up to 100 MB, and the response comes back as a single data string in JSON.

Two options control what you get. start_page and end_page limit extraction to a page range, which keeps responses small when you need one section of a long report. The four coordinate parameters (starting_x_coordinate through ending_y_coordinate) accept values from 0 to 100 and restrict extraction to a region of each page, useful for pulling a header block or a single column. Set preserve_paragraphs to true to keep paragraph breaks in the output.

Teams use the PDF to Text API to extract text from PDF documents for search indexing, to feed contracts and reports into AI pipelines, to migrate archived content into a CMS, and to run compliance checks across document sets.

This endpoint reads the text layer a PDF already carries. For scanned pages and photographed documents, run the OCR Document Data Extraction API first, then send the result here.

It pairs with the Extract Text from Word API when a batch mixes DOC and DOCX files, and with the Table Extraction API when row and column structure matters. All three are callable by AI agents through ApyHub MCP.

▣ ENDPOINT 01 / 02
POST
submit url: extracted data
https://api.eu.apyhub.com/apyhub/extract-text-from-pdf/url

QUICKSTART

GUIDE

Quickstart

Extract text from a PDF by sending its URL in a minimal JSON request.

curl -X POST "https://api.eu.apyhub.com/apyhub/extract-text-from-pdf/url" \
  -H "apy-token: $APY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://assets.apyhub.com/samples/sample.pdf"}'

What you'll get back

Returns a JSON object with a data string field containing the extracted text from the PDF.

{
  "data": "Chapter 1. Sample PDF text content extracted from the document."
}
TRY ITLIVE · 50 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.
body*

About this endpoint

What it does

Extracts text from a PDF located at a remote URL and returns the extracted content as a string. You can optionally limit the extraction to a page range and define coordinate bounds for the extracted area.

Request Body

ParameterTypeDescription
urlStringRemote PDF URL. Must be a valid URI.
end_pageIntegerLast page to extract from. Default: 0. Minimum: 0.
start_pageIntegerFirst page to extract from. Default: 1. Minimum: 1.
ending_x_coordinateIntegerEnding X coordinate for text extraction bounds. Default: 0. Range: 0 to 100.
ending_y_coordinateIntegerEnding Y coordinate for text extraction bounds. Default: 0. Range: 0 to 100.
preserve_paragraphsBooleanPreserves paragraph breaks in the extracted text. Default: false.
starting_x_coordinateIntegerStarting X coordinate for text extraction bounds. Default: 0. Range: 0 to 100.
starting_y_coordinateIntegerStarting Y coordinate for text extraction bounds. Default: 0. Range: 0 to 100.

Response

Returns a JSON object with a data string field containing the extracted text. Success responses are represented as a JSON object shaped like { data: string }.

ParameterTypeDescription
dataStringExtracted text content from the PDF.
▣ ENDPOINT 02 / 02
POST
upload file: extracted data
https://api.eu.apyhub.com/apyhub/extract-text-from-pdf/file

QUICKSTART

GUIDE

Quickstart

Upload a PDF file to extract its text.

curl -X POST "https://api.eu.apyhub.com/apyhub/extract-text-from-pdf/file" \
  -H "apy-token: $APY_TOKEN" \
  -F "file=@/path/to/document.pdf"

What you'll get back

Returns a JSON object with a data string field containing the extracted PDF text.

{
  "data": "Chapter 1. Sample PDF text content extracted from the document."
}
TRY ITLIVE · 50 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.
Max 100MB total per request (all files combined). Larger? Use this API's URL-based endpoint instead, if it has one.
body*
PDF file (.pdf).
Last page (0 = all pages).

About this endpoint

What it does

Uploads a PDF file and extracts text from it. You can optionally limit extraction to a page range and a coordinate bounding box, and control whether paragraphs are preserved in the extracted text.

Request Body

ParameterTypeDescription
fileBinaryPDF file (.pdf). Binary upload.
end_pageIntegerLast page to extract. 0 means all pages. Default: 0.
start_pageIntegerFirst page to extract. Default: 1.
ending_x_coordinateIntegerEnding X coordinate for extraction bounds. Minimum 0, maximum 100. Default: 0.
ending_y_coordinateIntegerEnding Y coordinate for extraction bounds. Minimum 0, maximum 100. Default: 0.
preserve_paragraphsENUMWhether to preserve paragraphs in the extracted text. Allowed values: true, false. Default: false.
starting_x_coordinateIntegerStarting X coordinate for extraction bounds. Minimum 0, maximum 100. Default: 0.
starting_y_coordinateIntegerStarting Y coordinate for extraction bounds. Minimum 0, maximum 100. Default: 0.

Response

Returns a JSON object with a data string field containing the extracted text from the PDF.

ParameterTypeDescription
dataStringExtracted text content from the uploaded PDF file.

Max 100MB total per request (all files combined). Larger? Use this API's URL-based endpoint instead, if it has one.

▣ COMMON ERRORS

Errors any endpoint can return

400bad_request

Required parameter missing or malformed body.

401unauthorized

API key missing, revoked, or not authorized for this service.

429rate_limited

Your plan's per-second rate exceeded. Retry with exponential backoff.

503upstream_busy

Backend temporarily unavailable. Try again in a few seconds.