apyhub
Back
▣ DATA EXTRACTION · SEO

AI-ready Clean Data Extractor API

Hosted on ApyHub

What it does

The Clean Data Extractor is a web data extraction API that turns any web page into clean, structured data an LLM, search index, or pipeline can use. Send a page url in a POST body and get back a data object with title, author, published_date, summary, headings, sections, and the page content as clean_markdown and raw_markdown.

It pulls out the parts of a page that are hard to scrape: tables as headers and rows, links tagged as internal or external, images with alt text, and code_blocks. clean_markdown drops cookie banners and header navigation, while raw_markdown keeps the full page. A quality object returns a score from 0 to 1 with warnings for pages that only partly extracted, and metadata adds language, word count, and read time. Use the score to decide which pages to keep.

Use the Clean Data Extractor to feed web pages into RAG and embeddings, give AI agents readable page content, import docs and help centers into a knowledge base, or collect tables and links for research and SEO audits. As a website content extractor it returns one consistent JSON shape for every page, so your parser stays the same across sites. AI agents can call it through ApyHub MCP to read a page mid-task.

To find every page on a site first, run the Sitemap Extraction API. For article body text with author and date only, use the Article Extractor API, and for plain visible text, the Extract Text from Website API.

POST
Extract structured content from a webpage
https://api.eu.apyhub.com/apyhub/webpage-extractor-api

QUICKSTART

GUIDE

Quickstart

Fetch structured content from a webpage by sending its URL.

curl -X POST "https://api.eu.apyhub.com/apyhub/webpage-extractor-api" \
  -H "apy-token: $APY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://apyhub.com"}'

What you'll get back

Returns a JSON object with a data object containing the extracted page content. The data object can include fields like title, summary, links, images, tables, headings, sections, metadata, and page_type, along with other structured content from the page.

{
  "data": {
    "title": "Example page title",
    "summary": "Short page summary",
    "page_type": "article"
  }
}
TRY ITLIVE · 50 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.
body*
HTTP/HTTPS URL of the page to scrape

About this endpoint

What it does

Extracts structured content from a webpage given its URL and returns a JSON object containing the scraped page data under data.

Request Body

ParameterTypeDescription
urlStringHTTP/HTTPS URL of the page to scrape. Format: URI.

Response

Returns a JSON object with a required data object field containing structured page content. The data object may include page-level fields such as links, title, author, images, tables, quality, summary, category, headings, metadata, sections, page_type, code_blocks, story_cards, raw_markdown, clean_markdown, published_date, and content_markdown.

ParameterTypeDescription
dataObjectStructured page content extracted from the webpage.
▣ COMMON ERRORS

Errors any endpoint can return

400bad_request

Required parameter missing or malformed body.

401unauthorized

API key missing, revoked, or not authorized for this service.

429rate_limited

Your plan's per-second rate exceeded. Retry with exponential backoff.

503upstream_busy

Backend temporarily unavailable. Try again in a few seconds.