apyhub
Back
▣ DATA EXTRACTION

Extract Article Content from Web Page API

What it does

The Article Extractor API turns a web article into clean text and metadata. Send one page as url to the GET endpoint, or an array of urls to the POST batch endpoint, and get back the readable body in articleBody with navigation, ads, and page chrome stripped out. Each result also includes title, author, publicationName, publicationDate, platform, wordCount, and readingTime, plus coverImage on platforms such as Medium.

Batch responses return a data array with one record per URL and a count field. articleBody is plain text with paragraphs separated by line breaks, ready for search, embeddings, or a language model. It works best on article pages such as blog posts, news stories, and Medium posts. Dates follow the source page's format, so normalize publicationDate before you sort on it.

Use the article extractor API to feed news and blog content into a RAG pipeline, import posts into a CMS, build a reader view or read-later app, or collect articles for research and content monitoring. AI agents can call it through ApyHub MCP to read a page before answering a question about it.

To condense the text, send it to the AI Summarizer API. For other languages, the Article Translation API extracts and translates in one call, and the Article Analysis API scores a page for SEO and content quality.

▣ ENDPOINT 01 / 02
GET
Extract article content from a web page
https://api.eu.apyhub.com/namastesumalya/extract-article-content

QUICKSTART

GUIDE

Quickstart

Fetch article metadata and extracted content by passing the article URL as a query parameter.

curl -X GET "https://api.eu.apyhub.com/namastesumalya/extract-article-content?url=https://example.com/article" \
  -H "apy-token: $APY_TOKEN"

What you'll get back

Returns a JSON object with article details in top-level fields such as url, title, author, platform, wordCount, articleBody, readingTime, publicationDate, and publicationName.

{
  "url": "https://example.com/article",
  "title": "Understanding Modern Content Pipelines",
  "author": "John Doe",
  "platform": "WordPress",
  "wordCount": 850,
  "articleBody": "This article explains how modern content pipelines work, including extraction, transformation, and delivery processes...",
  "readingTime": "4 min",
  "publicationDate": "2026-02-10T00:00:00Z",
  "publicationName": "Example Blog"
}
TRY ITLIVE · 300 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.

About this endpoint

What it does

Extracts article content from the web page at the given url and returns structured article metadata and text content.

Query Parameter(s)

AttributeTypeDescription
urlStringThe web page URL to extract article content from.

Response

Returns a JSON object with article fields including url, title, author, platform, wordCount, articleBody, readingTime, publicationDate, and publicationName.

ParameterTypeDescription
urlStringThe source page URL.
titleStringThe article title.
authorStringThe article author.
platformStringThe platform detected for the page.
wordCountIntegerThe article word count.
articleBodyStringThe extracted article body text.
readingTimeStringThe estimated reading time.
publicationDateStringThe publication date/time.
publicationNameStringThe publication name.
▣ ENDPOINT 02 / 02
POST
Batch extract article content from multiple web pages
https://api.eu.apyhub.com/namastesumalya/extract-article-content

QUICKSTART

GUIDE

Quickstart

Send one or more article URLs to extract their content.

curl -X POST "https://api.eu.apyhub.com/namastesumalya/extract-article-content" \
  -H "apy-token: $APY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://example.com/article1","https://example.com/article2"]}'

What you'll get back

Returns a JSON object with a data array of extracted article objects and a count integer for how many articles were processed.

{
  "data": [
    {
      "url": "https://example.com/article",
      "title": "Understanding Modern Content Pipelines",
      "author": "John Doe",
      "platform": "WordPress",
      "wordCount": 850,
      "articleBody": "This article explains how modern content pipelines work, including extraction, transformation, and delivery processes...",
      "readingTime": "4 min",
      "publicationDate": "2026-02-10T00:00:00Z",
      "publicationName": "Example Blog"
    }
  ],
  "count": 2
}
TRY ITLIVE · 500 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.
body*
urls

About this endpoint

What it does

Extracts article content from multiple web page URLs sent in the request body and returns an array of extracted article records, along with a total count.

Request Body

ParameterTypeDescription
urlsString ArrayA list of web page URLs to process.

Response

Returns a JSON object with a data array and a count integer field. Each item in data is an object containing the extracted article details for one URL, and count is the total number of returned items.

ParameterTypeDescription
dataObject ArrayAn array of article objects. Each object may include:
- url (String): The source page URL.
- title (String): The article title.
- author (String): The article author.
- platform (String): The publishing platform.
- wordCount (Integer): The article word count.
- articleBody (String): The extracted article text.
- readingTime (String): The estimated reading time.
- publicationDate (String): The publication date/time.
- publicationName (String): The publication name.
countIntegerThe total number of items in data.
▣ COMMON ERRORS

Errors any endpoint can return

400bad_request

Required parameter missing or malformed body.

401unauthorized

API key missing, revoked, or not authorized for this service.

429rate_limited

Your plan's per-second rate exceeded. Retry with exponential backoff.

503upstream_busy

Backend temporarily unavailable. Try again in a few seconds.