About this endpoint
What it does
Extracts article content and metadata from a webpage URL. The request sends a webpage url, and the response returns extracted article fields such as body, title, author, images, and related metadata.
Request Body
| Parameter | Type | Description |
|---|---|---|
| url | String | URL of the webpage to extract, audit, or validate content from. Must be an http or https URI. |
Response
Returns a JSON object with body, date, title, author, images, language, confidence, source_url, and word_count fields.
| Parameter | Type | Description |
|---|---|---|
| body | String | Extracted article body text. |
| date | String | Article date in date format. |
| title | String | Extracted article title. |
| author | String | Extracted article author. |
| images | String Array | Array of image URLs (uri format). |
| language | String | Detected language of the article. |
| confidence | Object | Confidence scores for extracted content. Contains body, title, author and date numeric fields. |
| confidence.body | Number | Confidence score for the extracted body. |
| confidence.title | Number | Confidence score for the extracted title. |
| confidence.author | Number | Confidence score for the extracted author. |
| confidence.date | Number | Confidence score for the extracted article date. |
| source_url | String | Source URL of the extracted article. uri format. |
| word_count | Integer | Word count of the extracted article. |



