How to extract URLs from text and validate them with an API (2026 guide)
Updated September 2026
01Introduction
To extract URLs from text, send the text to a URL extractor API and get back a structured list of every link it contains. The URLs Detector API does this in one call: it returns one JSON object per link, with the full URL and its protocol. It finds links with and without a scheme (docs.example.com/v2), plus ftp:, mailto: and tel: links. It is built by SharpAPI and available through ApyHub.
Most teams start with a regex, or with the URL parser in their language. A regex can find link-shaped strings in a paragraph, but URL grammar is large, and the patterns that cover it tend to be long and fragile. OWASP lists badly constructed patterns as a denial-of-service risk (OWASP, Regular expression Denial of Service). A parser such as JavaScript's new URL() or Python's urllib.parse checks a string you already know is a URL, and cannot find URLs inside a sentence.
Extraction is also only step one of URL validation. Links decay: Pew Research found that 38% of webpages that existed in 2013 were no longer accessible a decade later, and 23% of news pages contained at least one broken link (Pew Research Center, May 2024). This guide shows how to extract links from text with one API call, then chain two more calls into a broken link checker that also flags malicious URLs.
02What's new in the 2026 update
- A live test of the API in September 2026, with the real request and response below.
- A corrected description of the output. The URLs Detector API extracts and normalizes links. It does not check whether they load, so this guide now covers that step with a separate API.
- A three-call broken link checker pipeline.
- Access for AI agents through ApyHub MCP.
03What is a valid URL?
A valid URL is a string that follows the standard URL structure: a scheme, a host, and optionally a port, path, query string and fragment. In https://apyhub.com:443/catalog?q=url#top, the scheme is https, the host is apyhub.com, the port is 443, the path is /catalog, the query is q=url and the fragment is top.
URL validation usually means one of three checks, and each needs a different tool:
- Is it well formed? A built-in URL parser answers this.
- Does it load? Only an HTTP request answers this.
- Is it safe? Only a lookup against a threat database answers this.
A URL can pass the first check and fail the other two. https://example.com/deleted-page is well formed and returns a 404.
04How to extract URLs from text with an API
Submit the text you want to scan. The API works as an asynchronous job, so the first call returns a job ID.
bash
curl -X POST "https://api.eu.apyhub.com/sharpapi/detect-urls" \
-H "apy-token: $APY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"content":"Docs are at docs.example.com/v2/setup and the repo is https://github.com/acme/widget. Old link http//broken-site.org does not work. Ping me at mailto:[email protected] or call tel:+302101234567. Also see www.apyhub.com/catalog, ftp://files.example.net/export.csv and bit.ly/3xYz12."}'The response contains a job_id and a status_url. Poll the status URL: while the job runs, status is new, and when it finishes, the result array appears. In our test it was ready on the second poll. This is the actual response we got for the text above:
json
{
"data": {
"type": "api_job_result",
"id": "645f0983-546d-48ba-bfa6-c17b0e48b679",
"attributes": {
"status": "success",
"type": "content_detect_urls",
"result": [
{ "url": "http://docs.example.com/v2/setup", "protocol": "http" },
{ "url": "https://github.com/acme/widget", "protocol": "https" },
{ "url": "http://broken-site.org", "protocol": "http" },
{ "url": "mailto:[email protected]", "protocol": "mailto" },
{ "url": "tel:+302101234567", "protocol": "tel" },
{ "url": "http://www.apyhub.com/catalog", "protocol": "http" },
{ "url": "ftp://files.example.net/export.csv", "protocol": "ftp" },
{ "url": "http://bit.ly/3xYz12", "protocol": "http" }
]
}
}
}Three things to notice in that output:
- Bare domains get a scheme.
docs.example.com/v2/setupcame back ashttp://docs.example.com/v2/setup. That makes every result parseable, and it also means the API assumeshttpwhen the text gave no scheme. - Malformed links get repaired.
http//broken-site.org(missing colon) came back ashttp://broken-site.org. That helps extraction. If you need to reject malformed input as written, compare the returned URL against the original text. - Non-web schemes are kept and labeled.
mailto:,tel:andftp:links come back with their protocol, so you can route or drop them with one filter.
05Common challenges in extracting and validating URLs
1. Finding URLs inside text
User comments, support tickets, chat logs and scraped pages mix links with punctuation, brackets and line breaks. Links often lack a scheme. A regex that catches example.com/path also tends to catch file.txt or v2.0, and each fix adds another branch to the pattern.
ApyHub's API solution: send the raw text in the content field. The URLs Detector API returns only the links, each normalized to a full URL with its protocol. Trailing punctuation from the sentence (widget.) was stripped in our test.
2. Checking for broken links
A well-formed URL can point to a page that no longer exists, or to a short link that redirects three times before landing somewhere unexpected. The URLs Detector API does not request the links it finds, so it does not tell you whether a link is live.
ApyHub's API solution: pass each extracted URL to the Resolve Short URL API. It follows redirects (up to 30 hops, configurable with max_hops) and returns the final resolved URL plus every hop with its HTTP status. A final hop with a 4xx or 5xx status is a broken link. A resolved domain that differs from the visible domain is worth a second look.
3. Blocking malicious links
Platforms that accept user content need to stop phishing and malware links before other users click them. Format checks cannot catch this, because a malicious URL is usually well formed.
ApyHub's API solution: call the Generate Link Preview API with secure_mode set to true (the default). The URL is checked against a malicious-URL database before the page is fetched. A flagged link returns reported_malicious: true and a threat label such as malware. A clean link returns the page title, description, images and favicon, which you can reuse to render a link card.
4. Enforcing domain policies
Regulated teams in finance, healthcare and education often keep allowlists or blocklists for outbound links. The policy logic belongs in your code. The hard part is getting a clean, consistent list of URLs to apply it to.
ApyHub's API solution: every result has a full URL and a protocol field, so the policy check is a few lines of code. Parse the host with your language's URL parser and compare it against your list. Drop tel: and mailto: entries, or send email addresses to the Detect Email Address API.
06Build a broken link checker with three API calls
A link-checking step for user content, a CMS publish hook or a scraping pipeline chains three calls:
- Send the text to the URLs Detector API and keep results where
protocolishttporhttps. - Pass each link to the Resolve Short URL API and mark it broken if the final hop returns 4xx or 5xx.
- Pass each link to the Generate Link Preview API with
secure_modeon and block anything withreported_malicious: true. - Compare the host of each
resolvedURL against your allowlist or blocklist.
To check links on a webpage instead of in text, replace step 1 with the Extract Links from Webpage API, which takes a page URL and returns every hyperlink as an absolute URL.
07URL extractor API, parser or online tool: which to use
- Online URL extractor and link extractor tools work well for a one-off job: paste text, copy the links. They do not fit into an automated workflow.
- A built-in parser (
new URL()in JavaScript,urllib.parse.urlparse()in Python) is the right tool when you have one string from a form field and need to confirm it is well formed. Check for anhttporhttpsscheme and a host. A parser does not find URLs inside a paragraph, and does not tell you whether the page exists. - A URL extractor API fits when the text comes from users, scrapers or documents you do not control, and you need the links as data inside your application or agent.
08About ApyHub's URLs Detector API
The URLs Detector API is built by SharpAPI, which describes it as AI-powered, and is available in the ApyHub catalog. It has two endpoints: POST /sharpapi/detect-urls submits a content string and returns a job_id and status_url, and the status endpoint returns the job state and, when finished, an array of { url, protocol } objects. It detects http, https, ftp, sftp, ftps, mailto and tel links, with or without a scheme in the source text.
Every ApyHub endpoint, including this one, is available through ApyHub MCP. AI agents can discover the URLs Detector API, read its schema and call it directly, with no hand-written wrapper or tool definition. An agent that summarizes support tickets or reviews pull request descriptions can pull every link out of the text and run the broken link and safety checks above in the same run.
Connect ApyHub MCP to your agent
09Conclusion
Extracting URLs from text and validating them are separate jobs. The URLs Detector API handles extraction and returns a clean, normalized list. The Resolve Short URL API tells you where each link lands and whether it responds. The Generate Link Preview API flags known malicious links. Chained together, they form a broken link checker you can add to a comment form, a CMS publish hook or a scraping pipeline, all on one ApyHub subscription.
10Frequently asked questions
How do I extract URLs from text?
Send the text to a URL extractor API such as the URLs Detector API in the content field, then poll the returned status_url. The finished job returns an array with one object per link, each holding the full url and its protocol.
What is a URL extractor?
A URL extractor, also called a link extractor, is a tool that scans text or a webpage and returns the links it contains as a list. Online extractors handle one-off pastes. A URL extractor API returns the links as JSON, so you can use them in code.
What is a valid URL?
A valid URL follows the standard structure of scheme, host and optional port, path, query and fragment, such as https://apyhub.com/catalog. A URL without a scheme, such as apyhub.com/catalog, is not valid on its own, though browsers and extraction APIs often add http:// for you.
How do I validate a URL?
Decide which check you need. For format, use a built-in parser such as new URL() in JavaScript or urllib.parse.urlparse() in Python. For whether it loads, send an HTTP request or use the Resolve Short URL API. For safety, use the Generate Link Preview API with secure_mode on.
How do I check if a URL is broken?
Request it and read the HTTP status. The Resolve Short URL API does this for you, following redirects and returning the status of every hop. A final status in the 4xx or 5xx range means the link is broken.
Is there a broken link checker API?
Yes. On ApyHub you can build one from three calls: the URLs Detector API or the Extract Links from Webpage API to collect links, the Resolve Short URL API to check status, and the Generate Link Preview API to flag malicious links.
Does the URLs Detector API check whether a link works?
No. It extracts and normalizes links from text and does not request them. Pair it with the Resolve Short URL API to check reachability.
Can it find URLs that have no http or https prefix?
Yes. Links such as docs.example.com/v2/setup or www.apyhub.com/catalog are detected and returned with http:// added. If the site is HTTPS-only, upgrade the scheme yourself or rely on the redirect check to find the final URL.
What types of links does it detect?
It returns web links (http, https), file transfer links (ftp, sftp, ftps), email links (mailto) and phone links (tel). Each result carries a protocol field, so you can filter by type in one line.
Will it flag a malformed URL?
It tends to repair it. In our test, http//broken-site.org came back as http://broken-site.org. If you need to reject malformed input as the user typed it, compare each returned URL against the original text.
How do I extract all links from a webpage?
Use the Extract Links from Webpage API. Send a page URL and it returns every hyperlink as an absolute URL, with an optional secure_mode check against a malicious-URL database.
Should I use regex to extract or validate URLs?
For a single form field, a built-in parser is safer and easier to maintain than a regex. Complex URL patterns can also be slow on crafted input, a risk OWASP documents as regular expression denial of service. For finding URLs inside free text, an extraction API saves you from maintaining the pattern.
How much does the URLs Detector API cost?
Each detection job costs 1,000 atoms, and each status poll costs 1 atom. Atoms are ApyHub's unit of usage, where each call's cost reflects the compute behind it. The Resolve Short URL API costs 100 atoms per call and the Generate Link Preview API costs 50.
How do I test the URLs Detector API?
You have three options. Use Try it on the API page to send sample text from your browser. Use Voiden, the free, open-source API client from the ApyHub team, to save the submit, poll, redirect and safety calls as plain Markdown files in your Git repo and rerun them locally. Or send the curl request shown above from any terminal.
Can AI agents use this API?
Yes. The API is available through ApyHub MCP, so an agent can discover it, read its schema and call it without a custom wrapper.
11About ApyHub
ApyHub is a curated API catalog and trusted operational layer for developers and AI agents. It offers over 1,500 endpoints or capabilities, and the catalog keeps growing. Every API runs on a single subscription billed in atoms, and each API carries machine-readable certification (GDPR, SOC 2, ISO 27001). Every endpoint is MCP-ready by default. ApyHub is headquartered in Amsterdam, with offices in the Netherlands, Greece and India, and serves 65,000+ monthly developer workspaces. The free tier needs no credit card. API providers can list their APIs at apyhub.com/become-a-provider.
