Skip to main content

Web Fetch

Hand it a URL, get the readable content back as clean markdown. Navigation, ads, cookie banners and scripts are stripped out โ€” you get the article, not the page furniture.

This is the "read one page" primitive. Pair it with Web Search to build research flows: search โ†’ fetch the best hits โ†’ feed the text to a model.

Endpointsโ€‹

GET/fetch?url=https://nextjs.org/blog/next-16
Clean markdown of the main content.
GET/fetch?url=...&format=text
Plain text โ€” markdown syntax stripped.
GET/fetch?url=...&format=html
The isolated main-content HTML.
GET/fetch?url=...&format=json
Metadata only โ€” title, description, og:image, canonical.
GET/crawl?url=...&limit=5
Fetch a page and follow its own links.
GET/map?url=...&limit=100
Every link on a page, no fetching.

/v1/fetch, /v1/crawl and /v1/map are versioned aliases.

Parametersโ€‹

ParameterTypeDefaultDescription
urlstringโ€”โœ… Required. Full URL, or a bare hostname (example.com โ†’ https://example.com)
formatstringmarkdownmarkdown, text, html, or json (metadata only)
main0 | 111 isolates the main article; 0 keeps the whole page
maxnumber20000Max characters of content returned (500โ€“100000)
links0 | 10Also return every link found on the page
limitnumber5/crawl: how many extra pages to follow (1โ€“10) ยท /map: max links (1โ€“500)
same_host0 | 11For /crawl and /map: stay on the same domain

Responseโ€‹

{
"success": true,
"url": "https://nextjs.org/blog/next-16",
"final_url": "https://nextjs.org/blog/next-16",
"status_code": 200,
"content_type": "text/html",
"title": "Next.js 16",
"description": "Next.js 16 includes Cache Components, stable Turbopack...",
"content": "# Next.js 16\n\nAhead of our upcoming Next.js Conf...",
"format": "markdown",
"words": 1420,
"reading_time_min": 6,
"chars": 8903,
"lang": "en",
"site_name": "Next.js",
"image": "https://nextjs.org/og.png",
"canonical": "https://nextjs.org/blog/next-16",
"cached": false,
"took_ms": 711
}

/crawl returns the same shape for the root page plus a pages[] array with one entry per followed link.

Examplesโ€‹

# Read a page
curl "https://api.apimitra.in/fetch?url=https://nextjs.org/blog/next-16" \
-H "x-api-key: YOUR_KEY"

# Just the metadata โ€” cheap, no body parsing
curl "https://api.apimitra.in/fetch?url=example.com&format=json" \
-H "x-api-key: YOUR_KEY"

# Read a page and everything it links to on the same site
curl "https://api.apimitra.in/crawl?url=https://blog.rust-lang.org/&limit=5" \
-H "x-api-key: YOUR_KEY"

Search, then fetch the winnersโ€‹

import requests

HEADERS = {"x-api-key": "YOUR_KEY"}
BASE = "https://api.apimitra.in"

hits = requests.get(f"{BASE}/search",
params={"q": "postgres index types", "limit": 5},
headers=HEADERS).json()["results"]

context = []
for i, hit in enumerate(hits, 1):
page = requests.get(f"{BASE}/fetch",
params={"url": hit["url"], "max": 6000},
headers=HEADERS).json()
if page.get("content"):
context.append(f"[{i}] {page['title']}\nURL: {page['url']}\n{page['content']}")

print("\n\n".join(context)) # hand this to your model
main=1 is the default for a reason

Most pages are 80% chrome. Keeping main=1 usually cuts the payload by 5โ€“10ร— and gives a model far less to wade through. Set main=0 only when the layout itself matters.

Behaviour and limitsโ€‹

  • Caching โ€” a fetched URL is cached for 15 minutes; repeats return in about a millisecond.
  • Markdown output โ€” headings, links, images, lists, tables and fenced code blocks with language tags are preserved. Everything else is flattened to text.
  • Redirects are followed (up to 5), and final_url tells you where you landed.
  • Safety โ€” every hop, including redirects, is resolved and checked against private address ranges. localhost, 127.0.0.1, 10.x, 192.168.x, link-local and cloud metadata endpoints (169.254.169.254) are refused, so the endpoint cannot be used to probe the internal network.
  • SSRF and protocol โ€” only http and https are accepted.
  • Client errors return 400 with a readable reason; upstream failures return 502.
  • JavaScript-rendered pages โ€” content injected by client-side JS is not visible to this endpoint. Server-rendered pages, blogs, docs, news sites and Wikipedia-style pages work well; single-page apps that render entirely in the browser will return sparse content. Use format=json on those to at least get the metadata.
  • Binary files โ€” PDFs and other binary payloads are not parsed; the response carries an error field explaining this.
Response shape on failure

A missing or blocked URL still returns a well-formed JSON body with success: false and an error string, so callers can branch on one field instead of juggling status codes.