canonical: https://jentic.com/apis/diffbot.com/diffbot

# Diffbot Extract APIs

Jentic publishes the only available OpenAPI specification for Diffbot's Extract APIs, keeping it validated and agent-ready. The Diffbot Extract API turns a web page URL into normalized structured JSON. Its analyze operation detects the page type and routes it to the right extractor, and dedicated extractors return article, product, discussion, job, image, video, event, and list data. Custom Extraction APIs let you define and manage your own field extraction for pages the built-in extractors do not cover.

## For AI agents

Turn web page URLs into structured JSON: analyze a page of unknown type, or call type-specific extractors for article, product, discussion, job, image, video, and event data.

## Scope

Does not handle web crawling at scale, search indexing, or knowledge graph queries. Use for extracting structured data from individual web pages only.

## Capabilities

- Analyze a page of unknown type and route it to the right extractor
- Extract clean article text, title, and metadata from a news or blog URL
- Extract structured product data from a product page
- Pull discussion, job, event, image, and video data from their respective page types
- Extract a list of items from a listing or category page
- Create, retrieve, and delete Custom Extraction APIs for bespoke field extraction

## Use cases

### Agent Web Content Extraction

An AI agent that needs clean structured data from arbitrary URLs can send a page to the Diffbot Extract API's analyze operation, which identifies the page type and returns normalized JSON fields. For known page types the agent can call the article, product, or discussion extractor directly instead of parsing raw HTML.

Example prompt: Send a news URL to the analyze operation and return the extracted article title, author, and body text

### Product Page Extraction

Catalog and price-monitoring tools can extract structured product data from retailer pages with the Diffbot Extract API's product operation. It returns normalized product fields from the page so downstream logic does not depend on each site's markup.

Example prompt: Extract structured product data from a given product URL and record the fields returned

### Custom Field Extraction

Teams with bespoke extraction needs can define a Custom Extraction API with the Diffbot Extract API's custom operations, then call it to pull their chosen fields from a target page. The create, retrieve, and delete operations manage those custom definitions.

Example prompt: Create a custom extraction API definition, then retrieve the list of defined custom APIs

## Key endpoints

| Method | Path | Description |
| --- | --- | --- |
| GET | `/analyze` | Analyze a page of unknown type |
| GET | `/article` | Extract an article |
| GET | `/product` | Extract product data |
| GET | `/discussion` | Extract discussion data |
| GET | `/image` | Extract image data |
| GET | `/custom` | Retrieve Custom APIs |
| POST | `/custom` | Create or update a Custom API |
| DELETE | `/custom` | Delete a Custom API |

## Key resources

- **Analyze** — Detect a page's type and route it to the appropriate extractor
- **Type extractors** — Extract article, product, discussion, job, image, video, event, and list data
- **Custom APIs** — Create, retrieve, and delete Custom Extraction API definitions

## AI readiness

This API is usable in Jentic One now. Its AI-readiness score against Jentic's framework shows where it stands today and where improvements would make it even easier for agents to use.

- **Score:** 43 / 100
- **Maturity:** Foundational
- **Dimensions:**
  - Foundational Compliance: 89 / 100
  - Developer Experience & Jentic Compatibility: 62 / 100
  - AI-Readiness & Agent Experience: 35 / 100
  - Agent Usability: 94 / 100
  - Security: 15 / 100
  - AI Discoverability: 64 / 100
- **View full report:** https://jentic.com/apis/diffbot.com/diffbot/scorecard
- **How the score is calculated:** https://docs.jentic.com/reference/api-readiness-framework/overview/
- **More about the dimensions:** https://docs.jentic.com/reference/api-readiness-framework/specification/#dimensional-model-overview

### Score it yourself

Every API in the directory is allowlisted, so you can re-score it with no key required.

- **Score your own API:** https://jentic.com/scorecard.md
- **Scoring CLI agent skill:** https://github.com/jentic/jentic-api-scorecard/blob/main/skills/jentic-api-scorecard/SKILL.md

```sh
npx @jentic/api-scorecard-cli score <openapi-url>
```

## Why Jentic

- **Setup:** Wiring the Diffbot Extract API by hand means managing its API token on the query string, choosing between the analyze operation and the type-specific extractors, and handling each page type's response yourself. Through Jentic you install once, import the Diffbot Extract API from the API Directory, store the token once, and your agent calls it.
- **Permission scoping:** The Diffbot Extract API carries the target page URL as a query parameter, so rules bound which operations your agent may call rather than which site it reads. Limit it to the extraction operations it needs, such as analyze and article, and leave the custom-API create and delete operations out unless you add them.
- **Credential handling:** Your Diffbot token is stored once, encrypted, by your own Jentic One instance and injected at execution time. It never enters the agent's prompt, logs, or context.
- **Discovery method:** Agents search Jentic by intent such as 'extract an article from a URL' or 'analyze a web page', and Jentic returns the matching Diffbot operation with its input schema so the agent calls the right endpoint without reading the reference docs.

## Related APIs

- **Firecrawl** — Firecrawl turns URLs into clean markdown or structured data, an alternative to the Diffbot Extract API's JSON extraction.
- **Apify** — Apify runs configurable scraping actors, a more general platform than the Diffbot Extract API's typed extractors.
- **ZenRows** — ZenRows fetches pages past anti-bot defenses, complementing the Diffbot Extract API's structured parsing.

## FAQ

### Why is there no official OpenAPI spec for Diffbot's Extract APIs?

Diffbot does not publish an OpenAPI specification for its Extract APIs. Jentic generates and maintains this spec so that AI agents and developers can call the Diffbot Extract API via structured tooling. It is validated against the live API and kept up to date. To run it on your own infrastructure, install Jentic One from its GitHub repo.

### What authentication does the Diffbot Extract API use?

The Diffbot Extract API authenticates with an API token passed as a `token` query parameter, per its OpenAPI spec. Through Jentic the token is stored encrypted by your own Jentic One instance and injected at call time, so it never enters the agent's prompt or logs.

### Can I extract article data from a URL with the Diffbot Extract API?

Yes. The article operation returns normalized article fields such as text, title, and author from a page URL, and the analyze operation can first detect the page type if it is unknown. There are also dedicated extractors for products, discussions, jobs, images, videos, and events.

### What are the rate limits for the Diffbot Extract API?

The OpenAPI spec does not specify rate limits for the Diffbot Extract API. Check the provider's documentation at https://docs.diffbot.com for the current limits that apply to your token.

### Can I limit what my agent is allowed to do with the Diffbot Extract API?

Yes. The target page URL travels as a query parameter, so a rule bounds which operations the agent may call rather than which site it can read. Write a rule that allows only the extraction operations you need, such as analyze and article, and leave the custom-API create and delete operations out unless the agent needs them, and every call it makes is logged.

### How do I extract page data with the Diffbot Extract API through Jentic?

Search Jentic for 'extract article data from a URL' and it returns the Diffbot article operation with its input schema. Import the Diffbot Extract API from the Jentic API Directory, store your token once, and your agent can analyze pages and call the type-specific extractors without hand-wiring the token on each request.

### Is there a Diffbot Extract API MCP server?

You don't need an MCP server to give your agent the Diffbot Extract API. Jentic connects it directly from the API Directory: import it, store your token once, and your agent calls the analyze and extraction operations. Nothing extra loads into the agent's context until an operation is actually used.
