Install Jentic One Beta
Jentic One is a self-hosted execution layer for AI agents. It lets your agent call the Extract APIs, or any other public or private API you need. You set the rules, the agent never sees your credentials, and every call is logged.
Two steps, two machines. Install the instance in a safe environment, then register your agent from wherever it runs.
Step 1: Jentic One Host machine
# On the machine that will host your Jentic One instance:
curl -fsSL "https://jentic.com/install.sh?src=apis&api=%2Fapis%2Fdiffbot.com%2Fdiffbot" | shStep 2: Agent machine
# On the machine where your agent runs (keep this separate from the instance):
curl -fsSL "https://jentic.com/install.sh?src=apis&api=%2Fapis%2Fdiffbot.com%2Fdiffbot" | sh
jentic register # connects your agent to your Jentic One instanceJentic One is in public beta. The setup above keeps your agent separate from the instance, which is what you want before using real credentials: an agent running as the same OS user as Jentic One can read its stored keys directly. Just evaluating? A single local install is fine to start. See the secure deployment guide for the tiers.
What an agent can do with Extract APIs.
Analyze a page of unknown type and route it to the right extractor
Extract clean article text, title, and metadata from a news or blog URL
Extract structured product data from a product page
Pull discussion, job, event, image, and video data from their respective page types
Extract a list of items from a listing or category page
GET STARTED
Create, retrieve, and delete Custom Extraction APIs for bespoke field extraction
Patterns agents use Extract APIs for, with concrete tasks.
★ Agent Web Content Extraction
An AI agent that needs clean structured data from arbitrary URLs can send a page to the Diffbot Extract API's analyze operation, which identifies the page type and returns normalized JSON fields. For known page types the agent can call the article, product, or discussion extractor directly instead of parsing raw HTML.
Send a news URL to the analyze operation and return the extracted article title, author, and body text
Product Page Extraction
Catalog and price-monitoring tools can extract structured product data from retailer pages with the Diffbot Extract API's product operation. It returns normalized product fields from the page so downstream logic does not depend on each site's markup.
Extract structured product data from a given product URL and record the fields returned
Custom Field Extraction
Teams with bespoke extraction needs can define a Custom Extraction API with the Diffbot Extract API's custom operations, then call it to pull their chosen fields from a target page. The create, retrieve, and delete operations manage those custom definitions.
Create a custom extraction API definition, then retrieve the list of defined custom APIs
13 endpoints — jentic publishes the only available openapi specification for diffbot's extract apis, keeping it validated and agent-ready.
METHOD
PATH
DESCRIPTION
/analyze
Analyze a page of unknown type
/article
Extract an article
/product
Extract product data
/discussion
Extract discussion data
/image
Extract image data
/custom
Retrieve Custom APIs
/custom
Create or update a Custom API
/custom
Delete a Custom API
/analyze
Analyze a page of unknown type
/article
Extract an article
/product
Extract product data
/discussion
Extract discussion data
/image
Extract image data
/custom
Retrieve Custom APIs
/custom
Create or update a Custom API
/custom
Delete a Custom API
What agents get from Jentic-routed access to this vendor.
Setup
Wiring the Diffbot Extract API by hand means managing its API token on the query string, choosing between the analyze operation and the type-specific extractors, and handling each page type's response yourself. Through Jentic you install once, import the Diffbot Extract API from the API Directory, store the token once, and your agent calls it.
Permission scoping
The Diffbot Extract API carries the target page URL as a query parameter, so rules bound which operations your agent may call rather than which site it reads. Limit it to the extraction operations it needs, such as analyze and article, and leave the custom-API create and delete operations out unless you add them.
Credential isolation
Your Diffbot token is stored once, encrypted, by your own Jentic One instance and injected at execution time. It never enters the agent's prompt, logs, or context.
Intent-based discovery
Agents search Jentic by intent such as 'extract an article from a URL' or 'analyze a web page', and Jentic returns the matching Diffbot operation with its input schema so the agent calls the right endpoint without reading the reference docs.
Alternatives and complements available in the Jentic catalogue.
Specific to using Extract APIs through Jentic.
Why is there no official OpenAPI spec for Diffbot's Extract APIs?
Diffbot does not publish an OpenAPI specification for its Extract APIs. Jentic generates and maintains this spec so that AI agents and developers can call the Diffbot Extract API via structured tooling. It is validated against the live API and kept up to date. To run it on your own infrastructure, install Jentic One from its GitHub repo.
What authentication does the Diffbot Extract API use?
The Diffbot Extract API authenticates with an API token passed as a `token` query parameter, per its OpenAPI spec. Through Jentic the token is stored encrypted by your own Jentic One instance and injected at call time, so it never enters the agent's prompt or logs.
Can I extract article data from a URL with the Diffbot Extract API?
Yes. The article operation returns normalized article fields such as text, title, and author from a page URL, and the analyze operation can first detect the page type if it is unknown. There are also dedicated extractors for products, discussions, jobs, images, videos, and events.
What are the rate limits for the Diffbot Extract API?
The OpenAPI spec does not specify rate limits for the Diffbot Extract API. Check the provider's documentation at https://docs.diffbot.com for the current limits that apply to your token.
Can I limit what my agent is allowed to do with the Diffbot Extract API?
Yes. The target page URL travels as a query parameter, so a rule bounds which operations the agent may call rather than which site it can read. Write a rule that allows only the extraction operations you need, such as analyze and article, and leave the custom-API create and delete operations out unless the agent needs them, and every call it makes is logged.
How do I extract page data with the Diffbot Extract API through Jentic?
Search Jentic for 'extract article data from a URL' and it returns the Diffbot article operation with its input schema. Import the Diffbot Extract API from the Jentic API Directory, store your token once, and your agent can analyze pages and call the type-specific extractors without hand-wiring the token on each request.
Is there a Diffbot Extract API MCP server?
You don't need an MCP server to give your agent the Diffbot Extract API. Jentic connects it directly from the API Directory: import it, store your token once, and your agent calls the analyze and extraction operations. Nothing extra loads into the agent's context until an operation is actually used.
Know of an official OpenAPI document? Contribute it →
For Agents
Turn web page URLs into structured JSON: analyze a page of unknown type, or call type-specific extractors for article, product, discussion, job, image, video, and event data.
Use for: Analyze a web page when I don't know its type, Extract the article text and author from a news URL, Get structured product data from a product page, Pull the discussion posts from a forum thread
Not supported: Does not handle web crawling at scale, search indexing, or knowledge graph queries. Use for extracting structured data from individual web pages only.
Jentic publishes the only available OpenAPI specification for Diffbot's Extract APIs, keeping it validated and agent-ready. The Diffbot Extract API turns a web page URL into normalized structured JSON. Its analyze operation detects the page type and routes it to the right extractor, and dedicated extractors return article, product, discussion, job, image, video, event, and list data. Custom Extraction APIs let you define and manage your own field extraction for pages the built-in extractors do not cover.
This API is usable in Jentic One now. Its AI-readiness score against Jentic's framework shows where it stands today and where improvements would make it even easier for agents to use.
Base layer of spec validity and structural soundness.
Aggregated quality score from linter diagnostics, weighted by severity.
Percentage of `$ref` references that resolve successfully.
Checks whether the API description parses successfully and conforms to its declared specification (e.g., OpenAPI).
Structural correctness score based on schema issues using logarithmic dampening.
Clarity, completeness, and ingestion readiness for developers and tooling.
How richly the API is illustrated with examples.
Percentage of examples that conform to their schemas.
Percentage of operations with complete response definitions (success, client error, server error).
Health of API ingestion, bundling, and resolution within Jentic pipelines.
Semantic breadth, depth, and agent comprehension for AI systems.
Coverage of descriptions across API elements.
Coverage of RFC 9457 Problem Details for error responses.
Coverage, uniqueness, and casing consistency of operationIds for AI inference.
Coverage of summaries across operations/tags/info.
Functional utility, complexity comfort, and AI orchestration readiness.
Agent comfort level based on API operational and structural complexity.
Trust, risk posture, and security compliance.
Average quality of security schemes based on authentication method strength (weakest link for OAuth2).
Findability, semantic richness, and reasoning readiness.
Clarity and depth of descriptions across API elements.
Score it yourself
Every API in the directory is allowlisted, so you can re-score it with no key required.
npx @jentic/api-scorecard-cli score <openapi-url>