Install Jentic One Beta
Jentic One is a self-hosted execution layer for AI agents. It lets your agent call the BentoCloud API, or any other public or private API you need. You set the rules, the agent never sees your credentials, and every call is logged.
Two steps, two machines. Install the instance in a safe environment, then register your agent from wherever it runs.
Step 1: Jentic One Host machine
# On the machine that will host your Jentic One instance:
curl -fsSL "https://jentic.com/install.sh?src=apis&api=%2Fapis%2Fbentoml.com%2Fbentoml" | shStep 2: Agent machine
# On the machine where your agent runs (keep this separate from the instance):
curl -fsSL "https://jentic.com/install.sh?src=apis&api=%2Fapis%2Fbentoml.com%2Fbentoml" | sh
jentic register # connects your agent to your Jentic One instanceJentic One is in public beta. The setup above keeps your agent separate from the instance, which is what you want before using real credentials: an agent running as the same OS user as Jentic One can read its stored keys directly. Just evaluating? A single local install is fine to start. See the secure deployment guide for the tiers.
What an agent can do with BentoCloud API.
Deploy machine-learning models as inference endpoints on a cluster
List and manage bentos and model repositories
Create, start, and terminate serving endpoints
Manage clusters, deployments, and their revisions
Handle organization secrets, members, and usage metrics
Patterns agents use BentoCloud API for, with concrete tasks.
★ Model deployment automation
An MLOps pipeline needs to promote a packaged model to a serving endpoint. The BentoCloud API creates a deployment on a cluster from a bento and exposes it as an endpoint, so a pipeline can ship a new model version programmatically.
Create a deployment via POST /api/v1/clusters/{clusterName}/deployments from a selected bento
Inference endpoint lifecycle
A platform team needs to start, scale, and retire serving endpoints as demand shifts. The endpoint operations create an endpoint, start it, and terminate it so capacity tracks load without manual console work.
List endpoints via GET /api/v1/endpoints, then terminate an idle one via POST /api/v1/endpoints/{endpointUID}/terminate
Model and bento inventory
A registry integration needs to know which packaged models exist. The bento and model operations list all bentos and models across repositories so a backend can audit what is available to deploy.
Call GET /api/v1/bentos and GET /api/v1/models to inventory available artifacts
222 endpoints — the bentocloud api deploys and operates machine-learning models in production: managing bentos and model repositories, creating deployments and endpoints on clusters, and scaling inference services.
METHOD
PATH
DESCRIPTION
/api/v1/deployments
List deployments in the organization
/api/v1/clusters/{clusterName}/deployments
Create a deployment on a cluster
/api/v1/bentos
List all bentos
/api/v1/models
List all models
/api/v1/endpoints
List organization endpoints
What agents get from Jentic-routed access to this vendor.
Setup
Wiring the BentoCloud API by hand means setting the X-YATAI-API-TOKEN header on every request and shaping the deployment, endpoint, and cluster payloads yourself. Through Jentic you install once, import the API from the API Directory, store your credential once, and your agent calls it.
Permission scoping
The BentoCloud API addresses clusters, deployments, and endpoints by id in the URL path, so you can limit the agent to the operations it needs, such as listing bentos or creating a deployment. You credit the agent only with the operations you allow, so terminating a deployment stays out unless you add it.
Credential isolation
Your the BentoCloud credential is stored once, encrypted, by your own Jentic One instance and injected at execution time. It never enters the agent's prompt, logs, or context.
Intent-based discovery
Agents search Jentic by intent such as 'deploy a model endpoint', and Jentic returns the matching BentoCloud operation with its input schema so the agent calls the right endpoint without browsing the reference docs.
Alternatives and complements available in the Jentic catalogue.
Specific to using BentoCloud API through Jentic.
What authentication does the BentoCloud API use?
The BentoCloud API authenticates with an API token sent in the X-YATAI-API-TOKEN header. Through Jentic the token is stored encrypted in your own instance and injected into each request at run time.
Is there a the BentoCloud API MCP server?
You don't need an MCP server to give your agent the BentoCloud API. Jentic connects it directly from the API Directory: import it, store your credential once, and your agent calls it. That keeps another server's tool definitions out of your agent's context while still giving it the operation set from the spec.
Can the BentoCloud API deploy and scale model endpoints?
Yes. It creates deployments on clusters from packaged bentos, exposes them as endpoints, and lets you start and terminate endpoints so serving capacity can track demand.
Can I limit what my agent is allowed to do with the BentoCloud API?
Yes. Because you run Jentic One yourself, your own rules decide which the BentoCloud API operations and credentials the agent may use. You grant only the operations it needs, such as deploying a model, so deleting a cluster stays off limits unless you add it. The choice of what the agent may call is yours.
How do I connect my agent to the BentoCloud API through Jentic?
Search Jentic by intent such as 'deploy a model endpoint' to discover the operation, load its input schema, and your agent calls it with your stored credential injected at run time. To run it on your own infrastructure, install Jentic One from its GitHub repo.
GET STARTED
AI agent model operations
An AI agent managing inference infrastructure can deploy and scale models without hand-wiring the BentoCloud token. Through Jentic the agent discovers the deployment operations by intent and calls them with the stored credential.
Search Jentic for 'deploy a model endpoint', load the schema, and create a deployment on a cluster
/api/v1/endpoints
Create a serving endpoint
/api/v1/deployments
List deployments in the organization
/api/v1/clusters/{clusterName}/deployments
Create a deployment on a cluster
/api/v1/bentos
List all bentos
/api/v1/models
List all models
/api/v1/endpoints
List organization endpoints
/api/v1/endpoints
Create a serving endpoint
For Agents
Deploy and manage machine-learning inference endpoints on BentoCloud, list bentos and models, and scale deployments across clusters.
Use for: I need to deploy a model as an inference endpoint, List all the bentos in my organization, Create a deployment on a cluster, Terminate a running endpoint
Not supported: Does not handle model training, data labeling, or foundation-model inference itself - use for deploying and operating model serving on BentoCloud only.
The BentoCloud API deploys and operates machine-learning models in production: managing bentos and model repositories, creating deployments and endpoints on clusters, and scaling inference services. It covers organizations, clusters, secrets, and usage metrics for running models at scale. Requests authenticate with an API token sent in the X-YATAI-API-TOKEN header.
This API is usable in Jentic One now. Its AI-readiness score against Jentic's framework shows where it stands today and where improvements would make it even easier for agents to use.
Base layer of spec validity and structural soundness.
Aggregated quality score from linter diagnostics, weighted by severity.
Percentage of `$ref` references that resolve successfully.
Checks whether the API description parses successfully and conforms to its declared specification (e.g., OpenAPI).
Structural correctness score based on schema issues using logarithmic dampening.
Clarity, completeness, and ingestion readiness for developers and tooling.
How richly the API is illustrated with examples.
Percentage of examples that conform to their schemas.
Percentage of operations with complete response definitions (success, client error, server error).
Health of API ingestion, bundling, and resolution within Jentic pipelines.
Semantic breadth, depth, and agent comprehension for AI systems.
Coverage of descriptions across API elements.
Coverage of RFC 9457 Problem Details for error responses.
Coverage, uniqueness, and casing consistency of operationIds for AI inference.
Coverage of summaries across operations/tags/info.
Functional utility, complexity comfort, and AI orchestration readiness.
Agent comfort level based on API operational and structural complexity.
Trust, risk posture, and security compliance.
Average quality of security schemes based on authentication method strength (weakest link for OAuth2).
Findability, semantic richness, and reasoning readiness.
Clarity and depth of descriptions across API elements.
Score it yourself
Every API in the directory is allowlisted, so you can re-score it with no key required.
npx @jentic/api-scorecard-cli score <openapi-url>