For Agents
Generate multimodal content, stream responses, create embeddings, count tokens, fine-tune models, and build retrieval corpora with Google Gemini models. Authenticated with a Google API key.
Use for: I need to generate a response from a text prompt, Stream a long generated response token by token, Create an embedding for a piece of text, Count the tokens in a prompt before sending it
Not supported: Does not host a vector database, run image or video rendering pipelines, or manage Google Cloud billing. Use for Gemini model inference, embeddings, tuning, and retrieval only.
The Google Gemini API generates model responses from multimodal prompts that combine text, images, audio, video, and code. It supports single-shot and streamed content generation, grounded answers, text embeddings, and token counting, and lets you fine-tune models and run requests in batches. Retrieval features manage corpora and file search stores, cached content reuses large contexts, and the file service holds media for prompts. Requests authenticate with a Google API key sent in the x-goog-api-key header or as a query key.
Install Jentic One Beta
Jentic One is a self-hosted execution layer for AI agents. It lets your agent call the Gemini API, or any other public or private API you need. You set the rules, the agent never sees your credentials, and every call is logged.
Two steps, two machines. Install the instance in a safe environment, then register your agent from wherever it runs.
Step 1: Jentic One Host machine
# On the machine that will host your Jentic One instance:
curl -fsSL "https://jentic.com/install.sh?src=apis&api=%2Fapis%2Fgenerativelanguage.googleapis.com%2Fgemini-api" | shStep 2: Agent machine
# On the machine where your agent runs (keep this separate from the instance):
curl -fsSL "https://jentic.com/install.sh?src=apis&api=%2Fapis%2Fgenerativelanguage.googleapis.com%2Fgemini-api" | sh
jentic register # connects your agent to your Jentic One instanceJentic One is in public beta. The setup above keeps your agent separate from the instance, which is what you want before using real credentials: an agent running as the same OS user as Jentic One can read its stored keys directly. Just evaluating? A single local install is fine to start. See the secure deployment guide for the tiers.
What an agent can do with Gemini API.
Generate content from text, image, audio, and video prompts, with streaming responses
Create text embeddings for semantic search, individually or in batches
Count tokens in a prompt before sending it
Fine-tune models on your own data and generate from the tuned model
Build retrieval corpora and file search stores and import documents into them
Cache large contexts so they can be reused across requests
Upload and manage files used as inputs to multimodal prompts
Patterns agents use Gemini API for, with concrete tasks.
★ Agent-Driven Content Generation
An AI agent connected through Jentic can call Gemini for generation without a developer wiring the Google API key and the model query parameter. The agent lists the available models, counts the tokens in its prompt to stay in budget, and generates or streams a response. Jentic injects the API key at call time so the credential never reaches the agent.
List the available models, count the tokens in a prompt, then generate a response and return it
Retrieval-Augmented Generation
Teams grounding answers in their own documents can build a retrieval store on Gemini. The agent creates a file search store, imports files into it, and generates grounded answers that cite the stored content. This keeps responses anchored to a controlled corpus rather than open-ended generation.
Create a file search store, import a document into it, then generate a grounded answer that uses the stored content
Embeddings for Semantic Search
Developers building semantic search can turn text into vectors with Gemini and store them for retrieval. The agent batch-embeds a set of documents and returns the vectors ready to index in a vector database. This powers search and recommendation over a corpus of text.
Batch-embed a set of documents with an embedding model and return the vectors for indexing
86 endpoints — the google gemini api generates model responses from multimodal prompts that combine text, images, audio, video, and code.
METHOD
PATH
DESCRIPTION
/v1beta/models/{modelsId}:generateContent
Generate a model response
/v1beta/models/{modelsId}:streamGenerateContent
Stream a model response
/v1beta/models/{modelsId}:embedContent
Create a text embedding
/v1beta/models/{modelsId}:countTokens
Count tokens in a prompt
/v1beta/models
List available models
/v1beta/tunedModels
Create a tuned model
/v1beta/cachedContents
Create cached content
/v1beta/models/{modelsId}:generateContent
Generate a model response
/v1beta/models/{modelsId}:streamGenerateContent
Stream a model response
/v1beta/models/{modelsId}:embedContent
Create a text embedding
/v1beta/models/{modelsId}:countTokens
Count tokens in a prompt
/v1beta/models
List available models
/v1beta/tunedModels
Create a tuned model
/v1beta/cachedContents
Create cached content
This API is usable in Jentic One now. Its AI-readiness score against Jentic's framework shows where it stands today and where improvements would make it even easier for agents to use.
Base layer of spec validity and structural soundness.
Aggregated quality score from linter diagnostics, weighted by severity.
Percentage of `$ref` references that resolve successfully.
Checks whether the API description parses successfully and conforms to its declared specification (e.g., OpenAPI).
Structural correctness score based on schema issues using logarithmic dampening.
Clarity, completeness, and ingestion readiness for developers and tooling.
How richly the API is illustrated with examples.
Percentage of examples that conform to their schemas.
Percentage of operations with complete response definitions (success, client error, server error).
Health of API ingestion, bundling, and resolution within Jentic pipelines.
Semantic breadth, depth, and agent comprehension for AI systems.
Coverage of descriptions across API elements.
Coverage of RFC 9457 Problem Details for error responses.
Coverage, uniqueness, and casing consistency of operationIds for AI inference.
Coverage of summaries across operations/tags/info.
Functional utility, complexity comfort, and AI orchestration readiness.
Agent comfort level based on API operational and structural complexity.
Trust, risk posture, and security compliance.
Average quality of security schemes based on authentication method strength (weakest link for OAuth2).
Findability, semantic richness, and reasoning readiness.
Clarity and depth of descriptions across API elements.
Score it yourself
Every API in the directory is allowlisted, so you can re-score it with no key required.
npx @jentic/api-scorecard-cli score <openapi-url>What agents get from Jentic-routed access to this vendor.
Setup
Wiring Gemini by hand means managing a Google API key, passing the model in a query parameter, and matching each query to the right generate, embed, or tune operation yourself. Through Jentic you install once, import Gemini from the API Directory, store the key once, and your agent calls it.
Permission scoping
You choose which Gemini operations the agent may call, such as content generation and token counting, so destructive ones like deleting a tuned model, corpus, or cached content are not included unless you add them. Gemini carries the model in a query parameter and resource ids in the path, so rules bound which operations your agent may call.
Credential isolation
Your Google API key is stored once, encrypted, by your own Jentic One instance and injected at execution time. It never enters the agent's prompt, logs, or context.
Intent-based discovery
Agents search Jentic by intent such as 'generate a response from a prompt' or 'create a text embedding', and Jentic returns the matching Gemini operation with its input schema so the agent calls the right endpoint without browsing the reference docs.
Alternatives and complements available in the Jentic catalogue.
Specific to using Gemini API through Jentic.
What authentication does the Gemini API use?
Per its OpenAPI spec, the Gemini API uses a Google API key, sent either in the x-goog-api-key header or as a key query parameter. Through Jentic the key is stored encrypted by your own instance and injected at call time, so it never reaches the agent.
Is there a Gemini MCP server?
You don't need an MCP server to give your agent Gemini. Jentic connects it directly from the API Directory: import it, store your credential once, and your agent calls operations like generating content or creating an embedding on demand, without loading another server's tool definitions into its context.
Can I limit what my agent is allowed to do with Gemini?
Yes. Write a rule that allows only the generation and embedding operations, such as generating content and counting tokens, so the agent can produce responses but cannot delete a tuned model, corpus, or cached content unless you add those operations, and every call it makes is logged. This matches an inference bot that generates without managing resources.
Can I generate multimodal responses with the Gemini API?
Yes. Call the content generation operation with a prompt that combines text with images, audio, or video, and the model returns a response. Use the streaming operation when you want tokens as they are produced.
What are the rate limits for the Gemini API?
The OpenAPI spec does not specify rate limits. Check the Gemini API documentation at https://ai.google.dev/gemini-api/docs/rate-limits for the current per-model and per-tier limits before scaling up generation.
How do I generate content with the Gemini API through Jentic?
Search Jentic for 'generate a response from a prompt', which returns the model list and content generation operations with their input schemas. The agent picks a model, counts tokens, and generates a response, with your stored key injected at call time. To run it on your own infrastructure, install Jentic One from its GitHub repo.
GET STARTED