canonical: https://jentic.com/apis/generativelanguage.googleapis.com/gemini-api

# Google Gemini API

The Google Gemini API generates model responses from multimodal prompts that combine text, images, audio, video, and code. It supports single-shot and streamed content generation, grounded answers, text embeddings, and token counting, and lets you fine-tune models and run requests in batches. Retrieval features manage corpora and file search stores, cached content reuses large contexts, and the file service holds media for prompts. Requests authenticate with a Google API key sent in the x-goog-api-key header or as a query key.

## For AI agents

Generate multimodal content, stream responses, create embeddings, count tokens, fine-tune models, and build retrieval corpora with Google Gemini models. Authenticated with a Google API key.

## Scope

Does not host a vector database, run image or video rendering pipelines, or manage Google Cloud billing. Use for Gemini model inference, embeddings, tuning, and retrieval only.

## Capabilities

- Generate content from text, image, audio, and video prompts, with streaming responses
- Create text embeddings for semantic search, individually or in batches
- Count tokens in a prompt before sending it
- Fine-tune models on your own data and generate from the tuned model
- Build retrieval corpora and file search stores and import documents into them
- Cache large contexts so they can be reused across requests
- Upload and manage files used as inputs to multimodal prompts

## Use cases

### Agent-Driven Content Generation

An AI agent connected through Jentic can call Gemini for generation without a developer wiring the Google API key and the model query parameter. The agent lists the available models, counts the tokens in its prompt to stay in budget, and generates or streams a response. Jentic injects the API key at call time so the credential never reaches the agent.

Example prompt: List the available models, count the tokens in a prompt, then generate a response and return it

### Retrieval-Augmented Generation

Teams grounding answers in their own documents can build a retrieval store on Gemini. The agent creates a file search store, imports files into it, and generates grounded answers that cite the stored content. This keeps responses anchored to a controlled corpus rather than open-ended generation.

Example prompt: Create a file search store, import a document into it, then generate a grounded answer that uses the stored content

### Embeddings for Semantic Search

Developers building semantic search can turn text into vectors with Gemini and store them for retrieval. The agent batch-embeds a set of documents and returns the vectors ready to index in a vector database. This powers search and recommendation over a corpus of text.

Example prompt: Batch-embed a set of documents with an embedding model and return the vectors for indexing

## Key endpoints

| Method | Path | Description |
| --- | --- | --- |
| POST | `/v1beta/models/{modelsId}:generateContent` | Generate a model response |
| POST | `/v1beta/models/{modelsId}:streamGenerateContent` | Stream a model response |
| POST | `/v1beta/models/{modelsId}:embedContent` | Create a text embedding |
| POST | `/v1beta/models/{modelsId}:countTokens` | Count tokens in a prompt |
| GET | `/v1beta/models` | List available models |
| POST | `/v1beta/tunedModels` | Create a tuned model |
| POST | `/v1beta/cachedContents` | Create cached content |

## Key resources

- **Models** — List models and generate, stream, embed, and count tokens
- **Tuned Models** — Create, read, update, and delete tuned models and generate from them
- **Corpora & File Search Stores** — Build retrieval stores, import files, and manage documents and permissions
- **Cached Contents** — Create, read, update, and delete cached contexts for reuse
- **Files** — Create, read, and delete files used as prompt inputs
- **Batches** — Enqueue and manage batch generation and embedding requests

## AI readiness

This API is usable in Jentic One now. Its AI-readiness score against Jentic's framework shows where it stands today and where improvements would make it even easier for agents to use.

- **Score:** 54 / 100
- **Maturity:** Foundational
- **Dimensions:**
  - Foundational Compliance: 74 / 100
  - Developer Experience & Jentic Compatibility: 56 / 100
  - AI-Readiness & Agent Experience: 43 / 100
  - Agent Usability: 88 / 100
  - Security: 33 / 100
  - AI Discoverability: 61 / 100
- **View full report:** https://jentic.com/apis/generativelanguage.googleapis.com/gemini-api/scorecard
- **How the score is calculated:** https://docs.jentic.com/reference/api-readiness-framework/overview/
- **More about the dimensions:** https://docs.jentic.com/reference/api-readiness-framework/specification/#dimensional-model-overview

### Score it yourself

Every API in the directory is allowlisted, so you can re-score it with no key required.

- **Score your own API:** https://jentic.com/scorecard.md
- **Scoring CLI agent skill:** https://github.com/jentic/jentic-api-scorecard/blob/main/skills/jentic-api-scorecard/SKILL.md

```sh
npx @jentic/api-scorecard-cli score <openapi-url>
```

## Why Jentic

- **Setup:** Wiring Gemini by hand means managing a Google API key, passing the model in a query parameter, and matching each query to the right generate, embed, or tune operation yourself. Through Jentic you install once, import Gemini from the API Directory, store the key once, and your agent calls it.
- **Permission scoping:** You choose which Gemini operations the agent may call, such as content generation and token counting, so destructive ones like deleting a tuned model, corpus, or cached content are not included unless you add them. Gemini carries the model in a query parameter and resource ids in the path, so rules bound which operations your agent may call.
- **Credential handling:** Your Google API key is stored once, encrypted, by your own Jentic One instance and injected at execution time. It never enters the agent's prompt, logs, or context.
- **Discovery method:** Agents search Jentic by intent such as 'generate a response from a prompt' or 'create a text embedding', and Jentic returns the matching Gemini operation with its input schema so the agent calls the right endpoint without browsing the reference docs.

## Related APIs

- **OpenAI** — GPT models for text, chat, embeddings, and vision
- **Anthropic** — Claude models for chat, reasoning, and long context
- **Mistral** — Open-weight and hosted models for text and code
- **Pinecone** — Managed vector database for storing and querying embeddings

## FAQ

### What authentication does the Gemini API use?

Per its OpenAPI spec, the Gemini API uses a Google API key, sent either in the x-goog-api-key header or as a key query parameter. Through Jentic the key is stored encrypted by your own instance and injected at call time, so it never reaches the agent.

### Is there a Gemini MCP server?

You don't need an MCP server to give your agent Gemini. Jentic connects it directly from the API Directory: import it, store your credential once, and your agent calls operations like generating content or creating an embedding on demand, without loading another server's tool definitions into its context.

### Can I limit what my agent is allowed to do with Gemini?

Yes. Write a rule that allows only the generation and embedding operations, such as generating content and counting tokens, so the agent can produce responses but cannot delete a tuned model, corpus, or cached content unless you add those operations, and every call it makes is logged. This matches an inference bot that generates without managing resources.

### Can I generate multimodal responses with the Gemini API?

Yes. Call the content generation operation with a prompt that combines text with images, audio, or video, and the model returns a response. Use the streaming operation when you want tokens as they are produced.

### What are the rate limits for the Gemini API?

The OpenAPI spec does not specify rate limits. Check the Gemini API documentation at https://ai.google.dev/gemini-api/docs/rate-limits for the current per-model and per-tier limits before scaling up generation.

### How do I generate content with the Gemini API through Jentic?

Search Jentic for 'generate a response from a prompt', which returns the model list and content generation operations with their input schemas. The agent picks a model, counts tokens, and generates a response, with your stored key injected at call time. To run it on your own infrastructure, install Jentic One from its GitHub repo.
