Google Gemini Integration
Integrations

Paperclip + Gemini

Host Gemini agents on Paperclip — 3.1 Pro with 2M-token context, 3.5 Flash, and Flash-Lite, on a managed runtime from $3.99/mo.

Gemini agent hosting, in one answer

Google sells Gemini tokens through the Gemini API and, for GCP shops, agent tooling through Vertex AI. Paperclip gives your Gemini agents the rest of the production stack: a supervised runtime that restarts crashed processes, persistent memory between runs, per-run monitoring, and model routing across providers. Bring a Google AI API key and deploy in five minutes — 2M-token context for huge documents, one hosted bill instead of a GCP project.

Gemini agent models and prices (September 2026)

Gemini 3.1 Pro

Flagship with up to 2M-token context. $2.00 in / $12.00 out per 1M tokens. For huge documents and multimodal agents.

Fast $$
Gemini 3.5 Flash

Promo-priced at $1.50 in / $9.00 out per 1M tokens (rates double January 1, 2027). Balanced default for most agents.

Fast $$
Gemini Flash-Lite

High-volume tier at $0.30 in / $2.50 out per 1M tokens. For classification, routing, and extraction at scale.

Very Fast $

How to deploy Gemini agents

1

Get your Google AI API key

Go to aistudio.google.com → Get API key.

2

Add it to HostAgentes

Dashboard → Agent Settings → Environment Variables → Add GOOGLE_AI_API_KEY.

3

Select your model

Choose Gemini 3.1 Pro, 3.5 Flash, or Flash-Lite from the model dropdown — or let model routing pick per request.

4

Deploy

Click deploy. Your Gemini-powered agent is live with persistent memory, auto-restarts, and run-level monitoring.

Gemini agent capabilities

2M-token context

Gemini 3.1 Pro takes up to 2M tokens — the largest window of any provider we host.

Native multimodal

Text, image, video, and audio go in the same prompt.

Google Search grounding

Ground agent answers with live search results at the API level.

Thinking tokens

Extended reasoning bills as output tokens — factor it into agent budgets.

Cost optimization tip

Route high-volume tasks to Flash-Lite at $0.30 per 1M input tokens and keep huge-context work on 3.1 Pro. Watch two price cliffs: 3.5 Flash's promo rate doubles January 1, 2027, and the Pro rate doubles above 200K tokens. HostAgentes lets you switch models per agent — no redeployment needed.

HostAgentes vs Vertex AI Agent Builder

Vertex requires a GCP project, IAM, and service configuration before your first agent runs. Paperclip needs an API key: the same agent can route runs to Gemini, Claude, GPT-5.6, or Mistral, with memory, restarts, and monitoring handled identically regardless of model.

Model lineup and prices verified against Google's public Gemini API pricing in September 2026. Provider prices change; re-check before committing volume.

Frequently asked questions

Doesn't Google already host agents with Vertex AI Agent Builder?

Google's agent tooling (Vertex AI Agent Builder, Agentspace) runs agents inside Google Cloud, which means GCP projects, IAM, and service config. HostAgentes is the managed runtime around the Gemini API: bring a Google AI API key — no GCP account — and Paperclip handles process supervision, persistent memory, auto-restarts, per-run monitoring, and routing to Claude, GPT-5.6, or Mistral when a task suits another model.

Which Gemini model should my agent use?

Start with Gemini 3.5 Flash at $1.50 per 1M input tokens — but note its promo rate doubles on January 1, 2027. Use Flash-Lite ($0.30 input) for high-volume simple tasks and 3.1 Pro ($2.00 input, 2M context) for huge documents and multimodal workloads. HostAgentes lets you switch models per agent without redeploying.

How much does it cost to run Gemini agents?

A typical agent run spends about 1,000 input and 100 output tokens, so one million runs cost about $0.55 on Flash-Lite, $2.40 on 3.5 Flash (promo), and $3.20 on 3.1 Pro at September 2026 list prices. Gemini's thinking tokens bill as output, which can push real output bills higher.

Can I use both Gemini and Claude or OpenAI on the same agent?

Yes. HostAgentes supports all four providers with model routing — send long-document and multimodal tasks to Gemini and instruction-heavy reasoning to Claude or GPT-5.6, switching without redeployment.

Deploy a Gemini-powered agent in 5 minutes

Bring your Google AI API key. We handle the rest.

Get Started