[!IMPORTANT] CRITICAL: Unified SDK, Latest Models & GEAP Targeting * Unified SDK: Use the Google Gen AI SDK (google-genai >= 2.3.0for Python,@google/genai >= 2.3.0for JS/TS). SDK versions>= 2.0.0are the minimum floor at which the steps schema activates, but>= 2.3.0is the supported/recommended floor. Legacy SDKs likegoogle-cloud-aiplatform,@google-cloud/vertexai, andgoogle-generativeaiare strictly unsupported for Interactions. * Mandatory Instruction: When explaining client initialization or writing import code, you must explicitly state/warn the user that legacy packages likegoogle-cloud-aiplatformorgoogle-generativeaiare strictly unsupported for interactions. * Latest Models Only: Usegemini-3.5-flash(fast, balanced, multimodal — the recommended default),gemini-3.1-pro-preview(complex reasoning, coding, research), orgemini-3.1-flash-lite(cost-efficient, high-frequency lightweight tasks). Refer to the latest model versions to check for new updates. Legacy models (gemini-3-flash-preview,gemini-2.5-*,gemini-2.0-*,gemini-1.5-*) are deprecated and do not support interactions. * Mandatory Instruction: In any interaction response, you must warn the user that legacy models likegemini-2.5-*,gemini-2.0-*, orgemini-1.5-*are deprecated and unsupported for the Interactions API. * GEAP requires a provisioned agent (no direct base-model calls yet): On Gemini Enterprise Agent Platform (GEAP), direct/base-model calls (model="...") via the Interactions API are not supported yet. You must target a provisioned agent or endpoint with theagent="<AGENT_ID>"parameter instead ofmodel="...". The code examples in this skill useagent=...for this reason. (This is the primary difference from the ai.google.dev documentation for Interactions, which usesmodel=...— whilemodel=...is valid for other Gemini API contexts, it is not supported on the Agent Platform.) Provision an agent per the Agent Platform docs and pass its ID asagent. * Turn-Scoped Parameters: Parameters liketools,system_instruction, andgeneration_configare turn-scoped. They MUST be passed with each interaction request.
gcloud auth application-default login
gcloud services enable aiplatform.googleapis.com
export GOOGLE_GENAI_USE_ENTERPRISE=true
export GOOGLE_CLOUD_PROJECT="your-project-id"
export GOOGLE_CLOUD_LOCATION="global"
from google import genai
# The SDK automatically picks up the environment variables
client = genai.Client()
import { GoogleGenAI } from "@google/genai";
// The SDK automatically picks up the environment variables
const ai = new GoogleGenAI();
from google import genai
import google.auth
_, project_id = google.auth.default()
client = genai.Client(enterprise=True, project=project_id, location="global")
import { GoogleGenAI } from "@google/genai";
const ai = new GoogleGenAI({
enterprise: {
project: "your-project-id",
location: "global"
}
});
steps list.interaction = client.interactions.create(
agent="your-agent-id", # GEAP: target a provisioned agent, not a base model
input="Explain serverless computing in one sentence."
)
# Use the output_text convenience accessor (combined text from the trailing model_output steps)
print(interaction.output_text)
const interaction = await ai.interactions.create({
agent: "your-agent-id", // GEAP: target a provisioned agent, not a base model
input: "Explain serverless computing in one sentence."
});
console.log(interaction.output_text);
previous_interaction_id.# Turn 1: Introduce ourselves
# Interactions are stored by default (store=True); pass store=False to disable
# server-side retention (which also disables previous_interaction_id and background).
turn1 = client.interactions.create(
agent="your-agent-id",
input="Hi! My name is John. I am working on AI agents.",
store=True
)
print(f"Turn 1: {turn1.output_text}")
# Turn 2: Refer back to the stored turn state
turn2 = client.interactions.create(
agent="your-agent-id",
input="What is my name?",
previous_interaction_id=turn1.id
)
print(f"Turn 2: {turn2.output_text}")
// Turn 1 (interactions are stored by default; pass store: false to disable)
const turn1 = await ai.interactions.create({
agent: "your-agent-id",
input: "Hi! My name is John. I am working on AI agents.",
store: true
});
// Turn 2
const turn2 = await ai.interactions.create({
agent: "your-agent-id",
input: "What is my name?",
previousInteractionId: turn1.id
});
console.log(turn2.output_text);
stream=True returns an iterable chunk generator.# The stream yields typed events, not full interaction snapshots. The sequence is:
# interaction.created -> (step.start -> step.delta(s) -> step.stop)+ -> interaction.completed
for event in client.interactions.create(
agent="your-agent-id",
input="Write a short poem about debugging.",
stream=True
):
if event.event_type == "step.delta":
if event.delta.type == "text":
print(event.delta.text, end="", flush=True)
elif event.event_type == "interaction.completed":
print()
// The stream yields typed events, not full interaction snapshots. The sequence is:
// interaction.created -> (step.start -> step.delta(s) -> step.stop)+ -> interaction.completed
const responseStream = await ai.interactions.create({
agent: "your-agent-id",
input: "Write a short poem about debugging.",
stream: true
});
for await (const event of responseStream) {
if (event.event_type === "step.delta") {
if (event.delta.type === "text") {
process.stdout.write(event.delta.text);
}
} else if (event.event_type === "interaction.completed") {
console.log();
}
}
response_format)response_format argument directly takes the target schema structure.from pydantic import BaseModel, Field
class Book(BaseModel):
title: str = Field(description="The title of the book")
author: str = Field(description="The book's author")
year_published: int
interaction = client.interactions.create(
agent="your-agent-id",
input="Recommend one famous sci-fi book.",
response_format=Book
)
# The text will be a valid JSON matching the Book schema
print(interaction.output_text)
import { Type } from "@google/genai";
const BookSchema = {
type: Type.OBJECT,
properties: {
title: { type: Type.STRING, description: "The title of the book" },
author: { type: Type.STRING, description: "The book's author" },
yearPublished: { type: Type.INTEGER }
},
required: ["title", "author", "yearPublished"]
};
const interaction = await ai.interactions.create({
agent: "your-agent-id",
input: "Recommend one famous sci-fi book.",
responseFormat: BookSchema
});
console.log(interaction.output_text);
import json
def get_stock_price(ticker: str) -> float:
"""Gets the stock price for a given ticker symbol."""
if ticker.upper() == "GOOG":
return 175.50
return 100.0
# Turn 1: Pass tools to the model
interaction = client.interactions.create(
agent="your-agent-id",
input="What is the stock price of GOOG?",
tools=[get_stock_price]
)
# In the flat steps schema, a tool request is a top-level step of type
# "function_call" with flat `name` and `arguments` fields (no nested tool_calls).
for step in interaction.steps:
if step.type == "function_call" and step.name == "get_stock_price":
ticker_arg = step.arguments.get("ticker")
price = get_stock_price(ticker_arg)
# Turn 2: Submit the result back as a function_result step. Reference the
# originating call via call_id=step.id, and pass tools again (turn-scoped).
final_turn = client.interactions.create(
agent="your-agent-id",
input=[
{
"type": "function_result",
"name": step.name,
"call_id": step.id,
"result": [{"type": "text", "text": json.dumps(price)}],
}
],
tools=[get_stock_price],
previous_interaction_id=interaction.id
)
print(final_turn.output_text)
import { Type } from "@google/genai";
// Define local tool
function getStockPrice({ ticker }: { ticker: string }): number {
if (ticker.toUpperCase() === "GOOG") {
return 175.50;
}
return 100.00;
}
// Turn 1: Pass tools to the model
const toolDeclaration = {
functionDeclarations: [{
name: "getStockPrice",
description: "Gets the stock price for a given ticker symbol.",
parameters: {
type: Type.OBJECT,
properties: {
ticker: { type: Type.STRING, description: "The stock ticker symbol" }
},
required: ["ticker"]
}
}]
};
const interaction = await ai.interactions.create({
agent: "your-agent-id",
input: "What is the stock price of GOOG?",
tools: [toolDeclaration]
});
// In the flat steps schema, a tool request is a top-level step of type
// "function_call" with flat `name` and `arguments` fields (no nested toolCalls).
const fcStep = interaction.steps.find(s => s.type === "function_call");
if (fcStep && fcStep.name === "getStockPrice") {
const tickerArg = fcStep.arguments.ticker as string;
const price = getStockPrice({ ticker: tickerArg });
// Turn 2: Submit the result back as a function_result step. Reference the
// originating call via call_id=fcStep.id, and pass tools again (turn-scoped).
const finalTurn = await ai.interactions.create({
agent: "your-agent-id",
input: [{
type: "function_result",
name: fcStep.name,
call_id: fcStep.id,
result: [{ type: "text", text: JSON.stringify(price) }]
}],
tools: [toolDeclaration],
previousInteractionId: interaction.id
});
console.log(finalTurn.output_text);
}
curl.POST https://aiplatform.googleapis.com/v1beta1/projects/{PROJECT_ID}/locations/{LOCATION}/interactions
global (or custom region if required).AGENT_ID="your-agent-id"
ACCESS_TOKEN=$(gcloud auth print-access-token)
curl -X POST "https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/global/interactions" \
-H "Authorization: Bearer ${ACCESS_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"agent": "'"${AGENT_ID}"'",
"input": [{
"type": "user_input",
"content": [{
"type": "text",
"text": "Explain serverless computing in one sentence."
}]
}]
}'
{
"id": "your-interaction-id",
"status": "completed",
"steps": [
{
"type": "model_output",
"content": [
{
"type": "text",
"text": "Serverless computing is a cloud execution model where the cloud provider dynamically manages the allocation and provisioning of servers, charging customers based on actual usage rather than pre-purchased capacity."
}
]
}
],
"usage": {
"total_tokens": 24751,
"total_input_tokens": 23894,
"total_output_tokens": 857
},
"created": "2026-05-08T10:44:43Z",
"updated": "2026-05-08T10:44:43Z",
"environment_id": "your-environment-id",
"object": "interaction"
}
previous_interaction_id in the JSON payload:curl -X POST "https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/global/interactions" \
-H "Authorization: Bearer ${ACCESS_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"agent": "'"${AGENT_ID}"'",
"store": true,
"previous_interaction_id": "YOUR_PREVIOUS_INTERACTION_ID",
"input": [{
"type": "user_input",
"content": [{
"type": "text",
"text": "Can you elaborate on that?"
}]
}]
}'
"stream": true in the payload:curl -X POST "https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/global/interactions" \
-H "Authorization: Bearer ${ACCESS_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"agent": "'"${AGENT_ID}"'",
"stream": true,
"input": [{
"type": "user_input",
"content": [{
"type": "text",
"text": "Write a long story about space travel."
}]
}]
}'
data: containing JSON updates with the event_type and step contents.Howcurlhandles streaming: By default, when"stream": trueis passed, the server responds withTransfer-Encoding: chunkedandContent-Type: text/event-stream(Server-Sent Events).curlwill automatically keep the connection open and print the incoming data chunks tostdoutin real time as they are pushed by the server. The user does not need to poll or pull further; the complete sequence of events streams continuously until completion.
Interaction response contains steps, an array of typed step objects
representing a structured timeline of the interaction turn. Read the current
step type rather than assuming the last step is text — the trailing step may
be a function_call or a thought.user_input: User input (text, audio, multimodal). Contains a content
array. (This is why REST input payloads use "type": "user_input", not
"role": "user".)model_output: Final model generation. Contains a content array with
text, image, audio, etc. (REST responses use "type": "model_output",
not "role": "model".)thought: Model reasoning / chain of thought. Has a signature field and
optional summary.function_call: Tool call request, with flat id, name, and arguments
fields (there is no nested tool_calls list).function_result: Tool result you send back, with call_id, name, and
result fields.google_search_call / google_search_result, code_execution_call /
code_execution_result, url_context_call / url_context_result,
mcp_server_tool_call / mcp_server_tool_result, file_search_call /
file_search_result: built-in and remote tool steps.content array on model_output and user_input steps)text: Text content (text field).image / audio / document / video: Content with data, mime_type,
or uri.output_text: The combined text from the trailing model_output steps.
Prefer this over hand-walking steps[-1].content[0].text, which breaks when
the last step is a tool call or a thought.| Event | Description |
|---|---|
interaction.created | Interaction created; includes metadata. |
step.start | A new step begins. Contains the step type and |
| : : initial metadata. : | |
step.delta | Incremental data for the current step. Contains a |
: : typed delta object (e.g. delta.type == "text" : | |
: : with delta.text). : | |
step.stop | The step is complete. Contains index. |
interaction.completed | Interaction finished. Contains final usage. |
store=True), which enables stateful
features like previous_interaction_id and background execution. Passing
store=False disables server-side retention and therefore also disables
previous_interaction_id and background — in that mode you must pass the full
conversation history in input on each turn.Start with one job and grow from there.