AI product

Gemini Interactions API

The Gemini Interactions API is Google's unified API for calling Gemini models and specialized agents. It maintains server-side interaction state through interaction IDs and the `previous_interaction_id` parameter, while also supporting stateless requests. Each Interaction resource contains a chronological sequence of execution steps, including model thoughts, tool calls and results, and final output. The API supports text and multimodal generation, structured and strongly typed outputs, built-in and custom tools, tool orchestration, observable execution steps, background execution for long-running tasks, and chained image, video, and audio generation using shared context. Google documents it as generally available and recommends it for new Gemini API projects.

Visit site Mentioned in 2 videos ↓

What Gemini Interactions API is used for

3 uses taken from transcripts — each links to the moment in the video.

  • Maintains server-side interaction state through interaction IDs, supports multimodal generation, typed outputs, tool use, and chained model calls.

  • An API for working with both Gemini models and agents through a unified interface. It provides server-side state with interaction IDs, strongly typed multimodal outputs, support for built-in and custom tools, and a steps data model for complex agent workflows.

  • Google's new API for both Gemini models and agents. It offers server-side state management, long-running background operations, built-in and custom tools across text, audio and images, and replaces user/model turns with a timeline of steps.

Videos mentioning Gemini Interactions API

2 in the library.