31

models

Google logo

Google Models

Google offers the broadest multimodal AI portfolio — language, image, and video from one provider. Gemini 2.5 Pro and Flash deliver frontier reasoning with massive context windows. Imagen 4 produces photorealistic images with accurate text rendering. The Veo family (Veo 2, 3, 3.1) generates cinematic video with realistic motion and natural audio. Access Gemini, Imagen, and Veo via Segmind APIs — no Google Cloud credentials needed. Chain them in Segmind Workflows for end-to-end content pipelines from strategy to video, fully automated.

Google
Text To Video

Gemini Omni 1.1

51.0s
Video To Video

Gemini Omni 1.1 Video Extend

126.4s
Video To Video

Gemini Omni 1.1 Video Edit

98.7s
Text To Audio

Lyria 3 Pro

34.4s
Text To Audio

Lyria 3

16.1s
LLM

Gemini 3.7 Flash

6.3s
Image To Image

Nano Banana 2 Lite

11.3s
Text To Video

Gemini Omni Flash

46.4s
Text To Audio

Gemini 3.1 Flash TTS

13.8s
Text To Embed

Gemini Embedding 2

1.1s
Text To Embed

Gemini Embedding 001

0.9s
LLM

Gemini 2.5 Flash Lite

1.9s
LLM

Gemini 3.1 Flash Lite

3.2s
LLM

Gemini 3 Flash

13.4s
LLM

Gemini 3.1 Pro

24.4s
Image To Image

Nano Banana 2

28.5s
Text To Audio

Gemini TTS 2.5 Flash

7.9s
Text To Audio

Gemini TTS 2.5 Pro

14.1s
LLM

Gemini 3 Pro

27.1s
Image To Image

Nano Banana Pro

41.3s
Image To Video

Veo 3.1 Fast

108.9s
Image To Video

Veo 3.1

107.7s
LLM

Gemini 2.5 Flash

10.6s
LLM

Gemini 2.5 PRO

3.2s
Text To Image

Nano Banana

7.6s
Text To Video

Veo 3 Fast

58.8s
Text To Video

Google Veo 3

84.8s
Text To Audio

Lyria 2

33.5s
Image To Text

Google Translate

0.8s
Image To Video

Google Veo 2 Image To Video

41.6s
Text To Video

Google Veo 2

48.9s