Grok Imagine Image 2

Text-to-image and image editing with crisp, legible text.

Example outputDefault output example

Grok Imagine Image 2 — Text-to-Image and Image Editing Model

What is Grok Imagine Image 2?

Grok Imagine Image 2 is xAI's second-generation image model for text-to-image generation and instruction-based image editing. Built around the goal of making images you can use in real work, it follows prompts closely, plans typography and layout the way a designer would, and preserves the elements you supply across generations and edits. It powers the Quality Mode experience in Grok and is available on Segmind through a single synchronous endpoint that returns your image directly — no polling.

The model was tuned for fidelity across photography, design, and illustration, with editing treated as a first-class capability rather than an add-on. At launch it ranked second in the world on both the text-to-image and image-editing Arena leaderboards, a clear jump over xAI's previous image model.

Key Features

  • Designer-grade text and layout. Plans type hierarchy before rendering, so dense visuals like posters, infographics, and tutorial sheets hold together and small text stays sharp.
  • Instruction-based editing. Provide up to 3 reference images and describe the change; the model edits what you mean while preserving supplied subjects.
  • Multi-reference composition. Combine a subject, a style, and a scene across reference images in one generation.
  • Flexible output. Choose 1k or 2k resolution, low or medium quality, up to 4 images per request, and jpeg, png, or webp formats.
  • Wide aspect-ratio range. From 1:1 square to 16:9 widescreen, 9:16 vertical, and tall or wide banners.

Best Use Cases

Grok Imagine Image 2 shines on production-oriented visuals: marketing posters, e-commerce product shots, editorial graphics, infographics, menus, packaging concepts, app icons, game assets, and professional headshots. Because editing is first-class, it fits iterative workflows where you generate a hero image, then refine a region, swap a color, or lift a subject onto a clean background without touching the rest of the frame. Multi-reference input makes it a strong pick for consistent characters, locations, and props across a visual set.

Prompt Tips and Output Quality

Write prompts like a design brief: name the subject, the layout, the exact on-image words in quotes, the style, and the lighting, in that order. Put text you want rendered inside quotation marks and say where it sits. Use medium quality and 2k resolution for final assets, and 1k for fast drafts. For edits, describe one scoped change at a time and name what each reference image contributes.

FAQs

Is Grok Imagine Image 2 good at rendering text? Yes. Text is one of its strongest suits — it plans typography and layout so posters, infographics, and small labels come out legible when you spell words in quotes.

Can it edit an existing image? Yes. Supply source images and describe the change in natural language; the model applies scoped edits while preserving what you provide. It accepts up to 3 reference images per request.

How does it compare to GPT Image 2? On the August 2026 Arena leaderboards, Grok Imagine Image 2 ranks second in both text-to-image and image editing, behind OpenAI's gpt-image-2.

What resolutions and formats are supported? Output is available at 1k or 2k resolution in jpeg, png, or webp, with low or medium quality tiers and up to 4 images per request.

Does it support multiple reference images? Yes. You can combine several reference images in one generation to control subject, style, and scene simultaneously.

Is it a video model? No. Grok Imagine Image 2 is an image generation and editing model, distinct from xAI's Grok Imagine video models.