# Grok Imagine Image 2 > Generate and edit images with designer-grade text rendering and precise region edits at up to 2K resolution. ## Overview - **Endpoint**: `https://api.segmind.com/v1/grok-imagine-image-2` - **Model ID**: `grok-imagine-image-2` - **Category**: Image-to-Image Transformation - **Type**: Synchronous (Direct response) - **Average Latency**: ~22.0s (30-day average) - **Average Cost**: $0.1144 per run (observed across past runs, not a price — see Pricing) - **Provider**: XAI (Grok) ## Pricing | quality | resolution | cost | | --- | --- | --- | | low | 1k | 0.054 | | low | 2k | 0.081 | | medium | 1k | 0.081 | | medium | 2k | 0.108 | Edit requests add $0.0135 per input image. ## API Information This model uses a **synchronous response pattern**: 1. Make a POST request with your parameters 2. Receive the output directly in the response (binary for images/videos/audio, JSON for text) 3. No polling required - response is immediate ### Input Schema The API accepts the following input parameters: - **`prompt`** (`string`, _required_): Prompt Text prompt for the image or edit to apply. Write like a design brief with exact quoted text. - Default: `"A cozy ramen shop on a narrow Tokyo alley at night in the rain, glowing red and blue neon signs reflecting on the wet pavement, steam rising from a fresh bowl of ramen on a wooden counter, warm lantern light spilling from the doorway, a lone customer seated inside, cinematic street photography, ultra-detailed, sharp focus, shallow depth of field, 35mm"` - **`quality`** (`string`, _optional_): Quality Rendering effort: low is faster, medium maximizes fidelity and in-image text. Use medium for final renders. - Default: `"medium"` - Options: "low" (low), "medium" (medium) - **`image_urls`** (`array`, _optional_): Input Image List Optional source images to edit, up to 3 (extras ignored); switches to editing mode. Omit for text-to-image. Up to 3 reference images, sent as a list. - Item type: string - **`aspect_ratio`** (`string`, _optional_): Aspect Ratio Output aspect ratio, ignored when an input image sets size. Use 16:9 landscape, 9:16 social, 1:1 square. - Default: `"16:9"` - Options: "1:1" (1:1), "3:4" (3:4), "4:3" (4:3), "9:16" (9:16), "16:9" (16:9), "2:3" (2:3), "3:2" (3:2), "9:19.5" (9:19.5), "19.5:9" (19.5:9), "9:20" (9:20), "20:9" (20:9), "1:2" (1:2), "2:1" (2:1), "auto" (auto) - **`resolution`** (`string`, _optional_): Resolution Output resolution, 1k or 2k. Use 1k for fast drafts, 2k for detailed final renders. - Default: `"1k"` - Options: "1k" (1k), "2k" (2k) - **`n`** (`integer`, _optional_): Number of Images Number of images per request, 1 to 4. Use 2-3 to compare variations, 1 when dialed in. - Default: `1` - Range: 1 to 4 - **`output_format`** (`string`, _optional_): Output Format Output file format. Use jpeg for small files, png for lossless, webp for a balance. - Default: `"png"` - Options: "jpeg" (jpeg), "png" (png), "webp" (webp) **Required Parameters Example**: ```json { "prompt": "A cozy ramen shop on a narrow Tokyo alley at night in the rain, glowing red and blue neon signs reflecting on the wet pavement, steam rising from a fresh bowl of ramen on a wooden counter, warm lantern light spilling from the doorway, a lone customer seated inside, cinematic street photography, ultra-detailed, sharp focus, shallow depth of field, 35mm" } ``` **Full Example**: ```json { "prompt": "A cozy ramen shop on a narrow Tokyo alley at night in the rain, glowing red and blue neon signs reflecting on the wet pavement, steam rising from a fresh bowl of ramen on a wooden counter, warm lantern light spilling from the doorway, a lone customer seated inside, cinematic street photography, ultra-detailed, sharp focus, shallow depth of field, 35mm", "quality": "medium", "image_urls": [], "aspect_ratio": "16:9", "resolution": "1k", "n": 1, "output_format": "png" } ``` ### Output Schema The API returns a synchronous response based on the model type: **For Image/Video/Audio Models**: - Response contains binary data (image/png, video/mp4, audio/mp3) - Content-Type header indicates the media type - Save the response body directly to a file **For Text Models**: - Response is JSON with the generated text - Structure varies by model **HTTP Response Codes**: - **200 - OK**: Request successful, output in response body - **400 - Bad Request**: Invalid parameters - **401 - Unauthorized**: Invalid or missing API key - **404 - Not Found**: Model not found - **406 - Not Acceptable**: Insufficient credits - **429 - Too Many Requests**: Rate limit exceeded - **500 - Server Error**: Internal server error ## About ### Grok Imagine Image 2 — Text-to-Image and Image Editing Model #### What is Grok Imagine Image 2? Grok Imagine Image 2 is xAI's second-generation image model for text-to-image generation and instruction-based image editing. Built around the goal of making images you can use in real work, it follows prompts closely, plans typography and layout the way a designer would, and preserves the elements you supply across generations and edits. It powers the Quality Mode experience in Grok and is available on Segmind through a single synchronous endpoint that returns your image directly — no polling. The model was tuned for fidelity across photography, design, and illustration, with editing treated as a first-class capability rather than an add-on. At launch it ranked second in the world on both the text-to-image and image-editing Arena leaderboards, a clear jump over xAI's previous image model. #### Key Features - **Designer-grade text and layout.** Plans type hierarchy before rendering, so dense visuals like posters, infographics, and tutorial sheets hold together and small text stays sharp. - **Instruction-based editing.** Provide up to 3 reference images and describe the change; the model edits what you mean while preserving supplied subjects. - **Multi-reference composition.** Combine a subject, a style, and a scene across reference images in one generation. - **Flexible output.** Choose 1k or 2k resolution, low or medium quality, up to 4 images per request, and jpeg, png, or webp formats. - **Wide aspect-ratio range.** From 1:1 square to 16:9 widescreen, 9:16 vertical, and tall or wide banners. #### Best Use Cases Grok Imagine Image 2 shines on production-oriented visuals: marketing posters, e-commerce product shots, editorial graphics, infographics, menus, packaging concepts, app icons, game assets, and professional headshots. Because editing is first-class, it fits iterative workflows where you generate a hero image, then refine a region, swap a color, or lift a subject onto a clean background without touching the rest of the frame. Multi-reference input makes it a strong pick for consistent characters, locations, and props across a visual set. #### Prompt Tips and Output Quality Write prompts like a design brief: name the subject, the layout, the exact on-image words in quotes, the style, and the lighting, in that order. Put text you want rendered inside quotation marks and say where it sits. Use medium quality and 2k resolution for final assets, and 1k for fast drafts. For edits, describe one scoped change at a time and name what each reference image contributes. ## Usage Guide ### How to Use Grok Imagine Image 2 Grok Imagine Image 2 handles both text-to-image generation and instruction-based editing from one synchronous endpoint. Send a prompt and you receive the image directly. The keys to great output are a well-structured prompt and the right quality, resolution, and reference settings for the job. #### Writing the prompt Treat the prompt like a design brief rather than a caption. Name the subject, the layout, the exact on-image words in quotes, the style, and the lighting, in that order. Because the model plans typography and layout before rendering, spelling words in quotation marks and stating where they sit produces sharp, legible text — ideal for posters, infographics, and packaging. #### Generation vs. editing Leave `image_urls` empty for pure text-to-image. To edit, pass one or more source images (up to 3; extras are ignored) and describe only the change you want. Providing any image switches the request into editing mode. For multi-reference work, name what each image contributes — one for the subject, one for the style, one for the scene — and reference them in the prompt. #### Choosing quality and resolution Use `quality: low` and `resolution: 1k` for fast drafts and idea exploration, then switch to `quality: medium` and `resolution: 2k` for final assets where fidelity and in-image text matter most. Set `aspect_ratio` to match the destination — 16:9 for widescreen, 9:16 for social/stories, 1:1 for square. Note that aspect ratio is ignored when an input image already sets the size. #### Batching and formats Set `n` to 2–4 to compare prompt variations in a single call, then drop back to 1 once the direction is locked. Choose `output_format` by need: png for lossless, webp for a smaller balanced file, jpeg for the smallest size. #### Recommended Settings - **Posters and infographics:** medium quality, 2k, aspect ratio matched to layout, exact text in quotes. - **Fast drafts and iteration:** low quality, 1k, n = 3 to compare options. - **Photo edits:** supply 1 source image, medium quality, one scoped change per request. - **Multi-reference composition:** up to 3 images, name each reference's role in the prompt. ## FAQ ### Is Grok Imagine Image 2 good at rendering text? Yes. Text is one of its strongest suits — it plans typography and layout so posters, infographics, and small labels come out legible when you spell words in quotes. ### Can it edit an existing image? Yes. Supply source images and describe the change in natural language; the model applies scoped edits while preserving what you provide. It accepts up to 3 reference images per request. ### How does it compare to GPT Image 2? On the August 2026 Arena leaderboards, Grok Imagine Image 2 ranks second in both text-to-image and image editing, behind OpenAI's gpt-image-2. ### What resolutions and formats are supported? Output is available at 1k or 2k resolution in jpeg, png, or webp, with low or medium quality tiers and up to 4 images per request. ### Does it support multiple reference images? Yes. You can combine several reference images in one generation to control subject, style, and scene simultaneously. ### Is it a video model? No. Grok Imagine Image 2 is an image generation and editing model, distinct from xAI's Grok Imagine video models. ## Usage Examples ### cURL ```bash curl -X POST "https://api.segmind.com/v1/grok-imagine-image-2" \ -H "x-api-key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "A cozy ramen shop on a narrow Tokyo alley at night in the rain, glowing red and blue neon signs reflecting on the wet pavement, steam rising from a fresh bowl of ramen on a wooden counter, warm lantern light spilling from the doorway, a lone customer seated inside, cinematic street photography, ultra-detailed, sharp focus, shallow depth of field, 35mm", "quality": "medium", "image_urls": [], "aspect_ratio": "16:9", "resolution": "1k", "n": 1, "output_format": "png" }' ``` ### Python ```python import requests import json api_key = "YOUR_API_KEY" url = "https://api.segmind.com/v1/grok-imagine-image-2" data = { "prompt": "A cozy ramen shop on a narrow Tokyo alley at night in the rain, glowing red and blue neon signs reflecting on the wet pavement, steam rising from a fresh bowl of ramen on a wooden counter, warm lantern light spilling from the doorway, a lone customer seated inside, cinematic street photography, ultra-detailed, sharp focus, shallow depth of field, 35mm", "quality": "medium", "image_urls": [], "aspect_ratio": "16:9", "resolution": "1k", "n": 1, "output_format": "png" } response = requests.post( url, json=data, headers={ 'x-api-key': api_key, 'Content-Type': 'application/json' } ) if response.status_code == 200: # For image/video/audio models, response.content contains the binary data with open('output.png', 'wb') as f: f.write(response.content) print('Generation complete, saved to output.png') else: print(f"Error: {response.status_code}") print(response.text) ``` ### JavaScript ```javascript const apiKey = 'YOUR_API_KEY'; const url = 'https://api.segmind.com/v1/grok-imagine-image-2'; const data = { "prompt": "A cozy ramen shop on a narrow Tokyo alley at night in the rain, glowing red and blue neon signs reflecting on the wet pavement, steam rising from a fresh bowl of ramen on a wooden counter, warm lantern light spilling from the doorway, a lone customer seated inside, cinematic street photography, ultra-detailed, sharp focus, shallow depth of field, 35mm", "quality": "medium", "image_urls": [], "aspect_ratio": "16:9", "resolution": "1k", "n": 1, "output_format": "png" }; const response = await fetch(url, { method: 'POST', headers: { 'x-api-key': apiKey, 'Content-Type': 'application/json', }, body: JSON.stringify(data), }); if (response.ok) { // For image/video/audio models, response contains binary data const blob = await response.blob(); const downloadUrl = URL.createObjectURL(blob); // Create download link const a = document.createElement('a'); a.href = downloadUrl; a.download = 'output.png'; a.click(); console.log('Generation complete'); } ``` ## Additional Resources ### Documentation - [Model Playground](https://www.segmind.com/models/grok-imagine-image-2) - [API Documentation](https://www.segmind.com/models/grok-imagine-image-2/api) - [Pricing Details](https://www.segmind.com/models/grok-imagine-image-2/pricing) - [Platform Documentation](https://docs.segmind.com/)