# Grok Imagine Image > Generate and edit photorealistic images from text with Grok Imagine Image, up to sharp 2K detail and reliable in-image text. ## Overview - **Endpoint**: `https://api.segmind.com/v1/grok-imagine-image` - **Model ID**: `grok-imagine-image` - **Category**: Image-to-Image Transformation - **Type**: Synchronous (Direct response) - **Average Latency**: ~19.4s (30-day average) - **Average Cost**: $0.1127 per run (observed across past runs, not a price — see Pricing) - **Provider**: XAI (Grok) ## Pricing | mode | resolution | cost | | --- | --- | --- | | standard | 1k | 0.027 | | standard | 2k | 0.027 | | quality | 1k | 0.0675 | | quality | 2k | 0.0945 | Edit requests add $0.0027 (Standard) or $0.0135 (Quality) per input image. ## API Information This model uses a **synchronous response pattern**: 1. Make a POST request with your parameters 2. Receive the output directly in the response (binary for images/videos/audio, JSON for text) 3. No polling required - response is immediate ### Input Schema The API accepts the following input parameters: - **`prompt`** (`string`, _required_): Prompt Describe the image to generate or the edit to apply. Be specific about subject, lighting, camera, and style. - Default: `"A majestic snow leopard perched on a rocky cliff at dawn, golden sunlight catching its thick fur, breath visible in the cold air, sweeping Himalayan mountain range in the background, ultra-detailed wildlife photography, sharp focus, dramatic natural lighting"` - **`mode`** (`string`, _optional_): Mode Standard is fast and economical; quality maximizes fidelity and text rendering. Also affects price. - Default: `"standard"` - Options: "standard" (Standard), "quality" (Quality) - **`image_urls`** (`array`, _optional_): Reference images Optional source images to edit. Up to 3 — reference multiple as , , in the prompt. Providing any image switches from generation to editing (and edit pricing). Omit for pure text-to-image. - Default: `[]` - Item type: string - **`aspect_ratio`** (`string`, _optional_): Aspect ratio Output aspect ratio. Ignored when an input image sets the size. - Default: `"16:9"` - Options: "1:1" (Square (1:1)), "3:4" (Portrait (3:4)), "4:3" (Landscape (4:3)), "9:16" (Vertical (9:16)), "16:9" (Widescreen (16:9)), "2:3" (Portrait (2:3)), "3:2" (Landscape (3:2)), "9:19.5" (Tall (9:19.5)), "19.5:9" (Wide (19.5:9)), "9:20" (Very tall (9:20)), "20:9" (Very wide (20:9)), "1:2" (Tall (1:2)), "2:1" (Wide (2:1)), "auto" (Auto) - **`resolution`** (`string`, _optional_): Resolution Output resolution. 1k for fast drafts, 2k for detailed final renders. - Default: `"2k"` - Options: "1k" (1K), "2k" (2K) - **`n`** (`integer`, _optional_): Number of images Images to generate per request, 1 to 4. Price scales with the number of images returned. - Default: `1` - Range: 1 to 4 - **`output_format`** (`string`, _optional_): Output format Output file format. - Default: `"png"` - Options: "jpeg" (JPEG), "png" (PNG), "webp" (WEBP) **Required Parameters Example**: ```json { "prompt": "A majestic snow leopard perched on a rocky cliff at dawn, golden sunlight catching its thick fur, breath visible in the cold air, sweeping Himalayan mountain range in the background, ultra-detailed wildlife photography, sharp focus, dramatic natural lighting" } ``` **Full Example**: ```json { "prompt": "A majestic snow leopard perched on a rocky cliff at dawn, golden sunlight catching its thick fur, breath visible in the cold air, sweeping Himalayan mountain range in the background, ultra-detailed wildlife photography, sharp focus, dramatic natural lighting", "mode": "standard", "image_urls": [], "aspect_ratio": "16:9", "resolution": "2k", "n": 1, "output_format": "png" } ``` ### Output Schema The API returns a synchronous response based on the model type: **For Image/Video/Audio Models**: - Response contains binary data (image/png, video/mp4, audio/mp3) - Content-Type header indicates the media type - Save the response body directly to a file **For Text Models**: - Response is JSON with the generated text - Structure varies by model **HTTP Response Codes**: - **200 - OK**: Request successful, output in response body - **400 - Bad Request**: Invalid parameters - **401 - Unauthorized**: Invalid or missing API key - **404 - Not Found**: Model not found - **406 - Not Acceptable**: Insufficient credits - **429 - Too Many Requests**: Rate limit exceeded - **500 - Server Error**: Internal server error ## About ### Grok Imagine Image — Text-to-Image and Image Editing Model #### What is Grok Imagine Image? Grok Imagine Image is xAI's text-to-image and image-editing model, part of the Grok Imagine family powered by the Aurora engine. Describe a scene in plain language to generate a new image, or supply a source image and describe the change to run image-to-image edits — both workflows live in one model. A `standard` mode delivers fast, economical generations for rapid iteration, while a `quality` mode targets maximum fidelity, sharper detail, more natural lighting, and stronger prompt adherence. Outputs scale up to 2K resolution across a wide set of aspect ratios, making the model a practical default for everything from quick social concepts to polished hero visuals. #### Key Features - Text-to-image generation and natural-language image editing in a single model - `standard` and `quality` modes to trade speed for fidelity - Output up to 2K resolution with aspect ratios from 1:1 and 16:9 to 9:16 and 2:1 - Reliable in-image text rendering, including multilingual scripts - Batch up to 4 images per request to compare prompt variations - Choice of jpeg, png, or webp output #### Best Use Cases Grok Imagine Image is built for fast ideation and creative experimentation. It shines for social media graphics, marketing concepts, product mockups, and posters where readable brand names, slogans, or signage need to sit inside the frame. Photographers and designers use the editing mode to restyle a photo, swap backgrounds, or add objects with a single instruction. Concept artists rely on the model to explore characters, environments, and moodboards quickly, then switch to `quality` mode for final, presentation-ready renders. #### Prompt Tips and Output Quality The model rewards natural-language scene descriptions over keyword stacks. Lead with the subject, keep prompts roughly 30 to 80 words, and describe light behavior, camera language, and film stock instead of vague adjectives like "8K" or "stunning." It does not use negative prompts, so phrase constraints positively (for example, "sharp focus, clean composition"). To place text, spell the exact wording in quotes; `quality` mode renders typography most reliably. Generate a small batch first, then iterate one element at a time. ## Usage Guide ### How to Use Grok Imagine Image Grok Imagine Image handles two workflows in one model: text-to-image generation and image-to-image editing. To generate, send a descriptive `prompt`. To edit, add a source `image` (URL or base64) and describe the change you want — the model edits that image instead of starting from scratch. #### Choosing the right parameters - **mode** — Use `standard` for fast, low-effort iteration and concept exploration. Switch to `quality` for final renders that need maximum fidelity, natural lighting, and reliable text rendering. - **resolution** — Use `1k` while you dial in the prompt, then `2k` for detailed final assets such as posters, ads, and hero images. - **aspect_ratio** — Pick `16:9` for landscape and headers, `9:16` for stories and reels, `1:1` for square social posts, and `2:1` or `3:2` for banners and editorial layouts. The ratio is ignored when an input image dictates the size. - **n** — Generate 2 to 4 images when testing a new prompt so you can compare compositions, then drop to 1 once the prompt is consistent. - **output_format** — Choose `jpeg` for small files, `png` for lossless quality, or `webp` for a balance of size and quality. #### Prompt structure that works Write natural sentences, not tag lists. Lead with the subject, then add environment, lighting behavior, camera language, and mood — roughly 30 to 80 words. Name a film stock or lens ("Kodak Portra 400," "85mm f/1.2") for a concrete look. Avoid negative prompts; the model responds to positive constraints like "sharp focus, minimal grain." To render text, put the exact words in quotes and specify the surface, for example a sign or label. #### Editing tips When editing, anchor your instruction to a specific region ("modify the wooden sign above the door") and protect what should stay ("keep the facial features exactly as they are"). Change one element at a time and regenerate to compare results. ## FAQ ### Does Grok Imagine Image support image editing? Yes. Provide a source image as a URL or base64 and describe the change; the model edits instead of generating from scratch. ### What is the difference between standard and quality mode? `standard` is fast and economical for iteration. `quality` uses the premium model for higher fidelity, sharper detail, and stronger text rendering. ### Can it render text inside images? Yes. Put the exact words in quotes in the prompt. Text rendering is strongest in `quality` mode and supports multiple languages. ### What resolutions and aspect ratios are supported? Resolution is 1k or 2k, with aspect ratios spanning 1:1, 16:9, 9:16, 4:3, 3:2, 2:1, and more. ### How many images can I generate at once? Set `n` from 1 to 4 to produce multiple variations in a single request. ### Is Grok Imagine Image good for brand work? It is strong for fast concepts, mockups, and social assets; review brand-sensitive output, as commercial safety is lower than some rivals. ## Usage Examples ### cURL ```bash curl -X POST "https://api.segmind.com/v1/grok-imagine-image" \ -H "x-api-key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "A majestic snow leopard perched on a rocky cliff at dawn, golden sunlight catching its thick fur, breath visible in the cold air, sweeping Himalayan mountain range in the background, ultra-detailed wildlife photography, sharp focus, dramatic natural lighting", "mode": "standard", "image_urls": [], "aspect_ratio": "16:9", "resolution": "2k", "n": 1, "output_format": "png" }' ``` ### Python ```python import requests import json api_key = "YOUR_API_KEY" url = "https://api.segmind.com/v1/grok-imagine-image" data = { "prompt": "A majestic snow leopard perched on a rocky cliff at dawn, golden sunlight catching its thick fur, breath visible in the cold air, sweeping Himalayan mountain range in the background, ultra-detailed wildlife photography, sharp focus, dramatic natural lighting", "mode": "standard", "image_urls": [], "aspect_ratio": "16:9", "resolution": "2k", "n": 1, "output_format": "png" } response = requests.post( url, json=data, headers={ 'x-api-key': api_key, 'Content-Type': 'application/json' } ) if response.status_code == 200: # For image/video/audio models, response.content contains the binary data with open('output.png', 'wb') as f: f.write(response.content) print('Generation complete, saved to output.png') else: print(f"Error: {response.status_code}") print(response.text) ``` ### JavaScript ```javascript const apiKey = 'YOUR_API_KEY'; const url = 'https://api.segmind.com/v1/grok-imagine-image'; const data = { "prompt": "A majestic snow leopard perched on a rocky cliff at dawn, golden sunlight catching its thick fur, breath visible in the cold air, sweeping Himalayan mountain range in the background, ultra-detailed wildlife photography, sharp focus, dramatic natural lighting", "mode": "standard", "image_urls": [], "aspect_ratio": "16:9", "resolution": "2k", "n": 1, "output_format": "png" }; const response = await fetch(url, { method: 'POST', headers: { 'x-api-key': apiKey, 'Content-Type': 'application/json', }, body: JSON.stringify(data), }); if (response.ok) { // For image/video/audio models, response contains binary data const blob = await response.blob(); const downloadUrl = URL.createObjectURL(blob); // Create download link const a = document.createElement('a'); a.href = downloadUrl; a.download = 'output.png'; a.click(); console.log('Generation complete'); } ``` ## Additional Resources ### Documentation - [Model Playground](https://www.segmind.com/models/grok-imagine-image) - [API Documentation](https://www.segmind.com/models/grok-imagine-image/api) - [Pricing Details](https://www.segmind.com/models/grok-imagine-image/pricing) - [Platform Documentation](https://docs.segmind.com/)