Text to Image
Generate images from text prompts with multiple AI models
Generate images from text descriptions using various AI models.
How It Works
- Write a text prompt describing your image
- Select AI model(s) from the model selector
- Choose settings (aspect ratio, quantity, etc.)
- Click Generate
Settings
- Aspect Ratio: Choose image dimensions (1:1, 16:9, 9:16, etc.)
- Resolution: Tier selector on supported models —
1K(~1024px),2K(Full HD), or4K(UHD, SeeDream 4.x only). Availability is model-specific; some models are flat across tiers (Flux 2, Nano Banana, Imagen, Qwen) and some charge more at higher tiers (e.g. SeeDream 5.0 at4Kcosts 2× the base). - Quantity: Number of images to generate (1-4)
- Model Selection: Choose one or multiple AI models
- Stylization: Control creative interpretation (model-specific)
- Reference Images: Upload images to guide generation — see "Reference Images vs Source Images" below for what this actually does
Available AI Models
- FLUX.2: High-quality, versatile image generation
- GPT Image 2: OpenAI text rendering, infographics, photorealism
- Qwen Image 2 (Standard / Pro): Strong prompt following and text
- Nano Banana 2 / Nano Banana 2 Pro: Google multi-reference edits with web-grounded text
- Recraft 4: Logos, vector-style, brand design
- Grok Imagine Quality: xAI high-fidelity image generation
- SeeDream 5.0 Lite / Seedream 4.5: ByteDance high-resolution generation (up to 4K)
- Higgsfield Soul: Stylized, aesthetic image generation
- Midjourney V8 + Niji V7: Artistic and anime-style generation
See the full Models catalog for every available model.
Features
Multi-Model Generation
- Generate with multiple models simultaneously
- Compare results from different AI models
- Each model has unique strengths and styles
Moodboards
- Create collections of reference images
- Apply moodboard style to generations
- Manage and reuse moodboards across projects
Image Presets
- Quick style templates
- Save time with pre-configured settings
- Apply consistent styles across generations
Reference Images vs Source Images — Read This Before You Upload
This is the single most common confusion in Kolbo. The two concepts look similar but produce completely different results.
Reference Images (this page — Text to Image)
Reference Images guide style, mood, and composition. The model studies them and generates a brand-new image inspired by them. It does not copy their pixels.
Use Reference Images when you want to:
- Match an aesthetic ("make it feel like this brand's other ads")
- Borrow a color palette or lighting style
- Imitate a composition or pose
Do NOT use Reference Images when you want to:
- Embed a real logo, icon, or watermark into the output — the model will redraw a similar-looking shape, not your actual logo
- Keep a specific person's face exactly the same — use Visual DNA, or switch to Image Editing with the face photo as a Source Image
- Composite multiple uploaded assets together — that's an Image Editing workflow
Source Images (Image Editing tool)
Source Images get pixel-accurately composited into the output. Use Image Editing when you need a specific uploaded asset to actually appear in the final image. See the Image Editing docs for the full workflow.
Rule of thumb: If the user looking at the output should be able to see the exact pixels of an asset you uploaded, you need Image Editing, not Text to Image.
Tips
- Be Specific: Detailed prompts get better results
- Try Different Models: Each model interprets prompts differently
- Reference vs Source: Reference Images inspire; Source Images embed. See above.
- Experiment: Test different settings and prompts