Kolbo.AIKolbo.AI Docs

Image Editing

Edit, transform, and composite images with pixel-accurate AI workflows

Transform existing images and pixel-accurately composite uploaded assets into new generations.

Overview

Image Editing operates on Source Images — one or more images whose actual pixel content gets composited into the output. Unlike Text to Image (which uses Reference Images for style inspiration only), Image Editing preserves and embeds the pixels of the assets you upload.

This makes it the right tool for any workflow where a specific uploaded asset — a logo, an icon, a face, a product shot — needs to actually appear in the final image, not just inspire its look.

Three Modes

Image Editing auto-detects what you're trying to do from the number and shape of the Source Images:

1. Single-Image Edit

One Source Image → edit / transform that image.

  • Change the sky to sunset
  • Remove the background
  • Swap an outfit color
  • Restyle to a new artistic look

The model preserves the rest of the image and modifies only what your prompt describes.

2. Composite Into a Base

Multiple Source Images, one of them clearly a base scene → composite the others into the base.

  • Put the person from image 1 into the background from image 2
  • Drop a product shot onto a different surface
  • Add a character into an existing scene

3. Generate a New Scene With Embedded Assets

Multiple Source Images with no obvious base → generate a brand-new scene that pixel-accurately embeds your supplied images at positions described in the prompt.

This is the canonical thumbnail / branded composition workflow. Upload a host's face photo plus 2–3 brand logos as Source Images, describe the scene and where each asset goes by ordinal position in the prompt ("the FIRST source image in the center as the hero, the SECOND source image in the bottom-left as the brand logo"), and the model generates the composition with the actual pixels of your uploads.

How It Works

  1. Upload Source Images: One or more files for the model to work with
  2. Describe the Change or Composition: Tell the AI what you want — referring to Source Images by ordinal position ("FIRST", "SECOND") when there are multiple
  3. Optionally Lock Pixels: Add "composite AS-IS, do not redraw or restyle" to the prompt when you need exact preservation (e.g., logos)
  4. AI Processing: AI composites and renders the output
  5. Refine: Iterate on the prompt or swap Source Images and regenerate

Resolution

Image editing models default to 1K. Some models support 2K and 4K tiers — the picker appears on supported models only. Most Flux / Nano Banana / Qwen edit models are flat across tiers; SeeDream 4.x supports up to 4K.

Source Images vs Reference Images — Critical Distinction

Source Images (this tool)Reference Images (Text to Image)
What the model doesComposites the actual pixels into the outputUses them as style/composition inspiration
Logo fidelityPixel-accurateApproximated — the model redraws a similar-looking shape
Right tool for...Branded compositions, thumbnails, asset embeddingMatching aesthetic / mood / palette

If you upload a logo and the output shows a similar-looking but slightly wrong version of it, you used the wrong tool. Move the logo from Reference Images to Source Images.

Faces & Character Consistency

For pixel-accurate face matching of a specific person:

  • Preferred: Use Image Editing with the face photo as Source Image #1, plus any other assets (logos, products) as the remaining Source Images. Do not also apply a Visual DNA of the same person — that produces face averaging (a similar-but-wrong face).
  • Also OK: Use Text to Image with a Visual DNA, but omit physical descriptions of the person from the prompt. Visual DNA always injects its saved description into the prompt automatically (that's how it carries identity) — re-describing the person in your prompt competes with the DNA and causes drift. Let the DNA do the work.
  • Wrong: Combining a face Source Image with a face Visual DNA for the same person. Pick one anchor.

For style / product / scene consistency (no specific identity to lock), Visual DNA works well in either tool.

Common Workflows

YouTube Thumbnail (face + brand logos, pixel-accurate)

  1. Upload the host's portrait + each brand logo to your Media Library
  2. Open Image Editing, select Nano Banana 2 (Image Editing variant)
  3. Set Source Images: host face first, logos after
  4. In the prompt, refer to each Source Image by ordinal position ("the FIRST source image as the hero, the SECOND source image as the brand logo in the bottom-left billing block")
  5. Add "composite AS-IS, do not redraw or restyle" for the logo positions
  6. Generate at 16:9. Cost ~8 credits.

Background Replacement

  1. Upload the subject photo as the single Source Image
  2. Describe the new background in the prompt
  3. Generate

Product Restyling

  1. Upload the product shot as a single Source Image
  2. Describe the new setting / lighting / styling
  3. Generate

Common Mistakes

  • Logos in Reference Images (Text to Image) → they get redrawn as approximations. Move them to Source Images in Image Editing.
  • Face Source + Face DNA for the same person → face averaging. Pick one.
  • Physical description in prompt + Visual DNA → the injected DNA description competes with your prompt and visual reference. Drop the physical description from your prompt when using a face Visual DNA.
  • Implicit model name (e.g. nano-banana-2 on the API) when you want the edit endpoint. Use the explicit nano-banana-2-image-editing identifier.
  • Hallucinated billing text in branded layouts: prompts that ask for a "billing block" sometimes invite extra invented labels. Specify "only the logo image and the wordmark text, nothing else" to keep the block minimal.

Iterating on Edits

When you want to refine an edit you just made, use the previous output as a new Source Image. Don't go back to Text to Image and re-prompt from scratch — you'll lose the work and burn credits.

Comparison with Other Tools

vs Text to Image:

  • Image Editing: Pixel-accurate compositing — assets you upload appear as-is
  • Text to Image: Style-only references — assets inspire but don't embed

vs Canvas:

  • Image Editing: AI-driven workflow, prompt-based
  • Canvas: Professional editing with manual tools

vs Training Lab:

  • Image Editing: One-off edits and composites with any model
  • Training Lab: Consistent character/style results from trained models

Use Cases

  • YouTube thumbnails with face + brand logos
  • Product photography with embedded packaging or labels
  • Marketing compositions that combine multiple assets
  • Background changes
  • Style transformations
  • Outfit / object swaps
  • A/B testing variations of a composition