What sets them apart

Rather than generating from scratch, these models take an existing image plus a plain-language instruction — 'remove the background,' 'change the lighting to sunset' — and apply a targeted edit while preserving everything else.

How they work

They're trained on paired before/after images with instructions, learning to localize the requested change without disturbing unrelated parts of the image.

Where they fit

Ideal for iterative creative workflows and preserving brand assets — logos, product shots, headshots — across edits, rather than regenerating an entire scene from a text prompt alone.

Choosing the right model for your agent? Model selection and integration is part of every build we do — let's talk through your use case.

Start a conversation →