Generative Image Models Are Getting Hard To Ignore
Published:
Image generation from text has improved remarkably quickly. Systems such as DALL-E 2, announced by OpenAI this year, can create images from written descriptions with a level of coherence that would have looked much less convincing only a short time ago.
The interesting part is not only visual quality. Natural language is becoming an interface for a creative model. A user can describe an unusual combination of objects, style, lighting, and setting, then receive an image that never existed before.
The models are still imperfect. Hands, text inside images, object relationships, and fine details can become strange. A beautiful result may fail when we look closely. But the rate of progress makes those failures difficult to dismiss.
These systems also create questions that are larger than image quality. Training requires enormous collections of images and text. Artists may wonder whether their work was included and how generated styles affect creative labor. Copyright law is not designed around machines that learn statistical patterns from huge datasets and then produce new images.
I do not think this means human creativity becomes unnecessary. A model does not have a childhood, personal memory, or intention in the human sense. But it can become a new creative tool, and tools change professions even when they do not replace the people using them.
The most interesting skill may become describing and selecting rather than drawing every element directly. That can lower barriers for some people and create anxiety for others.
We are still early enough that confident predictions are risky. The images are impressive. The social consequences are much harder to generate.