- calendar_today August 10, 2025
Through its official release, OpenAI launched “Images in ChatGPT,” which brings image generation capabilities directly into the ChatGPT user experience. The GPT-4o model enables this feature advancement, which allows users to generate images during conversations with ChatGPT, representing a major progression in AI content creation.
ChatGPT subscribers at all levels now benefit from access to advanced image generation capabilities, including the Plus, Pro, Team tiers, as well as the free edition. Users of the free tier face comparable restrictions to DALL-E 3 with a daily generation limit of three images, but Taya Christianson from OpenAI stated these restrictions could be modified according to demand. Dedicated custom GPTs will allow DALL-E enthusiasts to maintain their access.
OpenAI’s research lead Gabriel Goh explained that GPT-4o represents a major advance because it works with multiple data forms such as text, images, audio, and video as its “omnimodal” functionality allows. The model now has an improved “binding” capability, which successfully tackles a long-standing problem in AI image generation. GPT-4o shows improved performance by managing 15 to 20 objects without confusing their colors or shapes, unlike earlier models, which struggled with object-attribute relationships.
The system now provides superior text rendering capabilities. AI-generated images have historically shown problems with distorted or meaningless text. Goh explained that their development process involved extensive iterative testing over several months to achieve the final result. Despite the difficulty of perfect text rendering, especially for small text, the team has reached a consistency level that makes text in images reliably usable.
The system utilizes an autoregressive architecture, which sets it apart from standard diffusion models used in image generation. The sequential image generation process that starts from left to right and top to bottom resembles text production and improves text rendering and binding performances.
The company revealed the system’s multiple capabilities during their presentation through demonstrations that included producing detailed scientific diagrams, such as Newton’s prism experiment with exact labels, along with comic panels featuring coherent characters and text, and informational posters with precise wording. The demonstration included practical uses for the system, which produced transparent background images for stickers, restaurant menus, and logos.
The system’s ability to utilize world knowledge was highlighted by Jackie Shannon, who serves as ChatGPT’s lead product manager for multimodal applications. She explained that while she draws images within her personal capability limits, she utilizes her entire world knowledge base. The model incorporates world knowledge, which enables you to request an image of Newton’s prism experiment without needing to describe what it is to obtain the image.
OpenAI maintains that the improved image quality and advanced capabilities outweigh the increased generation time for users. Shannon acknowledged room for latency improvements and yet emphasized that the superior image quality and expanded capabilities alongside world knowledge compensate for any extra waiting time.
Addressing Ethical Concerns and Ensuring Responsible Deployment
OpenAI emphasized its commitment to preventing misuse by implementing strong security measures. To ensure ethical usage, the system blocks CSAM requests and sexual deepfake creation while also preventing watermark removal. All images produced by the system will carry standard C2PA metadata despite the lack of visual watermarks to indicate their source as OpenAI products. The company runs internal verification tools for images.
Shannon affirmed that while no system achieves perfection for these tasks, we keep updating protections and consider this our baseline approach. Users who generate images through ChatGPT own them and may use them as desired while adhering to our usage policies.
OpenAI has expanded ChatGPT capabilities through “Images in ChatGPT,” while establishing a new industry benchmark for accessible AI-powered image creation. Enhanced binding techniques combined with superior text rendering capabilities and strong protective measures prove OpenAI’s dedication to creating a tool that delivers power alongside responsible operation. The company’s adoption of an autoregressive approach represents an innovative departure from conventional diffusion models for image creation. Through its focus on user ownership and metadata integration, OpenAI underscores its dedication to transparency and ethical usage within the growing field of AI-generated content. This integration represents an important advancement in making sophisticated AI image creation accessible to everyone and mitigates potential dangers to provide enhanced user safety.





