Viyan

Viyan AI

OpenAI Releases ChatGPT Image Models 2.5

OpenAI updated its image generation API with two new models, gpt-image-2.5-sunburst and gpt-image-2.5-flare, focused on instruction following and reference image fidelity.

OpenAI has released ChatGPT Image Models 2.5, a model update designed to improve instruction following over multi-turn conversations and maintain subject fidelity in reference images. The rollout introduces two new model IDs to the API: gpt-image-2.5-sunburst and gpt-image-2.5-flare. Builders now have access to explicit controls for reference images, allowing the model to treat user-provided assets as a foundation for subsequent generations rather than treating each prompt as an isolated request. This update is worth noting because it shifts the burden of subject consistency from custom prompt engineering to native API parameters. By providing a base reference, you are no longer relying on the model to hallucinate a consistent identity from a text description alone.

Choosing Between Model Variants

OpenAI has split the 2.5 release into two distinct models. The primary difference lies in the trade-off between creative control and latency. gpt-image-2.5-sunburst is optimized for accuracy in editing, while gpt-image-2.5-flare is tuned for speed. If you are building an application where the user expects a specific character or logo to persist through several iterations, sunburst is the clear choice. If your product is a fast-paced generator where the user is iterating to find a specific style, the reduced latency of flare will likely offer a better experience.

Feature gpt-image-2.5-sunburst gpt-image-2.5-flare
Primary Use Case Precision editing Fast generation
Subject Consistency High Moderate
Instruction Following Maximum Efficient
Latency Higher Lower

Implementation Strategy

If you are building applications that rely on image editing, such as adding characters to a chart or modifying assets based on user-provided photos, this release changes your integration strategy. Previously, maintaining a consistent subject across edits required complex prompt engineering or iterative retries that often failed to keep the subject's core features intact. With these models, you can now pass existing assets directly into the generation pipeline. This lets you build more predictable workflows where the "add a raccoon scientist" prompt results in the raccoon appearing in the context of your specific provided chart.

A concrete example of why this matters is the creation of branded marketing assets. Suppose you have a base photo of a physical retail kiosk. A user wants to see how a new sign would look in that space. Using a legacy model, you might get a sign, but the color palette or the perspective of the kiosk itself would likely drift with every attempt. By utilizing the 2.5 architecture, the reference image acts as the persistent constraint for the session.

What remains unknown is the absolute limit of these models regarding long-term state tracking. While the API handles the reference image as a foundation, we do not have public documentation on how the internal weighting of those reference pixels decays during extremely long sessions. It is also unclear how these models handle conflicting instructions—such as when a user asks to change the color of an object while simultaneously demanding the model retain the exact original shading. The balance between "follow the instruction" and "respect the reference" is a subjective threshold, and it is likely that different prompts will require different levels of tuning on your end to find the sweet spot.

Sources