ChatGPT Images 2.5 shows high-fidelity output but struggles with dense Waldo scenes
Early Waldo Bench results for ChatGPT Images 2.5 indicate generally high-fidelity image generation. Waldo was found in one image, while some text remained highly detailed when enlarged; however, faces became garbled under strong zoom.
The beach scene was judged too easy and not dense enough for a full Waldo test—about one-quarter of a “real Waldo” scene. The assessment described the result as possibly state of the art, but not “super intelligence” for generating Waldo scenes.
The analysis suggested that modern AI may eventually solve Waldo Bench, potentially by pairing the system with Astra, tiling, and adding more reasoning. This was an informal assessment without quantitative metrics.
