Google has made Gemini Nano Banana 2.1 generally available through the Gemini API, positioning it as the latest high-efficiency model for image generation and conversational image editing. The release succeeds Gemini 3.1 Flash Image, also known as Nano Banana 2, which Google says will be shut down on October 29, 2026.

The model supports text, image, video and PDF inputs and produces image and text output. Google says the update improves visual quality, prompt adherence, text rendering, infographic layout and multi-turn character consistency while retaining a speed-and-cost profile intended to sit below Gemini 3 Pro Image. It supports 1K, 2K and 4K output and specifically addresses tiling artifacts in very wide or panoramic aspect ratios including 1:4, 4:1, 1:8 and 8:1.

For reference-driven editing, Google documents multi-image fusion with as many as 14 reference images, including consistency support for up to four characters and object fidelity for up to 10 objects. The model also supports grounding with Google Web and Image Search and exposes configurable thinking levels. Google's model card identifies Nano Banana 2.1 as part of the Gemini 3 family and says it is based on Gemini 3.6 Flash.

Google's own model card reports improvements over Nano Banana 2 and Gemini 3 Pro Image on several internal or Google-reported generation and editing evaluations, including overall preference, infographic tasks and character consistency. Those measurements provide useful detail about the intended improvements, but they are not independent benchmark confirmation. This run did not find a credible third-party evaluation published after the release, so performance comparisons should remain attributed to Google.

The release matters for developers because it combines a stable model identifier with a migration deadline for the previous Flash Image model. Teams using gemini-3.1-flash-image have a concrete replacement path but also a relatively short transition window before the October 29 shutdown. The practical questions now are how the quality improvements hold up across real production prompts, whether character and object consistency remain reliable over long editing sessions, and how the model's cost and latency compare in independent workloads.

References