Models · Note

Nano Banana: the AI image model named at 2:30 a.m.

Google’s hit picture-making AI got its silly name when a product manager mashed two nicknames together in the middle of the night. Here’s that story, and how AI turns words into a picture.

In late July 2025, a team at Google was getting a new AI model ready to launch. It could make pictures and edit them from plain written instructions. It already had its official name, Gemini 2.5 Flash Image. What it still needed was a code name.

The code name was for LMArena, a website where people test AI models without knowing which is which. You type a request, get answers from two unnamed models and vote for the better one. Only then does the site show which models you were comparing. Teams often put unfinished models there to see how real people rate them, so those models go in under a code name.

A name picked at 2:30 a.m.

“We pushed the codename conversation until the last minute,” Naina Raisinghani, a product manager on the team, told Google’s blog. At 2:30 a.m., another product manager messaged her to say they had to submit it. She suggested “something funny like ‘Nano Banana’”. The answer came back: “Yeah, sure. That’s completely nonsensical.”

The name wasn’t random. “Some of my friends call me Naina Banana, and others call me Nano because I’m short and I like computers,” she said. “So I just smushed my two nicknames together.” It suited the model too, she added, “because it was a Flash model”. Flash is Google’s name for its speedy Gemini models.

A four-panel comic. At 2:30 a.m. a woman with long dark hair, seen from behind, works at a laptop while her phone lights up with a message. Two thought bubbles float above her, one reading Naina Banana and one reading Nano. A pair of hands squishes them into one yellow badge reading Nano Banana. A laptop screen shows the reply: Yeah, sure. That’s completely nonsensical.

Then the internet went bananas

Nano Banana appeared on LMArena in early August. People were stunned by how well it edited photos. For example, it could keep a person looking like themselves. Then they saw the name, and in the words of Google’s blog, “social media went bananas.” After a few weeks of guessing, Google hinted on X that the model was its own.

The votes backed up the excitement. While it was being tested, this model alone drew more than 2.5 million votes, LMArena says, and it had the biggest points lead in the site’s history. On 26 August 2025, Google launched it under its official name, Gemini 2.5 Flash Image, “aka nano-banana”.

The silly name stuck. Google turned the run button for Nano Banana in its AI Studio yellow, added a banana emoji to the “Create image” option in the Gemini app and even made limited-edition banana swag. Later versions kept the name as well: Nano Banana Pro came out in November 2025, and Nano Banana 2 in February 2026.

How does an AI make a picture?

Here is what a tool like this can do. We gave OpenAI’s image generator, not Nano Banana, one sentence: “A photo of a tiny banana, no bigger than a grain of rice, balanced on someone’s fingertip at a desk late at night, with a laptop glowing in the background and a wall clock showing 2:30.”

A photo-like picture made by AI. A tiny yellow banana, much smaller than a fingernail, balances on the tip of a finger in front of a glowing laptop on a desk at night. A wall clock behind reads 2:30. Around the desk are a mug, a notebook, a stack of books and notes pinned to the wall, all with short handwritten slogans.

No camera took that picture. To make one, an image model first has to learn what things look like, and it learns from huge collections of pictures that come with words. One public collection, LAION-5B, holds 5.85 billion pictures paired with text. From examples like these, a model picks up which words go with which shapes, colours and styles.

Then it has to draw. One common way is called diffusion. The model starts with a picture of pure random noise, like the static on an old TV. It removes a little of that noise, then a little more, again and again. The words steer every step, so the fuzz slowly turns into the picture that was asked for.

A four-panel comic. 1. Words: a teenager types the words a tiny banana on a fingertip into a laptop. 2. Random noise: a small green robot painter stands beside a canvas covered in grey TV static. 3. Clean up a little, many times: the robot wipes the static with a cloth, and a blurry yellow banana and fingertip begin to show. 4. A picture: the canvas shows a clear banana resting on a fingertip.

Diffusion isn’t the only way. Some models build a picture out of small pieces, one after another, a bit like a chatbot writing a reply word by word. OpenAI has said the picture maker built into GPT-4o works this way.

So which way does Nano Banana work? Google hasn’t spelled it out. Its model card for Nano Banana 2, a kind of fact sheet, says it is built on Gemini 3 Flash, one of Google’s general Gemini models, and Google says it draws on Gemini’s knowledge of the real world. Neither says how the picture itself gets made.

One thing Google has said: every picture made or edited with the first Nano Banana carries an invisible watermark called SynthID, so it can be spotted as AI-made. As the tiny banana above shows, a picture made from one sentence can look a lot like a real photo.

Back to the field notes