Models · Note

“Take a deep breath” helped an AI with maths

Another AI found the odd instruction. What happened in the test is more interesting than the idea of a robot calming its nerves.

Picture an AI at a desk, stuck on a pile of school maths questions. What would you tell it? In a 2023 Google DeepMind study, one winning line was: “Take a deep breath and work on this problem step-by-step.”

Comic illustration of a small robot calmly sitting with a stack of maths worksheets while a teacher looks surprised.

The researchers tested a version of an AI model called PaLM 2 on 1,319 grade-school maths word problems. With no extra line, it got 34.0% right. The researchers made its answer start with “Let’s think step by step”, and the score rose to 71.8%. With the deep-breath line, it reached 80.2%. These are scores for this model on this test, not a promise that the line will help every AI.

Another AI wrote the line

The researchers had one AI suggest instructions and a second AI try them on maths questions. The first AI saw earlier suggestions and their scores, then wrote new ones. The researchers called this method OPRO. Each round meant suggesting new instructions and scoring them. The deep-breath line turned up in round 107. The team then checked promising instructions on all 1,319 test questions.

Comic illustration of one robot passing a blank instruction card to another robot solving a worksheet, with a curved arrow showing feedback.

Was it really the breathing?

The study did not test “take a deep breath” by itself. It tested the whole sentence. Another instruction, “Break this down,” scored 79.9% on the same test, only 0.3 percentage points below the winner. The robot in the comic can pose like it is meditating; the real model has no lungs to fill.

The useful idea is that small changes to an instruction can change a model's answer. A separate 2022 study found that adding “Let’s think step by step” took one model from 17.7% to 78.7% on a different maths test. Asking for steps can help, but the words that worked best in one experiment may not win on the next model or problem.

Back to the field notes