Foundations

Reasoning modes, explained

What happens when a model reasons before it answers: chain of thought, reasoning tokens, when they help, what they cost, and how Think harder works in Brello 1.0.

Brello Research9 min readVersion 1.0

Summary

A reasoning mode runs a language model so that it writes out a chain of thought before its final answer. The reasoning tokens are generated like any others, so they cost time and count against the output limit. Chain-of-thought prompting improved multi-step reasoning in large models, and models such as Qwen3 1.7B are now trained to switch between thinking and non-thinking modes. Written reasoning is useful but is not guaranteed to reflect how an answer was produced. In Brello 1.0, Think harder works with all three models, raises the output limit from 1,200 to 2,048 tokens and streams the reasoning into a “Thought process” panel above the answer.

  • Reasoning tokens are ordinary generated tokens: they take time to write and count against the same output limit as the answer.
  • Wei and colleagues found that chain-of-thought prompting helped only in sufficiently large models, around 100 billion parameters in their experiments; Qwen3 1.7B has a thinking mode built in.
  • Qwen’s model card recommends temperature 0.6 and top-p 0.95 for thinking mode and warns that greedy decoding “can lead to performance degradation and endless repetitions”.
  • Written reasoning is not guaranteed to be faithful: in experiments, models’ explanations left out biases that had changed their answers.
  • In Brello 1.0, Think harder allows up to 2,048 output tokens instead of 1,200 and shows the reasoning live in a “Thought process” panel above the answer.
Contents8 sections

01

What is a reasoning mode?

A reasoning mode is a way of running a language model so that it writes out intermediate reasoning, a chain of thought, before its final answer. The reasoning is ordinary generated text, marked so that an app can show it apart from the answer. It tends to help on problems with several steps, and it costs time and length.

Some models are trained to reason this way, and many can switch it on and off. Qwen’s model card describes Qwen3 1.7B as switching between a thinking mode “for complex logical reasoning, math, and coding” and a non-thinking mode “for efficient, general-purpose dialogue” 9. Google’s model card describes all Gemma 4 models as “designed as highly capable reasoners, with configurable thinking modes” 10. Brello 1.0 is built on those open models from Alibaba and Google, which do not endorse it, and calls its reasoning mode Think harder.

02

What are chain of thought and reasoning tokens?

Chain of thought is the sequence of intermediate steps a model writes on the way to an answer, and reasoning tokens are the tokens it spends writing them. They are generated exactly like answer tokens, one at a time, and they count against the same output limit.

The idea was first shown with prompting. Wei and colleagues gave large models a few worked examples that included intermediate steps, and found that this chain-of-thought prompting improved their performance on arithmetic, commonsense and symbolic reasoning 1. Kojima and colleagues then found that adding “Let’s think step by step” to a question had a similar effect without any examples 2.

Reasoning models are trained to produce long chains of thought on their own. DeepSeek-AI showed that reinforcement learning on problems with checkable answers, such as mathematics with known results, could teach a model to reason at length 3. Qwen3’s technical report describes combining thinking and non-thinking modes in one model, so that the same weights can answer directly or reason first, and a thinking budget that caps how many tokens the reasoning may use 4.

To keep the two apart, models mark where their reasoning starts and ends. Qwen3 puts it in a <think>…</think> block before the answer 9, and Gemma 4 writes it in a separate “thought” channel 10. An app reads those markers to decide which text is reasoning and which is answer.

03

When does reasoning help?

Mostly on problems that need several dependent steps, such as arithmetic word problems, logic puzzles and code. For simple questions and conversation, it mainly adds time.

Wei and colleagues found that chain-of-thought prompting gave larger gains on more complicated problems, and that the gains appeared only in sufficiently large models, around 100 billion parameters in their experiments; smaller models wrote fluent but illogical chains 1. Qwen’s own description points the same way, with a thinking mode for complex logical reasoning, maths and coding, and a non-thinking mode for general dialogue 9.

Training changes the picture for small models. Qwen3’s report describes passing reasoning ability from its largest models to its smallest through “strong-to-weak distillation” 4, and Qwen3 1.7B has a thinking mode built in. Spending more computation at answer time can also stand in for size: Snell and colleagues found that on problems a smaller model already solves some of the time, extra computation at inference could outperform a model 14 times larger 5.

A useful test is a question with a tempting wrong answer. In the bat-and-ball problem from Frederick’s Cognitive Reflection Test, here in pounds, a bat and a ball cost £1.10 together and the bat costs £1.00 more than the ball. The quick answer for the ball, 10p, is wrong; the right one is 5p 6. Writing out the steps gives a model the chance to check the quick answer before it commits to one. Figure 1 follows that reasoning through Brello 1.0.

How Brello keeps a model’s reasoning apart from its answer The text Qwen3 1.7B writes for one question with Think harder on: reasoning between a think tag and a closing think tag, then the answer. Beside it, what Brello shows: the question, a status line reading Reasoning, a Thought process panel that streams the last few lines of the reasoning, and the answer. Tokens inside the think block go to the panel, and tokens after the closing tag go to the answer. When the answer is done, the panel collapses to a Thought process row. A bar at the bottom counts output tokens: 95 for the reasoning and 17 for the answer, 112 of the 2,048 that Think harder allows, against 1,200 normally. Model output Qwen3 1.7B What Brello shows Reasoning Answer <think> The quick answer is 10p. Check it: the bat would be £1.10, the total £1.20. Let the ball be x; the bat is x + £1.00. 2x + £1.00 = £1.10, so 2x = £0.10. x = £0.05, and the bat is £1.05. </think> The ball costs 5p, and the bat costs £1.05. A bat and a ball cost £1.10. The bat costs £1.00 more than the ball. How much is the ball? Reasoning Think harder on Thought process The quick answer is 10p. Check it: the bat would be £1.10, the total £1.20. Let the ball be x; the bat is x + £1.00. 2x + £1.00 = £1.10, so 2x = £0.10. x = £0.05, and the bat is £1.05. The ball costs 5p, and the bat costs £1.05. Output tokens Reasoning 95 Answer 17 112 of 2,048 used 1,200 · normal limit 2,048 · Think harder limit Model output · Qwen3 1.7B <think> The quick answer is 10p. Check it: … x = £0.05, and the bat is £1.05. </think> The ball costs 5p, and the bat costs £1.05. What Brello shows A bat and a ball cost £1.10. The bat costs £1.00 more than the ball. How much is the ball? Reasoning Think harder on Thought process The quick answer is 10p. Check it: the bat would be £1.10, the total £1.20. Let the ball be x; the bat is x + £1.00. 2x + £1.00 = £1.10, so 2x = £0.10. x = £0.05, and the bat is £1.05. The ball costs 5p, and the bat costs £1.05. Output tokens 1,200 2,048 Reasoning 95 Answer 17
  1. You ask with Think harder on. The status line reads ‘Reasoning’, and the reply may use up to 2,048 output tokens instead of the usual 1,200.
  2. The model writes its reasoning first, inside a <think> block. Brello sends these tokens to the ‘Thought process’ panel, which shows the last few lines as they arrive.
  3. The closing </think> tag switches the stream to the answer. Everything after it goes to the reply, so the reasoning stays out of the answer text.
  4. The panel collapses to a ‘Thought process’ row. It expands to show the full reasoning. This reply used 112 of its 2,048 output tokens: 95 for the reasoning, tags included, and 17 for the answer.
Figure 1How Brello 1.0 keeps a model’s reasoning apart from its answer, shown with the <think> tags of Qwen3 1.7B, the base model of Brello Core. The question is the bat-and-ball problem from the Cognitive Reflection Test, in pounds 6; the reasoning text is illustrative, and token counts are from Qwen3’s tokeniser.

04

What does reasoning cost?

Time, output length and room in the context window. Reasoning tokens are generated one at a time like any others, so 1,000 tokens of reasoning take about as long to write as 1,000 tokens of answer, and the answer itself starts only when the reasoning ends.

Output limits apply to the reasoning too. In Brello 1.0, a normal answer can run to 1,200 tokens; with Think harder the limit is 2,048, and the reasoning is part of that output. On a phone those tokens take real time. litert-community measured the file Brello Core runs writing 40.5–40.7 tokens per second on a Samsung Galaxy S26’s GPU and 8.0–8.3 on its CPU 11. At those rates, 2,048 tokens would take about 50 seconds on the GPU and over four minutes on the CPU. These are the publisher’s measurements on one phone, not Brello’s, and most replies are far shorter.

Reasoning also takes space. In a 4,096-token window, a 2,048-token output leaves 2,048 tokens for the instructions, the earlier messages and any web passages; Context windows, explained shows the whole budget.

Long, sampled reasoning can also fail in a particular way: it can loop. Qwen’s model card warns that greedy decoding in thinking mode “can lead to performance degradation and endless repetitions”, and recommends a temperature of 0.6 and a top-p of 0.95 9. Brello’s Think harder samples with the same two values. Its repetition stopper checks every 48 characters for a block repeated 3 or more times over at least 120 characters, then cuts the reply after the first copy and stops the model.

05

What does showing the reasoning tell the user?

It shows the steps the model wrote, which lets people follow and check them. It does not prove that those steps produced the answer, so visible reasoning is best read as a useful account, not a faithful record.

Research on faithfulness explains the caution. Turpin and colleagues biased models towards particular answers, for example by reordering the options in worked examples so that the correct one was always “(A)”, and found that the models’ written explanations justified the biased answers without mentioning the bias 7. Lanham and colleagues tested how much models rely on their stated reasoning by editing it, for instance by cutting it short or inserting mistakes, and found that the reliance varied widely from task to task 8.

Brello 1.0 shows the reasoning while it is written and keeps it afterwards. With Think harder on, a “Thought process” panel appears above the answer. While the model reasons, the panel’s title shimmers, the panel shows the last few lines of live reasoning, fading out at the top, and the status line reads “Reasoning” instead of “Thinking”. When the answer is done, the panel collapses to a “Thought process” row that expands to show the full reasoning. For more on how Brello presents what the model is doing, see our note on designing visible intelligence.

06

How well do small on-device models reason?

Less reliably than large ones. Small models hold less knowledge and are weaker at long or complex reasoning, so their chains of thought help but still go wrong, and a fluent chain is no guarantee of a correct answer.

On-device models are far smaller than frontier cloud models: they can be wrong, have a knowledge cutoff and are weaker at long or complex reasoning. Brello’s system prompt tells the model to say when it is not sure, but that is no guarantee.

Phones add two constraints. The first is length. Qwen’s model card recommends an output length of 32,768 tokens for most queries 9, while Brello 1.0 allows 2,048 inside a 4,096-token window, so a chain that would run longer is stopped at the limit. The second is speed: every reasoning token is written at the phone’s pace (section 04).

Small models also slip out of their format. Brello removes stray markers, such as a leftover <think> tag or Gemma’s channel markers, or moves them to the thought panel. If Think harder is off and a model reasons anyway but never reaches an answer, Brello shows the reasoning as the answer instead of an empty reply.

Small language models, explained covers what these models give up for their size, and our note on how small models copy the shape of their instructions covers why their prompts are kept short.

07

How does Think harder work in Brello 1.0?

Think harder is Brello 1.0’s reasoning mode. When it is on, the model reasons before it answers, the reasoning streams into a “Thought process” panel above the answer, and the output limit rises from 1,200 to 2,048 tokens.

Each model is described on the Brello models page, and Brello 1.0 describes the app as a whole.

References

Reviewed . Model cards and model pages were checked on that date.

  1. Wei, J. et al. (2022). “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.” Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arxiv.org/abs/2201.11903
  2. Kojima, T., Gu, S. S., Reid, M., Matsuo, Y. and Iwasawa, Y. (2022). “Large Language Models are Zero-Shot Reasoners.” Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arxiv.org/abs/2205.11916
  3. DeepSeek-AI (2025). “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” arXiv preprint. arxiv.org/abs/2501.12948
  4. Qwen Team (2025). “Qwen3 Technical Report.” arXiv preprint. arxiv.org/abs/2505.09388
  5. Snell, C., Lee, J., Xu, K. and Kumar, A. (2024). “Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.” arXiv preprint. arxiv.org/abs/2408.03314
  6. Frederick, S. (2005). “Cognitive Reflection and Decision Making.” Journal of Economic Perspectives 19 (4): 25–42. doi.org/10.1257/089533005775196732
  7. Turpin, M., Michael, J., Perez, E. and Bowman, S. R. (2023). “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.” Advances in Neural Information Processing Systems 36 (NeurIPS 2023). arxiv.org/abs/2305.04388
  8. Lanham, T. et al. (2023). “Measuring Faithfulness in Chain-of-Thought Reasoning.” arXiv preprint. arxiv.org/abs/2307.13702
  9. Qwen. “Qwen3-1.7B.” Model card, Hugging Face. Accessed 5 October 2026. huggingface.co/Qwen/Qwen3-1.7B
  10. Google DeepMind. “Gemma 4 model card.” Google AI for Developers. Accessed 5 October 2026. ai.google.dev/gemma/docs/core/model_card_4
  11. litert-community. “Qwen3-1.7B.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/litert-community/Qwen3-1.7B

Version history

  1. 1.0First published.