01
What is a reasoning mode?
A reasoning mode is a way of running a language model so that it writes out intermediate reasoning, a chain of thought, before its final answer. The reasoning is ordinary generated text, marked so that an app can show it apart from the answer. It tends to help on problems with several steps, and it costs time and length.
Some models are trained to reason this way, and many can switch it on and off. Qwen’s model card describes Qwen3 1.7B as switching between a thinking mode “for complex logical reasoning, math, and coding” and a non-thinking mode “for efficient, general-purpose dialogue” 9. Google’s model card describes all Gemma 4 models as “designed as highly capable reasoners, with configurable thinking modes” 10. Brello 1.0 is built on those open models from Alibaba and Google, which do not endorse it, and calls its reasoning mode Think harder.
02
What are chain of thought and reasoning tokens?
Chain of thought is the sequence of intermediate steps a model writes on the way to an answer, and reasoning tokens are the tokens it spends writing them. They are generated exactly like answer tokens, one at a time, and they count against the same output limit.
The idea was first shown with prompting. Wei and colleagues gave large models a few worked examples that included intermediate steps, and found that this chain-of-thought prompting improved their performance on arithmetic, commonsense and symbolic reasoning 1. Kojima and colleagues then found that adding “Let’s think step by step” to a question had a similar effect without any examples 2.
Reasoning models are trained to produce long chains of thought on their own. DeepSeek-AI showed that reinforcement learning on problems with checkable answers, such as mathematics with known results, could teach a model to reason at length 3. Qwen3’s technical report describes combining thinking and non-thinking modes in one model, so that the same weights can answer directly or reason first, and a thinking budget that caps how many tokens the reasoning may use 4.
To keep the two apart, models mark where their reasoning starts and ends. Qwen3 puts it in a <think>…</think> block before the answer 9, and Gemma 4 writes it in a separate “thought” channel 10. An app reads those markers to decide which text is reasoning and which is answer.
03
When does reasoning help?
Mostly on problems that need several dependent steps, such as arithmetic word problems, logic puzzles and code. For simple questions and conversation, it mainly adds time.
Wei and colleagues found that chain-of-thought prompting gave larger gains on more complicated problems, and that the gains appeared only in sufficiently large models, around 100 billion parameters in their experiments; smaller models wrote fluent but illogical chains 1. Qwen’s own description points the same way, with a thinking mode for complex logical reasoning, maths and coding, and a non-thinking mode for general dialogue 9.
Training changes the picture for small models. Qwen3’s report describes passing reasoning ability from its largest models to its smallest through “strong-to-weak distillation” 4, and Qwen3 1.7B has a thinking mode built in. Spending more computation at answer time can also stand in for size: Snell and colleagues found that on problems a smaller model already solves some of the time, extra computation at inference could outperform a model 14 times larger 5.
A useful test is a question with a tempting wrong answer. In the bat-and-ball problem from Frederick’s Cognitive Reflection Test, here in pounds, a bat and a ball cost £1.10 together and the bat costs £1.00 more than the ball. The quick answer for the ball, 10p, is wrong; the right one is 5p 6. Writing out the steps gives a model the chance to check the quick answer before it commits to one. Figure 1 follows that reasoning through Brello 1.0.
- You ask with Think harder on. The status line reads ‘Reasoning’, and the reply may use up to 2,048 output tokens instead of the usual 1,200.
- The model writes its reasoning first, inside a
<think>block. Brello sends these tokens to the ‘Thought process’ panel, which shows the last few lines as they arrive. - The closing
</think>tag switches the stream to the answer. Everything after it goes to the reply, so the reasoning stays out of the answer text. - The panel collapses to a ‘Thought process’ row. It expands to show the full reasoning. This reply used 112 of its 2,048 output tokens: 95 for the reasoning, tags included, and 17 for the answer.
<think> tags of Qwen3 1.7B, the base model of Brello Core. The question is the bat-and-ball problem from the Cognitive Reflection Test, in pounds 6; the reasoning text is illustrative, and token counts are from Qwen3’s tokeniser.04
What does reasoning cost?
Time, output length and room in the context window. Reasoning tokens are generated one at a time like any others, so 1,000 tokens of reasoning take about as long to write as 1,000 tokens of answer, and the answer itself starts only when the reasoning ends.
Output limits apply to the reasoning too. In Brello 1.0, a normal answer can run to 1,200 tokens; with Think harder the limit is 2,048, and the reasoning is part of that output. On a phone those tokens take real time. litert-community measured the file Brello Core runs writing 40.5–40.7 tokens per second on a Samsung Galaxy S26’s GPU and 8.0–8.3 on its CPU 11. At those rates, 2,048 tokens would take about 50 seconds on the GPU and over four minutes on the CPU. These are the publisher’s measurements on one phone, not Brello’s, and most replies are far shorter.
Reasoning also takes space. In a 4,096-token window, a 2,048-token output leaves 2,048 tokens for the instructions, the earlier messages and any web passages; Context windows, explained shows the whole budget.
Long, sampled reasoning can also fail in a particular way: it can loop. Qwen’s model card warns that greedy decoding in thinking mode “can lead to performance degradation and endless repetitions”, and recommends a temperature of 0.6 and a top-p of 0.95 9. Brello’s Think harder samples with the same two values. Its repetition stopper checks every 48 characters for a block repeated 3 or more times over at least 120 characters, then cuts the reply after the first copy and stops the model.
05
What does showing the reasoning tell the user?
It shows the steps the model wrote, which lets people follow and check them. It does not prove that those steps produced the answer, so visible reasoning is best read as a useful account, not a faithful record.
Research on faithfulness explains the caution. Turpin and colleagues biased models towards particular answers, for example by reordering the options in worked examples so that the correct one was always “(A)”, and found that the models’ written explanations justified the biased answers without mentioning the bias 7. Lanham and colleagues tested how much models rely on their stated reasoning by editing it, for instance by cutting it short or inserting mistakes, and found that the reliance varied widely from task to task 8.
Brello 1.0 shows the reasoning while it is written and keeps it afterwards. With Think harder on, a “Thought process” panel appears above the answer. While the model reasons, the panel’s title shimmers, the panel shows the last few lines of live reasoning, fading out at the top, and the status line reads “Reasoning” instead of “Thinking”. When the answer is done, the panel collapses to a “Thought process” row that expands to show the full reasoning. For more on how Brello presents what the model is doing, see our note on designing visible intelligence.
06
How well do small on-device models reason?
Less reliably than large ones. Small models hold less knowledge and are weaker at long or complex reasoning, so their chains of thought help but still go wrong, and a fluent chain is no guarantee of a correct answer.
On-device models are far smaller than frontier cloud models: they can be wrong, have a knowledge cutoff and are weaker at long or complex reasoning. Brello’s system prompt tells the model to say when it is not sure, but that is no guarantee.
Phones add two constraints. The first is length. Qwen’s model card recommends an output length of 32,768 tokens for most queries 9, while Brello 1.0 allows 2,048 inside a 4,096-token window, so a chain that would run longer is stopped at the limit. The second is speed: every reasoning token is written at the phone’s pace (section 04).
Small models also slip out of their format. Brello removes stray markers, such as a leftover <think> tag or Gemma’s channel markers, or moves them to the thought panel. If Think harder is off and a model reasons anyway but never reaches an answer, Brello shows the reasoning as the answer instead of an empty reply.
Small language models, explained covers what these models give up for their size, and our note on how small models copy the shape of their instructions covers why their prompts are kept short.
07
How does Think harder work in Brello 1.0?
Think harder is Brello 1.0’s reasoning mode. When it is on, the model reasons before it answers, the reasoning streams into a “Thought process” panel above the answer, and the output limit rises from 1,200 to 2,048 tokens.
Each model is described on the Brello models page, and Brello 1.0 describes the app as a whole.
References
Reviewed . Model cards and model pages were checked on that date.
- Wei, J. et al. (2022). “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.” Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arxiv.org/
abs/ 2201.11903 - Kojima, T., Gu, S. S., Reid, M., Matsuo, Y. and Iwasawa, Y. (2022). “Large Language Models are Zero-Shot Reasoners.” Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arxiv.org/
abs/ 2205.11916 - DeepSeek-AI (2025). “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” arXiv preprint. arxiv.org/
abs/ 2501.12948 - Qwen Team (2025). “Qwen3 Technical Report.” arXiv preprint. arxiv.org/
abs/ 2505.09388 - Snell, C., Lee, J., Xu, K. and Kumar, A. (2024). “Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.” arXiv preprint. arxiv.org/
abs/ 2408.03314 - Frederick, S. (2005). “Cognitive Reflection and Decision Making.” Journal of Economic Perspectives 19 (4): 25–42. doi.org/
10.1257/ 089533005775196732 - Turpin, M., Michael, J., Perez, E. and Bowman, S. R. (2023). “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.” Advances in Neural Information Processing Systems 36 (NeurIPS 2023). arxiv.org/
abs/ 2305.04388 - Lanham, T. et al. (2023). “Measuring Faithfulness in Chain-of-Thought Reasoning.” arXiv preprint. arxiv.org/
abs/ 2307.13702 - Qwen. “Qwen3-1.7B.” Model card, Hugging Face. Accessed 5 October 2026. huggingface.co/
Qwen/ Qwen3-1.7B - Google DeepMind. “Gemma 4 model card.” Google AI for Developers. Accessed 5 October 2026. ai.google.dev/
gemma/ docs/ core/ model_card_4 - litert-community. “Qwen3-1.7B.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/
litert-community/ Qwen3-1.7B
Version history
- 1.0First published.



