01
What is a context window?
A context window is the maximum amount of text, counted in tokens, that a language model can take into account at once. Everything the model uses to write a reply has to fit inside it: its instructions, the conversation so far, any documents it has been given and the reply itself as it is written.
The limit depends on how a model was trained and on how it is run. Google’s model card gives Gemma 4 E2B and E4B a context length of 128K tokens 7, and Qwen’s gives Qwen3 1.7B 32,768 8. An app can run a model with a smaller window than the model supports, and Brello 1.0 runs all three of its models, which are built on those open models, with a window of 4,096 tokens. Brello is not affiliated with or endorsed by Google or Alibaba.
Table 1 shows the limit at each layer, from the base model to the converted file on the phone to the app.
| Model | Base model | LiteRT-LM file | Brello 1.0 |
|---|---|---|---|
| Brello ProGemma 4 E4B | 128K 7 | Up to 32k 11 | 4,096 |
| Brello VisionGemma 4 E2B | 128K 7 | Up to 32k 10 | 4,096 |
| Brello CoreQwen3 1.7B | 32,768 8 | 4,096 9 | 4,096 |
02
What is a token?
A token is the unit of text a model reads and writes: often a whole common word, sometimes a piece of a longer or rarer word, a number or a punctuation mark. A tokeniser turns text into tokens from a fixed vocabulary before the model sees it.
Tokenisers learn their vocabulary from data, building frequent sequences of characters up into single tokens; byte-pair encoding 2 and SentencePiece 3 are the best-known methods. Common words become one token and rarer ones are split into familiar pieces. Run through Qwen3 1.7B’s tokeniser, “Quantization makes models smaller.” becomes six tokens: Quant, ization, makes, models, smaller and the full stop. The name “Brello” becomes two, B and rello. Gemma 4’s tokeniser splits both the same way. Qwen3’s tokeniser has 151,669 entries 8, and the Gemma 4 model card gives a vocabulary of 262K 7.
Because tokens vary in length, a token count is not a word count. On a 3,235-character sample of plain English, Qwen3’s tokeniser produced 714 tokens and Gemma 4’s produced 753: about 4.5 and 4.3 characters per token. We counted with each model’s published tokeniser file 8 12. The ratio depends on the text, and the same content can take several times as many tokens in some languages as in others 4.
03
What fills the window?
Everything the model works from, plus everything it writes: the system prompt, earlier turns of the conversation, any retrieved text such as web results, the new message and, token by token, the reply.
Language models write one token at a time, and each new token is added to the sequence as input for the next, as in the decoder of the original transformer 1. The reply therefore draws on the same budget as the prompt. In a 4,096-token window, a prompt that leaves room for a 1,200-token reply can use at most 2,896 tokens, and one that leaves room for 2,048 can use at most 2,048.
Figure 1 fills Brello Pro’s window with each part at the limit Brello 1.0 sets for it, with web search and Think harder both on.
- Every reply is built in a window of 4,096 tokens. Each square is one token. Brello 1.0 uses the same window for all three of its models.
- Brello’s instructions take 114 tokens. The system prompt is four sentences, 56 tokens; the date (15) and the rule for citing sources (43) are added only when they are needed.
- The last six earlier messages follow, clipped. Your messages are cut to 300 characters and Brello’s to 600: at most 2,700 characters, about 630 tokens.
- Web passages fill up to 3,400 characters. With web search on, the best passages from up to four pages come to about 790 tokens.
- Your question goes in, and the answer is written into the same window. A normal answer can use up to 1,200 tokens, so it needs room left over.
- Think harder raises the output limit to 2,048 tokens. The reasoning counts against it too. With every limit reached, about 3,590 tokens are used and about 506 are left.
Show data
| Part | Limit in Brello 1.0 | Brello Pro, Brello Vision | Brello Core |
|---|---|---|---|
| System prompt | Four sentences | 56 | 54 |
| Date | Added for questions about time or with web results | 15 | 15 |
| Citation rule | Added with web results | 43 | 43 |
| Earlier messages | Last 6; yours to 300 characters, Brello’s to 600 | ≈630 | ≈600 |
| Web passages | 3,400 characters (Gemma 4), 2,600 (Qwen3) | ≈790 | ≈570 |
| Question | Illustrative | 8 | 8 |
| Output | 1,200 tokens; 2,048 with Think harder | 2,048 | 2,048 |
| Total | ≈3,590 | ≈3,338 | |
| Left of 4,096 | ≈506 | ≈758 |
Token counts for the prompt, date, citation rule and question come from each model’s published tokeniser. Earlier messages and web passages are estimated from their character limits at 4.3 characters per token for Gemma 4 and 4.5 for Qwen3, the averages we measured on a 3,235-character sample of English prose.
At those limits the parts add up to about 3,590 tokens, leaving about 506 of the 4,096. The estimate leaves out the short markers that separate turns and label sources, and the example question is short; a long pasted message would use more of what is left.
04
Why isn’t a longer window free?
Because every token in the window costs memory and computation. The model keeps a key and a value for every token in every layer, so memory grows in step with the window, and attention compares each new token with all the earlier ones, so its work grows faster still.
That store of keys and values is the key-value (KV) cache, and it saves the model from recomputing the past for every new token. Its size follows from the architecture. Qwen3 1.7B has 28 layers, each with 8 key-value heads of 128 dimensions 8, so every token needs 2 × 28 × 8 × 128 = 57,344 numbers. At 16 bits each, that is 114,688 bytes per token: about 470 MB for 4,096 tokens, and 3.76 GB for the model’s full 32,768. Runtimes can store the cache at lower precision, and this is arithmetic on the published architecture, not a measurement of Brello.
Qwen3 already shares each key-value head between two query heads, a technique called grouped-query attention 5. With a separate key and value for each of its 16 query heads, the cache would be twice as large.
A longer window also costs time. Attention compares each token with every token before it, so its work grows with the square of the length 1: twice the text means roughly four times the attention work. Before the first word of a reply appears, the model has to read the whole prompt. litert-community measured the 4-bit Qwen3 1.7B file reading a 202-token prompt at 532–568 tokens per second on a Samsung Galaxy S26’s GPU and 48–76 on its CPU 9; a prompt ten times as long takes at least ten times as long to read.
A longer window is not automatically a better one. Liu and colleagues found that language models used information best when it appeared at the beginning or end of a long input, and noticeably worse when it sat in the middle, even in models built for long contexts 6.
05
Why does an assistant ‘forget’ earlier messages?
Because the model keeps nothing between replies. The app sends the conversation back with each new message, and anything it leaves out, or anything that no longer fits, is invisible to the model.
A language model has no memory of its own from one request to the next: each reply is computed from what is in the window at that moment. Chat apps create the impression of memory by placing earlier turns back into the window, and once a conversation outgrows the window they have to drop or shorten something, usually the oldest turns.
Brello 1.0 carries the last six earlier messages into each reply. Your messages are clipped to 300 characters and Brello’s to 600, citation markers are removed and photos are noted as “[shared a photo]”. Each reply also opens a fresh model session, so web results from an earlier answer never crowd the window. The trade-off is that in a long chat, Brello no longer sees early details, even though the whole chat is still stored on your phone.
The instructions in the window are kept short as well. Our note on how small models copy the shape of their instructions explains why.
06
How does Brello 1.0 budget 4,096 tokens?
With a fixed limit for each part: a four-sentence system prompt, at most six clipped earlier messages, a character budget for web passages and an output cap of 1,200 tokens, or 2,048 with Think harder.
With every limit reached, these parts fill about 3,590 of Brello Pro’s 4,096 tokens (Figure 1). For how the web passages are chosen, see our note on answering from the web on the phone; for the three models and their limits, see the Brello models.
07
How is a context window different from memory?
A context window is working memory for a single reply, not long-term memory. What a model knows comes from its training; what it can use about you or your conversation is only what is placed in the window.
It helps to separate three things. Knowledge in the weights is learned during training, fixed afterwards and limited by a knowledge cutoff. The context window holds the text for one reply and is rebuilt for the next. Stored information, such as saved chats or documents, reaches the model only if the app retrieves it and places it in the window, which is the idea behind retrieval-augmented generation.
Brello 1.0 has no memory across chats. Each reply’s window holds its instructions, up to six earlier messages from the same chat and, when you search, passages from the web; nothing from other chats is included. Chats are stored only on your phone and are excluded from backups. For why small on-device models need this kind of care, see Small language models, explained; for how a longer output fits the same window, see Reasoning modes, explained.
References
Reviewed . Model cards, model pages and tokeniser files were checked on that date.
- Vaswani, A. et al. (2017). “Attention Is All You Need.” Advances in Neural Information Processing Systems 30 (NeurIPS 2017). arxiv.org/
abs/ 1706.03762 - Sennrich, R., Haddow, B. and Birch, A. (2016). “Neural Machine Translation of Rare Words with Subword Units.” Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL 2016). arxiv.org/
abs/ 1508.07909 - Kudo, T. and Richardson, J. (2018). “SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.” EMNLP 2018: System Demonstrations. arxiv.org/
abs/ 1808.06226 - Petrov, A., La Malfa, E., Torr, P. H. S. and Bibi, A. (2023). “Language Model Tokenizers Introduce Unfairness Between Languages.” Advances in Neural Information Processing Systems 36 (NeurIPS 2023). arxiv.org/
abs/ 2305.15425 - Ainslie, J. et al. (2023). “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.” EMNLP 2023. arxiv.org/
abs/ 2305.13245 - Liu, N. F. et al. (2024). “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the Association for Computational Linguistics 12. arxiv.org/
abs/ 2307.03172 - Google DeepMind. “Gemma 4 model card.” Google AI for Developers. Accessed 5 October 2026. ai.google.dev/
gemma/ docs/ core/ model_card_4 - Qwen. “Qwen3-1.7B.” Model card, configuration and tokeniser files, Hugging Face. Accessed 5 October 2026. huggingface.co/
Qwen/ Qwen3-1.7B - litert-community. “Qwen3-1.7B.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/
litert-community/ Qwen3-1.7B - litert-community. “gemma-4-E2B-it-litert-lm.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/
litert-community/ gemma-4-E2B-it-litert-lm - litert-community. “gemma-4-E4B-it-litert-lm.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/
litert-community/ gemma-4-E4B-it-litert-lm - Google. “gemma-4-E2B-it.” Model repository and tokeniser file, Hugging Face. Accessed 5 October 2026. huggingface.co/
google/ gemma-4-E2B-it
Version history
- 1.0First published.



