Summary
Brello Core runs Qwen3 1.7B, reads text only, downloads once at 977 MB and is recommended for phones with 4 GB of RAM.
| Tier | Fastest |
|---|---|
| Base model | Qwen3 1.7B by Alibaba (Qwen’s model card) |
| Licence | Apache 2.0 |
| Input | Text |
| Download | 977 MB (977,184,032 bytes) |
| Speed cache | ≈974 MB, written on first launch |
| Total storage | ≈1.95 GB |
| Free space required | ≈2.25 GB, including 300 MB of headroom |
| Recommended RAM | 4 GB; recommended when the phone reads less than 5.4 GB |
| Context window | 4,096 tokens |
| Maximum answer | 1,200 tokens; 2,048 with Think harder |
| Runtime | LiteRT-LM, on the GPU through OpenCL, with CPU fallback |
| Repository | litert-community/Qwen3-1.7B on Hugging Face |
| File | Qwen3-1.7B_dynamic_wi4b32_afp32.litertlm |
What Brello Core is for
Brello Core is meant for text questions on phones with 4 GB of RAM or more, where a small download and quick answers matter more than photo understanding.
Brello recommends it when the phone reads less than 5.4 GB of memory: a 4 GB phone reads about 3.7 GB, and phones below every model’s threshold also get Brello Core. On 6 GB and 8 GB phones Brello recommends Brello Vision, and on 12 GB phones Brello Pro. On larger phones you can still choose Brello Core in Settings, for its 977 MB download and ≈1.95 GB footprint.
Brello Core can’t answer questions about photos, which need Brello Vision or Brello Pro. Brello 1.0 takes no files and has no voice input, and the model’s knowledge stops at a training cutoff, so questions about recent events need web search, which is off by default.
The base model: Qwen3 1.7B
Qwen3 1.7B is an open model in the Qwen3 family from the Qwen team at Alibaba, released under Apache 2.0. Qwen’s model card, checked on 5 October 2026, gives the specifications in Table 2,1 and the Qwen3 Technical Report describes the family.3
| Parameters | 1.7B; 1.4B excluding embeddings |
|---|---|
| Layers | 28 |
| Attention heads | 16 for queries, 8 for keys and values (grouped-query attention) |
| Context length | 32,768 tokens |
| Modes | Thinking and non-thinking, in one model |
| Languages | More than 100 languages and dialects |
| Licence | Apache 2.0 |
Qwen describes a thinking mode, for complex logical reasoning, maths and coding, and a non-thinking mode, for efficient general-purpose dialogue, both within the same model.1 Qwen publishes benchmark evaluations for Qwen3 in its own materials. They describe the base models under Qwen’s conditions, not Brello Core, and Brello Research has not yet published its own.
The build Brello runs: dynamic INT4 weights
Brello downloads Qwen3 1.7B as one LiteRT-LM file, Qwen3-1.7B_dynamic_wi4b32_afp32.litertlm, from the litert-community/Qwen3-1.7B repository on Hugging Face. The file is 977,184,032 bytes and downloads without an account or token.
The repository’s model card calls it “a dynamic INT4 variant (block-32 weights, FP32 activations)”, made with the quantization recipe dynamic_wi4b32_afp32 that the file name records. It was converted through LiteRT Torch and quantised with AI Edge Quantizer, and its graph uses “composite ops for RoPE, fused QKV, and fused Gate/Up projections”.2 The same repository holds a larger 8-bit variant of 2.1 GB; Brello uses the INT4 file. ‘Quantization, explained’ describes how 4-bit weights save memory.
The card lists the INT4 variant with a 4,096-token context, the window Brello 1.0 configures, and gives its size as 932 MB. That is the same 977,184,032 bytes counted in binary units (1 MiB = 1,048,576 bytes). Brello and this page use decimal units, so it appears here as 977 MB.2
How Brello 1.0 runs it
Brello 1.0 runs Brello Core with LiteRT-LM on the phone’s GPU through OpenCL, and falls back to the CPU automatically if the GPU path fails; Settings shows “Running on GPU” or “Running on CPU”. On first load the runtime writes the ≈974 MB speed cache, which takes up to a minute, once.
Brello samples Brello Core at temperature 0.7, top-k 20 and top-p 0.8, the settings Qwen suggests for non-thinking mode. With Think harder on, it uses temperature 0.6 and top-p 0.95, which match the temperature and top-p Qwen suggests for thinking mode, and the reasoning appears in a “Thought process” panel.1 Qwen’s card warns that greedy decoding “can lead to performance degradation and endless repetitions”; Brello samples rather than decoding greedily, and also runs a repetition stopper that checks every 48 characters for a block repeated three or more times over at least 120 characters. Leaked markup such as <think>, <|im_end|> and /no_think is removed or moved to the thought panel.
Each reply opens a fresh session with a short system prompt and up to six earlier messages, yours clipped to 300 characters and Brello’s to 600. The prompt is short on purpose, because small models copy the shape of their instructions. If you have chosen to search the web, up to 2,600 characters of passages are added, less than the 3,400 used for the Gemma 4 models; the search text goes directly from your phone to a search engine, which sees the request and your IP address as it would for any web request.
Technical details (for reference)
| Setting | Brello Core | Qwen’s suggestion |
|---|---|---|
| Temperature | 0.7 | 0.7, non-thinking mode |
| Top-k | 20 | 20 |
| Top-p | 0.8 | 0.8, non-thinking mode |
| Think harder | Temperature 0.6, top-p 0.95 | Temperature 0.6, top-p 0.95, thinking mode |
| Context window | 4,096 tokens | 32,768 tokens (base model) |
| Maximum answer | 1,200 tokens; 2,048 with Think harder | 32,768 tokens for most queries |
| Web passages | Up to 2,600 characters | Not applicable |
Storage and memory
Brello Core uses about 1.95 GB of storage once installed, the least of the three models, and Brello recommends it for phones with 4 GB of RAM.
Before downloading, Brello checks for about 2.25 GB of free space: the 977 MB file, the ≈974 MB speed cache and 300 MB of headroom. If there is less, it shows “Not enough space” with the exact numbers. Removing Brello Core in Settings frees the model and its cache.
Brello Core’s memory threshold is 3.6 GB, which is 0.9 × 4 GB. A 4 GB phone reads about 3.7 GB and fits. Below 3.6 GB no model fits, and Brello still recommends Brello Core as the smallest. Figure 1 on the models page applies the rule to any memory reading.
Text only: photos need Brello Vision or Brello Pro
Brello Core reads text only, so questions about photos need Brello Vision or Brello Pro.
If you tap Camera or Photos while using Brello Core, Brello opens a sheet titled “Photos need Brello Vision” (or Brello Pro). It offers to switch to a vision model that is already installed, or to download one, and notes that “your current model keeps working meanwhile”: Brello Core keeps answering while the download runs. Only one download runs at a time, and switching between installed models in Settings is instant.
Limitations
Brello Core is the smallest of the three models and, like any on-device model, far smaller than frontier cloud models: it can be wrong.
- It is the least capable of the three. Smaller models are faster but less capable, and Brello Core is the smallest. Like the others, it is weaker at long or complex reasoning than frontier cloud models.
- It makes mistakes. Its knowledge stops at a training cutoff, and the system prompt’s instruction to admit uncertainty is no guarantee.
- It cannot see photos. Photo questions need Brello Vision or Brello Pro.
- Its memory of a chat is short. Each answer sees only the last six messages, clipped, and answers stop at 1,200 tokens, or 2,048 with Think harder.
- Languages are not evaluated. Qwen lists more than 100 languages and dialects for the base model, but Brello’s interface is in English only and Brello has not evaluated other languages.
Brello Research has not published benchmarks for Brello Core. The Brello 1.0 system card lists the known limitations of the whole app.
Source files and licence
Brello Core is built from one file, Qwen3-1.7B_dynamic_wi4b32_afp32.litertlm (977,184,032 bytes), from the litert-community/Qwen3-1.7B repository on Hugging Face.
The Qwen team at Alibaba releases Qwen3 1.7B under the Apache 2.0 licence, and the repository carries the same licence. Brello credits it on the licences page. Brello isn’t affiliated with or endorsed by Alibaba.
Version history
- Included in the first build of Brello 1.0, based on Qwen3 1.7B.
- Labelled “Fastest” when the three models were given tiers; licence label corrected to Apache 2.0.
- Model card published.
Sources
Facts about the base model and the build come from their publishers’ pages, checked on 5 October 2026. Facts about Brello describe Brello 1.0, version 1.0.0 (build 1).
- Qwen. “Qwen3-1.7B.” Model card, Hugging Face. huggingface.co/
Qwen/ . Checked 5 October 2026.Qwen3-1.7B - litert-community. “Qwen3-1.7B.” Model repository, Hugging Face. huggingface.co/
litert-community/ . Checked 5 October 2026.Qwen3-1.7B - Qwen Team. “Qwen3 Technical Report.” arXiv:2505.09388, 2025. arxiv.org/
abs/ .2505.09388