Summary
Brello Vision runs Gemma 4 E2B, accepts text and photos, downloads once at 2.59 GB and is recommended for phones with 6 GB of RAM.
| Tier | Balanced |
|---|---|
| Base model | Gemma 4 E2B by Google (Google’s model card) |
| Licence | Apache 2.0 |
| Inputs | Text and photos, one photo per message |
| Download | 2.59 GB (2,588,147,712 bytes) |
| Speed cache | ≈1.01 GB, written on first launch |
| Total storage | ≈3.60 GB |
| Free space required | ≈3.90 GB, including 300 MB of headroom |
| Recommended RAM | 6 GB; recommended when the phone reads at least 5.4 GB and less than 10.8 GB, or when its memory is unknown |
| Context window | 4,096 tokens |
| Maximum answer | 1,200 tokens; 2,048 with Think harder |
| Runtime | LiteRT-LM, on the GPU through OpenCL, with CPU fallback |
| Repository | litert-community/gemma-4-E2B-it-litert-lm on Hugging Face |
| File | gemma-4-E2B-it.litertlm |
What Brello Vision is for, and when Brello recommends it
Brello Vision is meant for phones with 6 GB of RAM or more, and it is the model Brello recommends for 6 GB and 8 GB phones and whenever it can’t read the phone’s memory.
It understands photos and text with a lighter download than Brello Pro: 2.59 GB instead of 3.66 GB, and about 3.60 GB of storage in total instead of 6.26 GB. A 6 GB phone reads about 5.6 GB and an 8 GB phone about 7.4 GB. Both clear Brello Vision’s 5.4 GB threshold but not Brello Pro’s 10.8 GB, so Brello Vision is the most capable model that fits them. On 12 GB phones Brello recommends Brello Pro, and on 4 GB phones Brello Core.
Brello 1.0 takes no files and has no voice input, and the model’s knowledge stops at a training cutoff, so questions about recent events need web search, which is off by default.
The base model: Gemma 4 E2B
Gemma 4 E2B is an open model in Google’s Gemma 4 family, built by Google DeepMind and released under Apache 2.0. Google’s model card, checked on 5 October 2026, gives the specifications in Table 2.1
| Parameters | 2.3B effective; 5.1B with embeddings |
|---|---|
| Layers | 35 |
| Context length | 128K tokens |
| Input | Text, image and audio |
| Output | Text |
| Training-data cutoff | January 2025 |
| Licence | Apache 2.0 |
The “E” stands for effective parameters. Per-Layer Embeddings give “each decoder layer its own small embedding for every token”, in tables that are large but used only for quick lookups, so Google counts 2.3B effective parameters against 5.1B in total.1
Brello 1.0 uses the model for text and photos only: Gemma 4 E2B also accepts audio, but Brello has no voice input or output. Benchmark results in Google’s card describe Gemma 4 E2B under Google’s own evaluation conditions. They are not measurements of Brello Vision, and Brello Research has not yet published its own.
The build Brello runs: mixed-precision weights
Brello downloads Gemma 4 E2B as one LiteRT-LM file, gemma-4-E2B-it.litertlm, from the litert-community/gemma-4-E2B-it-litert-lm repository on Hugging Face. The file is 2,588,147,712 bytes, is built from Google’s instruction-tuned checkpoint and downloads without an account or token.
According to the repository’s model card, the build uses a Gemma 4 mobile quantization scheme with “a mixture of 2bit, 4bit and 8 bit weights”. For text-only use, the card says, the weights can take as little as 0.8 GB of memory, while the runtime memory-maps the 1.12 GB of embedding parameters and loads the vision and audio models on demand.2 ‘Quantization, explained’ describes how lower-precision weights save memory.
The card states that the model “can support up to 32k context length”; Brello 1.0 configures a 4,096-token window. Its speed and memory figures are LiteRT community measurements on its own reference devices, taken with a 2,048-token context, not measurements of Brello on your phone.2
How Brello 1.0 runs it
Brello 1.0 runs Brello Vision with LiteRT-LM on the phone’s GPU through OpenCL, and falls back to the CPU automatically if the GPU path fails; Settings shows “Running on GPU” or “Running on CPU”.
While the model loads, the app shows “Starting Brello Vision”. On first load the runtime also writes the ≈1.01 GB speed cache, a weight cache tuned to the phone’s chip, which takes up to a minute, once. After that the model warms up while the app’s first frame draws, so it is usually ready by the time you type.
Each reply opens a fresh session with a short system prompt and up to six earlier messages, yours clipped to 300 characters and Brello’s to 600, within a 4,096-token window (‘Context windows, explained’). If you have chosen to search the web, up to 3,400 characters of passages from the pages Brello read are added; the search text goes directly from your phone to a search engine, which sees the request and your IP address as it would for any web request. A repetition stopper checks every 48 characters for a block repeated three or more times over at least 120 characters, and leaked markup such as <end_of_turn> and Gemma’s channel markers is removed or moved to the thought panel.
With Think harder on, the model reasons before it answers, the reasoning appears in a “Thought process” panel, and answers can run to 2,048 tokens.
Technical details (for reference)
| Temperature | 1.0 |
|---|---|
| Top-k | 64 |
| Top-p | 0.95 |
| Think harder | Temperature 0.6, top-p 0.95 |
| Context window | 4,096 tokens |
| Maximum answer | 1,200 tokens; 2,048 with Think harder |
| Web passages | Up to 3,400 characters |
| Acceleration | GPU through OpenCL; CPU fallback with an XNNPack weight cache |
Storage and memory
Brello Vision uses about 3.60 GB of storage once installed, and Brello recommends it for phones with 6 GB of RAM.
Before downloading, Brello checks for about 3.90 GB of free space: the 2.59 GB file, the ≈1.01 GB speed cache and 300 MB of headroom. If there is less, it shows “Not enough space” with the exact numbers. Removing Brello Vision in Settings frees the model and its cache.
The memory threshold is 5.4 GB, which is 0.9 × 6 GB. A 4 GB phone reads about 3.7 GB, below the threshold, so Brello recommends Brello Core there; if you choose Brello Vision anyway, Brello shows “Brello Vision may be too big”, explains that it “may run slowly or fail to start” and offers “Download anyway”. If it then crashes while loading, Brello switches to another installed model at the next launch and says so. Figure 1 on the models page applies the rule to any memory reading, and ‘Fitting a model to the phone in your pocket’ explains it.
Photos, understood offline
Brello Vision answers questions about photos on the phone, and a message with a photo never triggers a web search.
Attach one photo per message from the camera or the gallery through the + menu. The input bar shows a 72-pixel thumbnail and the hint “Ask about this photo”; if you send the photo without a question, Brello asks the model to “Describe this image in detail.” Photos are scaled to at most 1280 × 1280 pixels, at quality 88, and kept in Brello’s private storage with the chat.
If you are using Brello Core and tap Camera or Photos, Brello opens a sheet titled “Photos need Brello Vision”. It says that Brello Vision “understands images completely offline”, offers to switch to it if it is installed or to download it, and keeps Brello Core working during the download.
Gemma 4 E2B and E4B in Brello: what differs
Brello Vision and Brello Pro differ in the size of the base model, the download, the storage and the memory they need; inside Brello 1.0 they share the same inputs, settings, context window and answer limits.
| Property | Brello Vision | Brello Pro |
|---|---|---|
| Base model | Gemma 4 E2B | Gemma 4 E4B |
| Effective parameters | 2.3B | 4.5B |
| Parameters with embeddings | 5.1B | 8B |
| Layers | 35 | 42 |
| Download | 2.59 GB | 3.66 GB |
| Total storage | ≈3.60 GB | ≈6.26 GB |
| Recommended RAM | 6 GB | 12 GB |
| Tier in the app | Balanced | Most capable |
| Inputs in Brello | Text and photos | Text and photos |
| Settings, context and answer length | Same | Same |
Brello Research has not measured how their answers compare, so this page makes no claim about quality beyond the app’s own tier labels, “Balanced” and “Most capable”.
Limitations
Brello Vision is smaller than Brello Pro and, like any on-device model, far smaller than frontier cloud models: it can be wrong.
- It needs a phone with 6 GB of RAM or more. That is its recommended size, and it needs about 3.90 GB free before download.
- The first launch is slow. Building the speed cache takes up to a minute, once.
- It makes mistakes. Its knowledge stops at Gemma 4’s training-data cutoff of January 2025,1 smaller models are less capable than larger ones, and the system prompt’s instruction to admit uncertainty is no guarantee.
- Its memory of a chat is short. Each answer sees only the last six messages, clipped.
- Its inputs are limited. One photo per message, no files or documents, no voice, and an English-only interface.
Brello Research has not published benchmarks for Brello Vision. The Brello 1.0 system card lists the known limitations of the whole app.
Source files and licence
Brello Vision is built from one file, gemma-4-E2B-it.litertlm (2,588,147,712 bytes), from the litert-community/gemma-4-E2B-it-litert-lm repository on Hugging Face.
Google releases Gemma 4 E2B under the Apache 2.0 licence, and the repository carries the same licence. Brello credits it on the licences page. Brello isn’t affiliated with or endorsed by Google.
Version history
- Included in the first build of Brello 1.0, based on Gemma 4 E2B, with photo understanding.
- Labelled “Balanced” when the three models were given tiers; licence label corrected to Apache 2.0.
- Model card published.
Sources
Facts about the base model and the build come from their publishers’ pages, checked on 5 October 2026. Facts about Brello describe Brello 1.0, version 1.0.0 (build 1).
- Google. “Gemma 4 model card.” Google AI for Developers. ai.google.dev/
gemma/ . Checked 5 October 2026.docs/ core/ model_card_4 - litert-community. “gemma-4-E2B-it-litert-lm.” Model repository, Hugging Face. huggingface.co/
litert-community/ . Checked 5 October 2026.gemma-4-E2B-it-litert-lm