Summary
Brello Pro runs Gemma 4 E4B, accepts text and photos, downloads once at 3.66 GB and is recommended for phones with 12 GB of RAM.
| Tier | Most capable |
|---|---|
| Base model | Gemma 4 E4B by Google (Google’s model card) |
| Licence | Apache 2.0 |
| Inputs | Text and photos, one photo per message |
| Download | 3.66 GB (3,659,530,240 bytes) |
| Speed cache | ≈2.60 GB, written on first launch |
| Total storage | ≈6.26 GB |
| Free space required | ≈6.56 GB, including 300 MB of headroom |
| Recommended RAM | 12 GB; recommended when the phone reads at least 10.8 GB |
| Context window | 4,096 tokens |
| Maximum answer | 1,200 tokens; 2,048 with Think harder |
| Runtime | LiteRT-LM, on the GPU through OpenCL, with CPU fallback |
| Repository | litert-community/gemma-4-E4B-it-litert-lm on Hugging Face |
| File | gemma-4-E4B-it.litertlm |
What Brello Pro is for
Brello Pro is meant for phones with 12 GB of RAM or more, for people who want the most capable of Brello’s three models, including for questions about photos.
Brello recommends it when the phone reads at least 10.8 GB of memory, which in practice means 12 GB phones: they report about 11.2 GB. On a phone with less memory, Brello recommends Brello Vision instead, which also understands photos and is recommended for 6 GB. Brello 1.0 takes no files and has no voice input, and the model’s knowledge stops at a training cutoff, so questions about recent events need web search, which is off by default.
The base model: Gemma 4 E4B
Gemma 4 E4B is an open model in Google’s Gemma 4 family, built by Google DeepMind and released under Apache 2.0. Google’s model card, checked on 5 October 2026, gives the specifications in Table 2.1
| Parameters | 4.5B effective; 8B with embeddings |
|---|---|
| Layers | 42 |
| Context length | 128K tokens |
| Input | Text, image and audio |
| Output | Text |
| Training-data cutoff | January 2025 |
| Licence | Apache 2.0 |
The “E” stands for effective parameters. Google explains that Per-Layer Embeddings give “each decoder layer its own small embedding for every token”, and that these tables “are large but are only used for quick lookups, which is why the effective parameter count is much smaller than the total”.1
Brello 1.0 uses the model for text and photos only: Gemma 4 E4B also accepts audio, but Brello has no voice input or output. Benchmark results in Google’s card describe Gemma 4 E4B under Google’s own evaluation conditions. They are not measurements of Brello Pro, and Brello Research has not yet published its own.
The build Brello runs
Brello downloads Gemma 4 E4B as one LiteRT-LM file, gemma-4-E4B-it.litertlm, from the litert-community/gemma-4-E4B-it-litert-lm repository on Hugging Face. The file is 3,659,530,240 bytes, is built from Google’s instruction-tuned checkpoint and downloads without an account or token.
The repository’s model card says the file “includes a text decoder with 2.24 GB of weights and 0.67 GB of embedding parameters”. LiteRT-LM keeps the main weights in memory and memory-maps the embedding parameters, and it loads the vision and audio models only as needed.2
The same card states that the model “can support up to 32k context length”; Brello 1.0 configures a 4,096-token window. The card also publishes speed and memory figures for this file. Those are LiteRT community measurements on its own reference devices, taken with a 2,048-token context, not measurements of Brello on your phone.2
How Brello 1.0 runs it
Brello 1.0 runs Brello Pro with LiteRT-LM on the phone’s GPU through OpenCL, and falls back to the CPU automatically if the GPU path fails; Settings shows “Running on GPU” or “Running on CPU”.
On first load the runtime writes the ≈2.60 GB speed cache, a weight cache tuned to the phone’s chip, which takes up to a minute, once. After that the model warms up while the app’s first frame draws, so it is usually ready by the time you type. A status dot in the top bar shows its state: green when ready, amber while starting and red if it failed.
Each reply opens a fresh session with a short system prompt and up to six earlier messages, yours clipped to 300 characters and Brello’s to 600, within a 4,096-token window (‘Context windows, explained’). If you have chosen to search the web, up to 3,400 characters of passages from the pages Brello read are added; the search text goes directly from your phone to a search engine, which sees the request and your IP address as it would for any web request. Two guards watch the output: a repetition stopper that checks every 48 characters for a block repeated three or more times over at least 120 characters, and a clean-up that removes leaked markup such as <end_of_turn>, <eos> and Gemma’s channel markers, or moves it to the thought panel.
With Think harder on, the model reasons before it answers, the reasoning appears in a “Thought process” panel, and answers can run to 2,048 tokens. ‘Reasoning modes, explained’ covers the technique.
Technical details (for reference)
| Temperature | 1.0 |
|---|---|
| Top-k | 64 |
| Top-p | 0.95 |
| Think harder | Temperature 0.6, top-p 0.95 |
| Context window | 4,096 tokens |
| Maximum answer | 1,200 tokens; 2,048 with Think harder |
| Web passages | Up to 3,400 characters |
| Acceleration | GPU through OpenCL; CPU fallback with an XNNPack weight cache |
Storage and memory
Brello Pro uses about 6.26 GB of storage once installed, and Brello recommends it for phones with 12 GB of RAM.
Before downloading, Brello checks for about 6.56 GB of free space: the 3.66 GB file, the ≈2.60 GB speed cache and 300 MB of headroom. If there is less, it shows “Not enough space” with the exact numbers. Removing Brello Pro in Settings frees the model and its cache.
Brello recommends Brello Pro only when the phone reads at least 10.8 GB, which is 0.9 × 12 GB. A 12 GB phone reads about 11.2 GB and qualifies; an 8 GB phone reads about 7.4 GB and gets Brello Vision instead. You can still choose Brello Pro on a smaller phone: Brello shows “Brello Pro may be too big”, explains that it “may run slowly or fail to start” and offers “Download anyway”. If it then crashes while loading, Brello switches to another installed model at the next launch and says so. Figure 1 on the models page applies the rule to any memory reading, and ‘Fitting a model to the phone in your pocket’ explains it.
Photos, understood offline
Brello Pro answers questions about photos on the phone, and a message with a photo never triggers a web search.
Attach one photo per message from the camera or the gallery through the + menu. The input bar shows a 72-pixel thumbnail and the hint “Ask about this photo”; if you send the photo without a question, Brello asks the model to “Describe this image in detail.” Photos are scaled to at most 1280 × 1280 pixels, at quality 88, and kept in Brello’s private storage with the chat. Tapping one opens a full-screen viewer with pinch to zoom.
Brello Vision also understands photos, with a 2.59 GB download. Brello Core reads text only.
Limitations
Brello Pro is the largest of the three models and, like any on-device model, far smaller than frontier cloud models: it can be wrong.
- It needs a high-end phone. It is recommended only for 12 GB phones, needs about 6.56 GB free before download and is the largest download of the three, so some phones won’t have room for it.
- The first launch is slow. Building the speed cache takes up to a minute, once.
- It makes mistakes. Its knowledge stops at Gemma 4’s training-data cutoff of January 2025,1 it is weaker at long or complex reasoning than frontier models, and the system prompt’s instruction to admit uncertainty is no guarantee.
- Its memory of a chat is short. Each answer sees only the last six messages, clipped.
- Its inputs are limited. One photo per message, no files or documents, no voice, and an English-only interface.
Brello Research has not published benchmarks for Brello Pro. The Brello 1.0 system card lists the known limitations of the whole app.
Source files and licence
Brello Pro is built from one file, gemma-4-E4B-it.litertlm (3,659,530,240 bytes), from the litert-community/gemma-4-E4B-it-litert-lm repository on Hugging Face.
Google releases Gemma 4 E4B under the Apache 2.0 licence, and the repository carries the same licence. Brello credits it on the licences page. Brello isn’t affiliated with or endorsed by Google.
Version history
- Brello Pro added to Brello 1.0 as the “Most capable” tier, with a memory-based recommendation, a warning for models that may be too big and a free-space check before download.
- Model card published.
Sources
Facts about the base model and the build come from their publishers’ pages, checked on 5 October 2026. Facts about Brello describe Brello 1.0, version 1.0.0 (build 1).
- Google. “Gemma 4 model card.” Google AI for Developers. ai.google.dev/
gemma/ . Checked 5 October 2026.docs/ core/ model_card_4 - litert-community. “gemma-4-E4B-it-litert-lm.” Model repository, Hugging Face. huggingface.co/
litert-community/ . Checked 5 October 2026.gemma-4-E4B-it-litert-lm