Our approach: evaluate first, then ship
Before Brello can do something new, we will test how it could fail or be misused, hold the release until it meets criteria written before testing began, and publish what we found. This is commitment 04 of the Brello Charter, version 1.0, and it covers Brello Super Intelligence and every Brello release after 1.0.0. Brello 1.0, version 1.0.0, predates the charter and has no published evaluation. It has the safeguards below, each with its real parameters, and none is a guarantee.
- In Brello 1.0
- Shipped in version 1.0.0, 4 October 2026.
- Planned
- Design intent for Brello SI. No results exist yet.
Safeguards in Brello 1.0 today
In Brello 1.0
Brello 1.0 runs eight safeguards. Six are checks written in code; two are instructions to the model, which it may not follow.
| Safeguard | What it does | Kind |
|---|---|---|
| Uncertainty instruction | The system prompt is short on purpose, because small models copy the shape of long instructions. It ends: “If you are not sure about something, say so instead of guessing.” | Instruction |
| Inline citations | With web results, the model is told to “cite the sources you rely on inline using their numbers, like [1] or [2]”. Each citation in the answer becomes a link to its source page. | Instruction, with links in code |
| Asking before searching | Web search is off by default. With it off, a question that looks time-sensitive pauses at “Needs the web” and shows the “Search the web for this?” card. A question with a photo never searches. | Check in code |
| Repetition stopper | Every 48 characters, Brello checks whether the end of the reply is a block repeated 3 or more times over at least 120 characters. If it is, Brello cuts the reply after the first copy and stops the model. | Check in code |
| Leaked-markup clean-up | Stray <think>, <|im_end|>, <end_of_turn>, <eos>, /no_think and Gemma “channel” markers are removed, or routed to the “Thought process” panel. | Check in code |
| Empty replies | If the model reasoned but never answered while Think harder was off, the reasoning becomes the answer. A reply with no text shows “No response was generated. Try rephrasing.” | Check in code |
| Storage and memory checks | Before a download, Brello requires free space for the model, its speed cache and 300 MB of headroom, or shows “Not enough space” with exact numbers. On a phone with less memory than the model needs, it warns: “It may run slowly or fail to start.” You can still choose “Download anyway”. | Check in code |
| Crash-loop recovery | If a model crashes the app while loading, usually by running out of memory, Brello remembers. On the next launch it switches to another installed model and says why: “{Model} couldn't start on this phone · Using {Other}”. | Check in code |
Figure 1 follows one reply through the safeguards that act on its text: the two instructions, the repetition stopper and the markup clean-up.
- The reply starts from a short instruction. The model is asked to admit uncertainty and to cite web results, but it can ignore both, so the next steps are checks in code.
- The reply is checked as it streams. Every 48 characters, Brello looks for a block repeated 3 or more times over 120 or more characters. Four checks pass.
- Leaked markup is removed. Control tokens such as
<think>or<end_of_turn>are stripped or routed to the Thought process panel, so they never reach the answer. - A loop is caught at the next check. One 49‑character sentence has now appeared three times, so Brello cuts the reply after the first copy and stops the model.
- You see the cleaned reply, with its sources. Source cards sit above the answer, and each citation links to its page. The checks make a reply tidier, not correct.
None of these makes a small model correct. On-device models are far smaller than frontier cloud models: they can be wrong, have a knowledge cutoff and are weaker at long or complex reasoning. Why AI makes things up explains the underlying problem, and section 2 of Evaluate first, then ship maps where each safeguard acts in the life of a reply.
How we intend to evaluate Brello Super Intelligence
Planned
We intend to evaluate Brello SI in six areas, each re-tested with every release. Because Brello is designed so that we don’t see conversations, the evidence has to be gathered before release, from task sets, synthetic profiles and attacks written for the purpose.
| Area | The question it answers | How we intend to test it |
|---|---|---|
| Capability on real tasks | Does it do the work people bring to it, at the quality it implies? | Task sets built around research with sources, writing and multi-step tasks, written for the purpose and never taken from conversations. Checked automatically where possible, and by people otherwise. |
| Honesty and calibration | Does it say it isn’t sure when it should, and only then? | Comparing its confidence with its accuracy, and checking that each citation supports the sentence it is attached to. |
| Privacy leakage | Does it reveal personal information it shouldn’t? | Privacy audits that seed synthetic profiles with made-up personal details, then search every output and every outbound request for them. |
| Misuse and harm | Could someone use it to hurt other people, or themselves? | Red-teaming by people and by language models, before release. |
| Indirect prompt injection | Does it follow instructions planted in what it reads? | Pages seeded with instructions, with any change in behaviour counted as a failure. |
| Safety of actions | Does every consequential action wait for you? | Checking that the confirmation appears every time, shows exactly what will happen and can’t be bypassed by any phrasing or web page. |
When we publish a result, it will give the method, the sample size, the date and which direction is better. AI safety evaluations, explained introduces these methods in general terms, and section 3 of Evaluate first, then ship sets out each area in full.
Release gates
Planned
Each Brello SI release, and each new capability within one, is intended to pass four gates in order, with pass criteria written before testing starts. A failure at any gate will send the release back to the first gate, not to the one it failed, because a fix for one problem can cause another.
| Gate | To pass, the release will need |
|---|---|
| 1 Internal evaluations | Automated suites and human review across all six areas, with every threshold met, no unexplained regression and no hard stop. |
| 2 Adversarial red-teaming | Every finding from people and automated methods trying to break it resolved: fixed, mitigated, or accepted as a known limit with a written reason. |
| 3 Invited early access | A small group who know they are using an early system, with a clear way to report problems. Serious reports resolved and known limits written down. |
| 4 Wider availability | Its evaluation summary published. Reports from people using Brello SI and from independent researchers continue to be reviewed. |
The criteria will be of three types: a threshold, the minimum acceptable result in an area; no regressions, meaning nothing gets worse than in the previous release without a stated reason; and hard stops, failures that block a release however good everything else is. Completing a consequential action without your confirmation will be a hard stop. Because the criteria are fixed before testing, the bar can’t drift towards whatever a release already does.
Indirect prompt injection: Brello 1.0 and the Brello SI design
Brello SI is being designed to treat text from outside as data, never as instructions. We intend to make that hold with four layers, because no single defence against indirect prompt injection is reliable alone.
| Layer | In Brello 1.0 | Planned for Brello SI |
|---|---|---|
| Label retrieved text as data | Web passages reach the model as a numbered block labelled “Web results (retrieved {date})”, each with its source number. | The same: each passage in a marked, numbered block. |
| Keep data out of the instruction channel | Not in Brello 1.0. | Prompt assembly designed so retrieved text can’t reach the place instructions go. |
| Let people check | Inline citations link each claim to the page it came from. | The same, so a distorted answer can be traced to the page that caused it. |
| Confirm before acting | Brello 1.0 can only write an answer. | Sending, spending or deleting will wait for your confirmation, a check designed to sit outside the model. |
In Brello 1.0 a page cannot make Brello take an action, because Brello 1.0 can only write an answer. A page can still distort the answer that draws on it and mislead the person who reads it. That answer, clipped to 600 characters, stays in the context of later replies in the same chat for as long as it is among the six most recent messages. The web passages themselves are used for that reply only: each reply opens a fresh model session. No defence against prompt injection is complete, and we don’t claim these layers remove the risk. Prompt injection, explained introduces the attack, and section 4 of Evaluate first, then ship gives the design in full.
Confirmation before consequential actions: Brello 1.0 and the Brello SI design
Brello SI is being designed so that it will not complete an action that spends money, sends a message on your behalf or deletes something until you have confirmed it, shown in plain terms. This is commitment 05 of the charter.
| Action | Brello 1.0 | Planned for Brello SI |
|---|---|---|
| Spending money | Not possible. | Will wait until you have seen the amount and confirmed it. |
| Sending a message on your behalf | Not possible. | Will wait until you have seen who receives what, and confirmed it. |
| Deleting | Swiping left deletes one chat. “Delete all chats” asks first: “This permanently removes every conversation from this device.” | Will wait until you have seen what will be deleted, and confirmed it. |
We are designing the confirmation as a check outside the model, so that a model that has been misled could not complete the action on its own. A confirmation protects only people who read it, and one that appears too often invites approval by habit, so how often it appears is part of what we intend to evaluate.
Saying when it isn’t sure
Brello 1.0 tells its model to admit uncertainty and to cite its web sources, and we intend to evaluate whether Brello SI does both. This is commitment 06 of the charter.
| Instruction | Text in the system prompt | When it is included |
|---|---|---|
| Uncertainty | “If you are not sure about something, say so instead of guessing.” | Every reply |
| Citations | “Base your answer on them and cite the sources you rely on inline using their numbers, like [1] or [2]. If the results do not answer the question, say so briefly and answer from general knowledge.” | Replies that use web results |
An instruction is not a guarantee, and a small model can still be wrong. For Brello SI, each evaluation summary will report how often it says it isn’t sure when it should.
What we will publish
We will publish an evaluation summary with each Brello SI release and with each new capability in any Brello release, and we will publish Brello SI’s architecture before anyone outside the team uses it.
| Publication | What it contains | Status |
|---|---|---|
| Evaluation summary, with each release | What was tested and how, the pass criteria set before testing, the results including failures, and what the release still gets wrong. Anything withheld, such as a test that would work as instructions for an attack, is named with the reason. | None yet: no Brello SI release exists. |
| Brello SI architecture | What runs where, and what each layer can and cannot see. | Before anyone outside the team uses Brello SI (Charter, section 4). |
| Brello 1.0 system card | What the app does, the requests it makes and the permissions it holds. | Published |
| Research notes | Our methods and designs, with references. | Published: Evaluate first, then ship; Asking before going online |
Reporting a vulnerability
Report safety and security issues through our security and disclosure page, which sets out what to include and the safe harbour for research done in good faith. A machine-readable contact is published at /.well-known/security.txt.
| If you find | Do this |
|---|---|
| A vulnerability in Brello 1.0 or this website | Email our security contact with “Security report” in the subject. |
| A request from Brello 1.0 that our privacy page doesn’t list | Report it as a security issue, the same way. |
| A safeguard on this page that doesn’t behave as described | Report it the same way, with the steps that reproduce it. |
Further reading
- Evaluate first, then ship: the full evaluation plan for Brello SI.
- The Brello Charter: commitments 04, 05 and 06.
- Brello 1.0 system card: the technical account of the app.
- Asking before going online: consent for web search
- AI safety evaluations, explained, Prompt injection, explained and Why AI makes things up
- How Brello handles your questions, photos and answers
Changelog
- Version 1.0. First published, for Brello 1.0, version 1.0.0.