Privacy

Where AI assistants process and keep your questions

Where a prompt is processed, what is kept and for how long, who can read it, whether it trains models, and how to judge an assistant’s privacy claims.

Brello Research12 min readVersion 1.0

Summary

An AI assistant’s privacy comes down to a few concrete facts: where the model processes your prompt, what the service stores and for how long, whether conversations train models, who can read them, and what an account links them to. Cloud assistants send each prompt to the provider’s servers, where its policies and your settings decide retention, training and human review. On-device assistants process the prompt on the phone, so the device rather than a policy answers most of these questions, although web search still sends the search text and page requests to third parties. This explainer covers each question in general terms and ends with a checklist for judging privacy claims; our comparison pages quote what each provider says.

  • Where a prompt is processed decides who receives it: a cloud assistant sends every prompt to the provider’s servers, which decrypt it so the model can process it, while an on-device assistant processes it on the phone.
  • TLS encryption protects a prompt on its way to a cloud assistant, not at its destination: there, the provider’s policy and your settings decide how long it is kept, whether people review it and whether it trains models.
  • Training is hard to undo: a trained model cannot delete a conversation it learned from, and Carlini and colleagues extracted hundreds of verbatim sequences, including names and phone numbers, from GPT-2’s training data.
  • Android backs up app data to the user’s Google Drive by default, so an app that keeps chats on one phone has to opt out of cloud backup and of device-to-device transfer.
  • Brello 1.0 runs its model on the phone with no account and no server; with web search on, only the search text and page requests leave the phone, and those sites see a normal web request.
Contents10 sections

01

What happens to your prompt when you use an AI assistant?

When you use an AI assistant, your prompt is processed by a language model either on the provider’s servers or on your own device. Where it runs decides who receives your words; the provider’s policies and your settings then decide how long they are kept, whether they train future models and whether people may read them.

Five questions decide what that means in practice: where the model that reads your prompt runs, what is kept afterwards and for how long, whether it is used to train models, who can read it, and what your account connects it to. A sixth applies whenever an assistant uses the web: what the search engines and sites it contacts receive.

This explainer takes each question in turn and ends with a checklist for judging any assistant’s privacy claims. It describes how assistants work in general rather than any one provider; what each provider states about its own assistant, quoted and dated, is on our comparison pages. It does not rank assistants; the right trade-offs depend on what you ask.

02

Where is your prompt processed?

Most AI assistants process prompts on the provider’s servers, because the largest language models need data-centre hardware to run. A smaller number run a model on the phone itself, so the prompt is processed on the device it was typed on.

In the cloud case, the app sends your prompt to the provider, usually with recent turns of the conversation and any attached files or photos. The connection is encrypted with TLS, which stops anyone on the network path from reading or altering the message 1. TLS protects the message in transit, not at its destination: the provider’s systems decrypt the prompt so the model can read it, and from that point the provider’s policies and your settings govern what happens to it (sections 03 to 05).

On a device, the model’s weights are downloaded once and the computation, called inference, runs on the phone’s own processors. The prompt goes only to the local runtime. Android offers this as a system service: Google states that Gemini Nano “lets you deliver rich generative AI experiences without needing a network connection or sending data to the cloud” 2. Apps can also download and run their own models, as our guide to running a language model on an Android phone explains. The trade-off is size: a phone-sized model knows less and reasons less reliably than a frontier cloud model (see small language models, explained).

Some systems mix the two, answering simple requests on the device and sending harder ones to a server. Others aim to make server-side processing checkable by running it in hardware-isolated environments whose software can be verified remotely, the subject of our explainer on confidential computing. Either way, the useful question is which computer reads the prompt for each feature. On-device AI, explained covers the architecture in more depth, and Figure 1 follows one question through both designs.

Where a prompt is processed: a cloud assistant and an on-device assistant Two lanes. In the top lane, a cloud assistant sends the question “What does a high ALT reading mean?” from the phone over an encrypted internet connection to the provider’s servers, where the model decrypts and reads it. The chat is then saved with the user’s account, retention is set by policy and settings, people may read some chats, and chats may train future models. In the bottom lane, Brello 1.0, an on-device assistant, sends the same question only to the model on the phone, which runs on the GPU or CPU, and stores the chat on the phone, excluded from backups. Only if web search is on, the search text goes from the phone to a search engine, which receives it with the phone’s IP address, and up to four result pages are requested from their sites. Cloud assistant Your phone What does a high ALT reading mean? Sends the prompt and recent turns of the chat Encrypted in transit (TLS) Over the internet Provider’s servers Model Decrypts and reads the prompt History Retention Review Training Saved with your account Set by policy and settings People may read some chats May train future models On-device assistant · Brello 1.0 On this phone No Brello server What does a high ALT reading mean? Stored On the phone, excluded from backups Model Runs on the phone’s GPU or CPU What does a high ALT reading mean Search text Search engine Only if web search is on Not contacted Receives the search text and the phone’s IP address Sites: up to 4 page requests
  1. The same question goes to two kinds of assistant. What happens to it next depends first on where the model that reads it runs.
  2. A cloud assistant sends the prompt to the provider. TLS encryption protects it on the way, and the provider’s servers decrypt it so the model can read it.
  3. From there, policy decides. The chat is saved with your account, and the provider’s terms and your settings decide how long it is kept, whether people review it and whether it trains models.
  4. An on-device assistant answers on the phone. The prompt goes only to the local model, and Brello 1.0 keeps the chat in private app storage that is excluded from backups.
  5. Web search is the exception. With it switched on, or approved for one question, Brello 1.0 sends only the search text to a search engine, which sees it and the phone’s IP address, as for any web request.
Figure 1Where a prompt is processed decides who receives it. The cloud lane is general: retention, review and training differ by provider and setting (sections 03 to 05). The on-device lane shows Brello 1.0, which reads up to four result pages when it searches; the question is illustrative.

An on-device model changes who receives a prompt, not whether the app ever uses the network. It may still download models, check for updates, send analytics or search the web, which is why the checklist in section 08 asks what leaves the device besides the prompt.

03

What is stored, and for how long?

A cloud assistant usually keeps your conversations with your account until you delete them or an auto-delete period ends, and may keep some copies for longer, for safety review, legal reasons or model training.

Defaults differ by provider, by product and by setting. Some assistants keep your history until you delete it, others delete it automatically after a set period that you can often change, and some offer a temporary or incognito chat that does not appear in your history. Our comparison pages quote each provider’s own figures, with the date we checked them.

Deleting a chat is not always the end of it. It helps to separate three kinds of copy: the history you can see, the provider’s back-end copy, which can take some time to be removed after you delete a chat, and copies made for other purposes such as safety review, feedback or training, which follow their own schedules (sections 04 and 05).

On a phone, the question becomes whether the app copies its files anywhere else. Android backs up app data to the user’s Google Drive by default, and each app can opt out 3. Opting out is not quite the whole answer: Android’s documentation notes that for apps targeting Android 12 or later, on some manufacturers’ devices, turning backup off this way still allows device-to-device transfers 3. An app that wants chats to stay on one phone has to exclude them from transfers as well.

04

Are your conversations used to train models?

It depends on the provider, the product and your settings. Some consumer assistants use conversations to train models unless you opt out, others only if you opt in, and business and API products are usually covered by separate terms.

Providers describe these choices in their privacy documentation, and the details matter: which conversations a setting covers, whether feedback you submit can be used when training is off, and whether conversations flagged for safety review are treated differently. Our comparison pages set out, in each provider’s own words and with the date we checked, what OpenAI, Google, Anthropic, Microsoft and Perplexity say about training on consumer chats.

Training matters because it is hard to undo. A trained model does not store conversations as records that can be deleted; what it learns is spread across its parameters, so turning training off applies to future training, not to models that have already been trained. Models can also memorise rare text. Carlini and colleagues extracted hundreds of verbatim sequences from GPT-2’s training data, including names, phone numbers and email addresses, and found larger models more vulnerable than smaller ones 4; later work extended such attacks to production language models 5.

Some providers de-identify data before training. De-identification removes the link to your account. It does not necessarily remove personal details you typed into the conversation itself, which is why temporary or incognito modes, where they exist, are worth knowing about. For an on-device assistant the question changes: conversations that never reach the developer cannot train its models, so what matters is whether anything else, such as analytics, crash reports or feedback, carries them off the device.

05

Who can read your conversations?

With a cloud assistant, people working for the provider or its contractors may read a sample of conversations, to improve quality and enforce safety rules, and stored chats remain subject to the provider’s legal obligations. Anyone with access to your account or your unlocked device can read your history too.

Providers that use human review describe it in their privacy documentation, and the details differ: whether reviewers include contractors, whether a chat is disconnected from your account before review, how long reviewed chats are kept, and whether they are deleted when you delete your history. Some providers ask you not to enter anything confidential that you would not want a reviewer to read. Our comparison pages quote each provider’s own description.

Automated systems scan conversations too, and a flag can change how long one is kept: a chat flagged as breaking a provider’s usage rules can be kept for longer than ordinary history, and providers keep chats where the law requires it.

Shared links and connected apps also pass conversations to other people and services. An on-device assistant with no server removes the provider as a reader, but not the people who can pick up your phone.

06

What does an account link to your conversations?

An account ties each conversation to an identity, typically an email address and sometimes a phone number or payment details. That is what lets your history follow you across devices, and it is also what makes stored chats personal data about you.

Accounts make possible the features that depend on remembering you. Some assistants offer a memory feature that carries details from earlier chats into new ones, and those details can include sensitive information you mentioned in passing. Personalisation can be useful, but it means one conversation can shape later ones, so it is worth checking whether such a feature is on and how to clear it.

Because account-linked chats are personal data, data protection law gives you rights over them in many places. In the European Union, the General Data Protection Regulation includes a right of access (Article 15) and a right to erasure (Article 17) 6, and providers explain how to use them in their privacy policies.

Using an assistant without an account narrows what can be linked, but every service you connect to still sees the network address a request comes from. An assistant that needs no account and sends nothing to its developer leaves the developer nothing to link.

08

How can you judge an AI assistant’s privacy claims?

Judge a privacy claim by whether it names a mechanism you can check: where the model runs, what is stored and for how long, who can read it and what leaves the device, each with a scope and a date. A claim that names no mechanism, such as “private by design” on its own, tells you little.

  1. Where does the model run? Look for a plain statement for each feature: on the device, on the provider’s servers, or both.
  2. What is kept, and for how long? Look for numbers, in days or months, for your history, for back-end copies after you delete it and for copies kept for review or training.
  3. Is training on by default? Check whether conversations train models unless you opt out, only if you opt in, or not at all, and what a temporary or incognito mode still keeps.
  4. Who can read conversations? Check whether people review them, whether those people work for contractors, and how long reviewed chats are kept.
  5. What does the account link? Check what identity it requires, whether memory or personalisation uses past chats, and whether chats sync or back up.
  6. What else leaves the device? Analytics, crash reports, advertising, web search, link previews and site icons all send requests that can reveal what you were doing.
  7. Can you check it, and is it dated? Prefer published architecture, open-source code or independent verification to assurances, and look for a “last updated” date and a record of changes.

The same questions apply to Brello. The box below answers them for Brello 1.0, with the limitation that goes with each answer.

09

How do cloud assistants compare with Brello 1.0?

On these questions, cloud assistants differ mainly in their defaults and settings, while an on-device assistant answers most of them by where it runs. Table 1 sets out the general pattern for cloud assistants next to Brello 1.0.

Table 1The general pattern for consumer cloud assistants, next to Brello 1.0 as shipped (version 1.0.0). Individual providers and plans differ; business and API products have separate terms. The comparison pages quote each provider’s own documentation, with the date it was checked.
QuestionCloud assistant, in generalBrello 1.0
Where the prompt is processedOn the provider’s servers, which decrypt it so the model can process itOn the phone, by the model the app downloaded
History and retentionKept with your account until you delete it or an auto-delete period ends, as the provider’s policy and your settings decideOn the phone until you delete it, excluded from backups and device transfers
Training on conversationsDepends on the provider and your settings: on unless you opt out, only if you opt in, or not at allNone. No conversation reaches us, so none can be used
Human reviewSome providers have people review a sample of conversationsNone. There is no Brello server
AccountOften needed, and signing in links your chats to an identityNone
Web searchUsually sent from the provider’s servers, under its agreements with search partnersOff by default. When used, the search text and up to four page requests go directly from the phone, and those sites see its IP address

Each provider also offers controls a table can’t capture, and terms change, so check a provider’s own page before relying on a general pattern. Side-by-side comparisons with Brello 1.0, each sourced to the provider’s own documentation and dated, are on the comparison pages for Brello 1.0 and ChatGPT, Brello 1.0 and Gemini, Brello 1.0 and Claude, Brello 1.0 and Copilot and Brello 1.0 and Perplexity, all listed on the comparisons page.

Where Brello 1.0 is weaker matters too. Its models are far smaller than the models behind cloud assistants, so its answers can be wrong more often and it is weaker at long or complex reasoning; and with no sync, a lost or reset phone takes its chats with it. How Brello handles your questions, photos and answers sets out every data flow, and the glossary defines the terms used here.

References

Reviewed . Web pages were checked on that date.

  1. Rescorla, E. (2018). “The Transport Layer Security (TLS) Protocol Version 1.3.” RFC 8446, Internet Engineering Task Force. rfc-editor.org/rfc/rfc8446
  2. Google. “Gemini Nano.” Android Developers. Accessed 5 October 2026. developer.android.com/ai/gemini-nano
  3. Google. “Back up user data with Auto Backup.” Android Developers. Accessed 5 October 2026. developer.android.com/identity/data/autobackup
  4. Carlini, N. et al. (2021). “Extracting Training Data from Large Language Models.” 30th USENIX Security Symposium. arxiv.org/abs/2012.07805
  5. Nasr, M. et al. (2023). “Scalable Extraction of Training Data from (Production) Language Models.” arXiv preprint. arxiv.org/abs/2311.17035
  6. European Union (2016). “Regulation (EU) 2016/679 (General Data Protection Regulation).” Official Journal of the European Union, L 119, 4 May 2016. eur-lex.europa.eu/eli/reg/2016/679/oj
  7. W3C (2019). “Tracking Preference Expression (DNT).” W3C Working Group Note, 17 January 2019. w3.org/TR/tracking-dnt
  8. “Global Privacy Control (GPC).” Draft specification, W3C repository on GitHub. Accessed 5 October 2026. w3c.github.io/gpc

Version history

  1. 1.0First published.