1. Overview: when a task outgrows the device
Brello Super Intelligence, which Stuvio is developing, is being designed to send a task off the phone only when the phone can’t do it, and only to a server the phone has verified first. This paper sets out the design space for that verifiable private compute.
Brello 1.0, our app for Android and iPhone, has no server to verify. It runs Brello Pro, Brello Vision and Brello Core on the phone, so questions, photos and answers are processed on the device, with no account and no Brello server. If someone turns on web search, the search text goes directly to a search engine and up to four result pages are requested from their sites. Those services see a normal web request, including the phone’s IP address. Answering from the open web, without a server describes that pipeline.
A phone also sets hard limits. Brello’s models are downloads of 977 MB to 3.66 GB, recommended for phones with 4 to 12 GB of memory, and each answer is written within a 4,096-token context window. Small models are weaker at long or complex reasoning, as Fitting a model to the phone in your pocket explains. Deep reasoning, long documents and long-term personal context will need more computation than a phone has.
Brello SI is therefore being designed in three layers, set out in the Brello Charter (version 1.0, section 4). The device comes first. When a task needs more, it is intended to go to sealed compute: hardware-isolated, stateless servers designed to keep nothing, which the device is intended to verify before sending anything. The open web comes last, and only when the user chooses. This paper is about the middle layer.
The usual answer to “is this server private?” is a privacy policy. A policy can’t stop an engineer with production access, an attacker on the machine or a court order, and it can be rewritten. We want a device that checks which software is on the other end before sending any part of a task, and software that anyone can inspect to confirm it keeps nothing. The techniques behind both are confidential computing and remote attestation, introduced for a general reader in Confidential computing for AI, explained.
2. A threat model
Our threat model for sealed compute names six actors who might try to read a request: operator staff, a compromised host, legal compulsion, a network observer, other tenants on the same hardware and a targeted attacker. We include ourselves on purpose. A design that depends on Stuvio’s good behaviour gives no more assurance than a privacy policy.
Figure 1 introduces each actor in turn, shows what it would try, and names the property from Section 3 designed to stop it. Two of the six are only partly stopped, and the figure marks them in amber.
- Operator staff try to look inside a running node. The node is being designed with no shell, debugger or admin interface, for us or anyone else, and to write nothing from a request to logs.
- An attacker who controls the host tries to read the node’s memory. The processor is meant to keep that memory isolated, and altered software would change the node’s attestation.
- An order demands stored data, or a build that watches one person. Nothing would be retained to hand over, and devices would refuse any build that isn’t published in the transparency log.
- Someone on the network copies the traffic. Requests are designed to be sealed to the node’s key, so the observer would learn who is connecting, when and how much, but not what is asked. That remainder is why this case is amber.
- A workload on the same hardware probes shared caches and timing. Isolation is meant to protect the node’s memory, but side channels are only partly closed, so this case is amber too.
- An attacker with one node tries to draw one person’s requests to it. The device would pick the node, and the node is designed not to learn who is asking, so reaching one person would mean attacking many nodes.
- Operator staff, at Stuvio or wherever the hardware runs, should be unable to read a request while it runs and should find nothing afterwards.
- A compromised host, meaning an attacker who controls the host operating system or hypervisor, should find the node’s memory unreadable, and any change to the node’s software should be detectable.
- Legal compulsion to hand over data should find nothing retained. An order to change the software to watch one person would need a build that devices refuse unless it has been published.
- A network observer should see only encrypted traffic, although its size and timing remain visible.
- Other tenants on the same physical hardware may probe side channels such as timing and shared caches.
- A targeted attacker who controls one node should find it hard to make one person’s requests reach it, and an attempt at scale should be visible.
Three things sit outside this model. If the phone itself is compromised, nothing downstream can help, because the device is where trust starts. We assume the hardware’s isolation works as documented, an assumption Section 6 qualifies. And the model covers privacy only: whether an answer is accurate or safe is addressed in Evaluate first, then ship.
3. The properties we want
We derive six properties from the threat model. Each is meant to hold because of how sealed compute would be built, not because of how its operators behave.
- Isolation in use. While a node processes a request, its memory can’t be read or altered from outside: not by the host operating system, the hypervisor, other tenants or administrators.
- No privileged access. A running node has no shell, debugger or administrative interface, including for us. Changing a node means shipping a new, published release.
- No retention. A node uses a request only to compute its response, writes nothing from it to disk or logs, and erases it afterwards. Operational metrics are fixed in advance and contain no request content.
- Verifiable software. A device can learn exactly which software a node runs, and refuses any release that isn’t published in a public transparency log.
- Non-targetability. A node doesn’t learn who sent a request, and nobody can steer one person’s requests to a chosen node without an attack broad enough to notice.
- Sealed, minimal requests. The device sends only what the task needs, encrypted to the key of the node it has verified, so nothing in between can read it.
The first three describe what a node does. The fourth lets anyone check them. The last two limit what a node, or anyone on the way to it, could learn if something else failed. Four of the six follow requirements Apple published for its Private Cloud Compute: stateless computation, no privileged runtime access, non-targetability and verifiable transparency.1 We adopt that framing and credit it here.
Brello 1.0 already applies two of these ideas on the phone. Each reply opens a fresh model session, so old web results don’t accumulate from one reply to the next. And with web search off, Brello asks “Search the web for this?” before a time-sensitive question goes online, then sends only a refined search text, as Asking before going online describes.
Why encryption alone isn’t enoughEncryption protects data in transit and at rest, but a model has to read a request to answer it. Fully homomorphic encryption computes on encrypted data,2 but it remains far too slow for large language models at interactive speed. So a request is decrypted somewhere. The question is where, and who can see it while it is in use.
4. Five building blocks for verifiable compute
Five established techniques can supply these properties together: confidential computing, remote attestation, transparency logs, reproducible builds and oblivious relays. Each has known costs, and naming one here doesn’t mean Brello SI will use a particular vendor’s product. No hardware has been chosen.
4.1 Confidential computing
A trusted execution environment is a hardware-isolated region whose memory the rest of the machine, including the operating system and hypervisor, can’t read. The Confidential Computing Consortium defines confidential computing as “the protection of data in use by performing computation in a hardware-based, attested Trusted Execution Environment”.3 Early designs isolated a single process, and Costan and Devadas analyse one of them, Intel SGX, in detail.4 Newer designs isolate a whole virtual machine: AMD’s description of SEV-SNP, for example, adds integrity protection for a virtual machine’s encrypted memory against a malicious hypervisor.5 Research prototypes such as Graviton extend trusted execution to GPUs, which run most large models.6
The cost: the hardware vendor’s design, firmware and keys become part of what must be trusted. Side-channel research has repeatedly shown that isolation can leak,7,8 and protection for accelerators is less mature than for processors.
4.2 Remote attestation
Remote attestation lets a device learn, before it sends anything, which software an isolated environment is running. The hardware signs evidence that describes the loaded software, usually as cryptographic hashes called measurements. The IETF’s RATS architecture names the roles: an attester produces evidence, a verifier appraises it, and a relying party acts on the result.9 In our sketch the node is the attester and the phone is the relying party. Whether the phone should also be its own verifier is an open question (Section 7.2).
The cost: a measurement says which bytes are running, not whether they are trustworthy. Attestation turns “what is running?” into “who has checked it?”
4.3 Transparency logs
A transparency log is a public, append-only record that can’t be quietly rewritten. Certificate Transparency put the pattern into wide use for web certificates.10,11 Its log is a Merkle tree with signed heads. An inclusion proof shows that an entry is in the tree, and a consistency proof shows that a later tree extends an earlier one. Crosby and Wallach analysed tamper-evident logs of this kind in detail.12 In our sketch, each release’s measurement would be logged before any device accepted it.
The cost: a log proves that something was published, not that it is good, and it helps only if people watch it. Devices and independent witnesses must also compare the tree heads they see, so that a log can’t show different versions to different people.
4.4 Reproducible builds
A build is reproducible when anyone with the same source code, build environment and instructions can recreate bit-for-bit identical output.13,14 That is what ties a measurement, which is only a hash, to code that people can read. Ken Thompson’s Turing Award lecture remains the classic reminder that the build tools matter as much as the source.15
The cost: full reproducibility is hard for large software stacks, and drivers and firmware often ship only as binaries. Model weights can be hashed and published, though a hash says nothing about how a model behaves.
4.5 Oblivious relays
An oblivious relay splits what two parties learn about a request. In Oblivious HTTP,16 the client encrypts its request to the server’s public key with Hybrid Public Key Encryption (HPKE)17 and sends it through a relay run by a separate party. The relay sees the client’s IP address but not the content. The server sees the content but not the address.
The cost: the split holds only if relay and server don’t collude, which is why the specification requires that they are not operated by the same entity. The relay adds a network hop. And it hides the address, not the content: a request that names a person still names them.
4.6 How the approaches compare
Only confidential computing with attestation combines capacity beyond the phone with privacy properties that an outsider can check. It still asks users to trust the hardware vendor and whoever reviews the published code. Table 1 compares it with two common server-side approaches and with staying on the device.
Enforced, and checkable from outsidePartly, or only under conditionsNot enforced
| Property | Policy only | Encryption at rest | Confidential computing with attestation | On-device only |
|---|---|---|---|---|
| Isolation in use | No. Readable by the operator | No. Decrypted to be used | Yes. Enforced by hardware | Yes. No server |
| No privileged access | No. Promised | No. Promised | Yes. Checkable | Yes. No server |
| No retention | No. Promised | No. Kept, encrypted | Yes. Checkable | Yes. No server copy |
| Verifiable software | No. Not verifiable | No. Not verifiable | Partly. Trusts the hardware | Partly. By observing traffic |
| Non-targetability | No. Not enforced | No. Not enforced | Partly. Needs a relay | Yes. Nothing to target |
| Sealed, minimal requests | Partly. Encrypted to the operator | Partly. Encrypted to the operator | Yes. Sealed to a verified node | Yes. Stays on the device |
| Capacity beyond the phone | Yes. Yes | Yes. Yes | Yes. Yes | No. Limited to the phone |
| Who must still be trusted | The operator | The operator and its key handling | The hardware vendor and the people who review the code | The phone and the app |
5. How a request could flow (a design sketch)
In our current sketch, a device would verify a node’s attestation against the public log before it sent anything. It would then send the task, sealed to that node’s key, through an independent relay. Figure 2 shows the sequence across four parties: the transparency log, the device, the relay and the sealed node. It is a sketch of our direction, not a specification, and the details will change.
- Each release would be logged before any device accepts it. Its measurement, a hash of the exact software, would be appended to a public log that anyone can audit and no one can quietly rewrite.
- The device would send a fresh nonce, and the node would return signed evidence. The evidence would name the node’s software and public key and repeat the nonce, so an old answer couldn’t be replayed.
- The device is designed to check the evidence against the log before it sends anything. Genuine hardware, a matching nonce, a logged measurement and a bound key would all have to hold. If any check failed, nothing would be sent.
- The task would be sealed to the node’s key and sent through an independent relay. The relay would forward it without the device’s IP address, and couldn’t read what it carries.
- The node would decrypt the task and run the model in isolated memory. Nothing outside that memory, including the host and our own staff, is meant to be able to read it while it is in use.
- The response would return sealed, and the node would erase the task, the response and the keys. The node’s software is being designed with nowhere to write them.
- Publish. Each release would be built reproducibly, and its measurement appended to the transparency log before any device accepts it.
- Challenge. Through the relay, the device would send the node a fresh random value, a nonce. The node would return evidence signed by its hardware, containing its measurement, its public key and the nonce, so old evidence couldn’t be replayed.
- Verify. The device would check that the signature chains to genuine hardware, the nonce matches, the measurement is in the log and the key is bound to that measurement. It would also check that the log is consistent with the tree head it saw last. If any check failed, nothing would be sent.
- Send. The device would encrypt the minimal task to the node’s key with a fresh ephemeral key, and send it through the relay, which would forward it without the device’s IP address.
- Compute. The node would decrypt the task and run the model inside isolated memory.
- Reply and erase. The node would encrypt the response with a key derived from the request, return it through the relay, and erase the task, the response and the keys. Its software is being designed with nowhere to write them.
The checks in step 3 are small enough to show as code:
// Illustrative only: the checks a device would run before it sends anything.
// Through the relay
evidence = request_evidence(node, nonce)
// Genuine isolated hardware
require verify_signature(evidence, hardware_roots)
// Fresh, not replayed
require evidence.nonce == nonce
// A published release
require log.includes(evidence.measurement, proof)
// The same log everyone sees
require log.consistent(last_seen_head, proof.head)
// Bound by the signature above
key = evidence.public_key
// Fresh ephemeral key
sealed = hpke_seal(key, minimal(task))
// The first time the task leaves
send_via_relay(sealed)
Two details serve the threat model. The proofs in step 3 could travel with the evidence, so the device would never contact the log operator, which would reveal its address. And the device, not a server that knows who is asking, would choose among attested nodes, which supports non-targetability.
5.1 What each party can learn
The design is meant to ensure that no party other than the device sees both who is asking and what is asked. The relay would see the address, the node would see the content, and the log would hold only measurements of software. Table 2 sets this out for each party in the threat model.
| Party | Can learn | Should not be able to learn |
|---|---|---|
| Network observer | That a device is talking to the relay, when, and how much data moves | What is asked, or which node answers |
| Relay | The device’s IP address, and the size and timing of requests | What is asked, because requests would be sealed to the node |
| Sealed node | The task, while it processes it | The device’s IP address, or who sent the task, unless the task itself says |
| Transparency log | Measurements of published releases | Anything about requests or the people who make them |
| Operator staff, including us | Operational metrics fixed in advance | Request content, because there would be no runtime access and nothing would be kept |
| Other tenants | Timing and cache behaviour they can measure on shared hardware | Request content, although side channels are only partly closed (Section 6) |
Because every device would check the same log, serving one person a special build would mean publishing that build for everyone to see. And because a node is designed never to learn who is asking, an attacker who controls one node couldn’t tell whose requests it receives.
6. Limitations of this design space
This design narrows whom a user has to trust. It doesn’t remove trust, and several attacks remain open.
- It trusts the hardware vendor. Isolation and attestation rest on the processor’s design, firmware and signing keys. The Foreshadow attack extracted attestation keys from processors in the field and forged attestation responses with them,7 and a device can’t detect that kind of forgery on its own.
- Side channels are only partly closed. Timing, shared caches and speculative execution have leaked data across isolation boundaries.7,8 Mitigations such as dedicated hardware, constant-time code and prompt firmware updates cost capacity, and none is complete.
- Measurements aren’t meaning. Attestation proves which software runs. Whether it keeps nothing depends on people reading the published code and on reproducible builds, which some components, such as firmware, may not have.
- Metadata still leaks. A network observer and the relay can see that a device uses the service, when and how much. Padding and batching reduce this, at a cost in bandwidth and delay.
- Collusion breaks the split. If the relay and node operators share what they see, addresses and requests can be joined again.
- Content can identify. The relay hides addresses, not words. A task that names a person still names them.
- The device is the root of trust. A compromised phone can skip every check.
- Nothing here is built or measured. We don’t know the design’s latency, cost or failure rate, and this paper gives no performance figures because we have none.
7. Open questions
Six questions about sealed compute remain open. We intend to answer each in public before it handles anyone’s data.
7.1 Making verification meaningful to people
A phone can check a signature in milliseconds, but a person can’t read a measurement. Most people will rely on others, such as researchers who rebuild releases and monitors who watch the log. The hard part is showing what was checked without turning “verified” into one more badge to take on faith. Brello 1.0 is our starting point. It states its privacy behaviour in plain words in Settings (Figure 3), and each finished answer ends with a line such as “2.4s · On-device” that shows whether the web was used. For sealed compute, plain words would be honest only if the checks behind them were public. On-device AI, explained covers how to judge claims like these.

7.2 Who verifies, and how
The phone could appraise evidence itself, or rely on a separate verifier, as the RATS architecture allows.9 A separate verifier simplifies the phone’s work but adds a party to trust, and one that might learn who is asking. How proofs and tree heads reach the device without revealing it is part of the same question.
7.3 Cost and latency
Isolation, attestation and an extra network hop all add cost and delay, and stateless nodes give up some of the caching that conventional serving uses between turns. We don’t yet know the price, and we won’t guess at numbers. Part of the answer is the first layer: a task that stays on the device needs no remote verification.
7.4 Updates and revocation
Every new model or software version means a new measurement, a new log entry and a new build to reproduce. Security fixes need to ship quickly, while outside review takes time. How to balance the two, how long an old release should stay acceptable and how to withdraw a flawed one are all open.
7.5 Isolating accelerators
Large models run on accelerators rather than general-purpose processors, and trusted execution for accelerators is newer, with less tooling and less public analysis.6 We need to know which guarantees hold along the whole path, from processor to accelerator and back, before relying on them.
7.6 Abuse without identity
A service that doesn’t know who is asking still has to resist abuse and manage load. Privacy Pass tokens, which let a client show it is entitled to make a request without revealing which client it is, are one direction.18 Whether they are enough, and how they fit with the relay, we haven’t answered.
8. What we intend to publish
We will publish a description of sealed compute, including what runs where and what each party can and cannot see, before anyone outside the team uses it (Brello Charter, version 1.0, section 4). Alongside it, we intend to publish:
- the measurement of every release a device will accept, in a public transparency log;
- source code and build instructions to reproduce those measurements, with a plain account of anything that can’t yet be reproduced;
- the device-side verification code, so that the checks themselves can be audited;
- the side-channel mitigations we rely on, and the ones we don’t;
- a way for security and privacy researchers to report what they find, on our security and disclosure page.
These follow the charter’s commitment that independent researchers should be able to verify what runs, where it runs and what it keeps (Brello Charter, version 1.0, section 5). If a property in this paper proves out of reach, we will say which one and why. Until sealed compute meets this bar, the device stays the default, as it is in Brello 1.0: no Brello server, and nothing on our side to verify. How Brello handles your questions, photos and answers sets out Brello 1.0’s data flow in full.
References
- Apple Security Research. “Private Cloud Compute: A new frontier for AI privacy in the cloud.” 10 June 2024. security.apple.com/
blog/ (accessed 5 October 2026).private-cloud-compute - C. Gentry. “Fully Homomorphic Encryption Using Ideal Lattices.” In Proceedings of the 41st Annual ACM Symposium on Theory of Computing (STOC), pages 169–178, 2009.
- Confidential Computing Consortium. “A Technical Analysis of Confidential Computing.” Version 1.3, November 2022. confidentialcomputing.io (accessed 5 October 2026).
- V. Costan and S. Devadas. “Intel SGX Explained.” IACR Cryptology ePrint Archive, Report 2016/086, 2016. eprint.iacr.org/
2016/ 086 - AMD. “AMD SEV-SNP: Strengthening VM Isolation with Integrity Protection and More.” White paper, January 2020. amd.com (accessed 5 October 2026).
- S. Volos, K. Vaswani and R. Bruno. “Graviton: Trusted Execution Environments on GPUs.” In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI), pages 681–696, 2018. usenix.org
- J. Van Bulck, M. Minkin, O. Weisse, D. Genkin, B. Kasikci, F. Piessens, M. Silberstein, T. F. Wenisch, Y. Yarom and R. Strackx. “Foreshadow: Extracting the Keys to the Intel SGX Kingdom with Transient Out-of-Order Execution.” In Proceedings of the 27th USENIX Security Symposium, pages 991–1008, 2018. usenix.org
- P. Kocher, J. Horn, A. Fogh, D. Genkin, D. Gruss, W. Haas, M. Hamburg, M. Lipp, S. Mangard, T. Prescher, M. Schwarz and Y. Yarom. “Spectre Attacks: Exploiting Speculative Execution.” In 2019 IEEE Symposium on Security and Privacy, 2019.
- H. Birkholz, D. Thaler, M. Richardson, N. Smith and W. Pan. “Remote ATtestation procedureS (RATS) Architecture.” RFC 9334, IETF, January 2023. rfc-editor.org/
rfc/ rfc9334 - B. Laurie, A. Langley and E. Kasper. “Certificate Transparency.” RFC 6962, IETF, June 2013. rfc-editor.org/
rfc/ rfc6962 - B. Laurie, E. Messeri and R. Stradling. “Certificate Transparency Version 2.0.” RFC 9162, IETF, December 2021. rfc-editor.org/
rfc/ rfc9162 - S. A. Crosby and D. S. Wallach. “Efficient Data Structures for Tamper-Evident Logging.” In Proceedings of the 18th USENIX Security Symposium, 2009. usenix.org
- Reproducible Builds. “Definitions.” reproducible-builds.org/
docs/ (accessed 5 October 2026).definition - C. Lamb and S. Zacchiroli. “Reproducible Builds: Increasing the Integrity of Software Supply Chains.” IEEE Software, 39(2):62–70, 2022. arxiv.org/
abs/ 2104.06020 - K. Thompson. “Reflections on Trusting Trust.” Communications of the ACM, 27(8):761–763, 1984.
- M. Thomson and C. A. Wood. “Oblivious HTTP.” RFC 9458, IETF, January 2024. rfc-editor.org/
rfc/ rfc9458 - R. Barnes, K. Bhargavan, B. Lipp and C. Wood. “Hybrid Public Key Encryption.” RFC 9180, IRTF, February 2022. rfc-editor.org/
rfc/ rfc9180 - A. Davidson, J. Iyengar and C. A. Wood. “The Privacy Pass Architecture.” RFC 9576, IETF, June 2024. rfc-editor.org/
rfc/ rfc9576
Cite this work
Brello Research. “Private compute you can verify: the design space.” Stuvio, 5 October 2026. https://brello.ai/research/verifiable-private-compute/
@misc{brello2026verifiableprivate,
title = {Private compute you can verify: the design space},
author = {{Brello Research}},
year = {2026},
month = {oct},
url = {https://brello.ai/research/verifiable-private-compute/},
note = {Stuvio}
}
Version history
- 1.0First published as a design note. Nothing described here has been built.



