How it works

Quantization, explained

How storing weights in 8, 4 or even 2 bits shrinks a language model, what block-wise scales do, what it costs in accuracy, and the files Brello 1.0 runs.

Brello Research11 min readVersion 1.0

Summary

Quantization compresses a neural network by storing its weights, and sometimes its activations, in fewer bits. Weight memory is roughly the number of weights times the bits per weight, so 1.7 billion weights take 3.4 GB at 16 bits and about 0.96 GB at 4 bits with one 16-bit scale per block of 32. Block-wise scales confine the damage an outlier does to its own block. Accuracy loss depends on the bit width, the method and the model, and we have not measured it for the files Brello uses. Brello 1.0 runs three pre-quantized LiteRT-LM files; Brello Core’s is a 977 MB 4-bit, block-32 build of Qwen3 1.7B.

  • Weight memory is roughly the number of weights times the bits per weight: 1.7 billion weights take 3.4 GB at 16 bits, 1.7 GB at 8 bits and 0.85 GB at 4 bits.
  • A 4-bit code can take 16 values; block-wise quantization gives each block of, for example, 32 weights its own scale, which adds 0.5 bits per weight when the scale is stored in 16 bits.
  • Brello Core runs litert-community’s “dynamic INT4 variant (block-32 weights, FP32 activations)” of Qwen3 1.7B, a 977,184,032-byte file; the same repository’s INT8 build is 2,056,729,520 bytes.
  • litert-community says the Gemma 4 E2B file Brello Vision runs “uses a mixture of 2bit, 4bit and 8 bit weights”.
  • Dettmers and Zettlemoyer found 4-bit precision “almost universally optimal for total model bits and zero-shot accuracy”; Brello has not measured the accuracy cost of its own files.
Contents8 sections

01

What is quantization?

Quantization is a way of compressing a neural network by storing its numbers, chiefly its weights, in fewer bits: 8 or 4 bits each instead of 16 or 32. At 4 bits, a language model’s weights take about a quarter of the memory they need at 16 bits, which is how models with billions of weights fit on a phone.

A language model is mostly a very large set of learned numbers, called weights or parameters. Qwen3 1.7B, the open model from Alibaba that Brello Core is built on, has 1.7 billion of them, and its published weights use bfloat16, a 16-bit number format 9. Writing each new token involves nearly all of those weights, so their size decides how much memory a model needs and, on a phone, much of how fast it runs. Brello is not affiliated with or endorsed by Alibaba or Google.

Quantization replaces each weight with a small integer and keeps a scale factor, so that an approximation of the original value can be rebuilt when the model runs 1. The approximation is the cost: every weight is rounded, and rounding error is what can reduce accuracy. The word, also spelt quantisation, comes from signal processing, where it means mapping a continuous range of values onto a limited set of levels.

This page works through the arithmetic, the common formats and what research says about accuracy, then looks at the three files Brello 1.0 runs. For why phones run small models in the first place, see Small language models, explained.

02

How do bits translate into model size?

Weight memory is the number of weights multiplied by the bits stored for each, divided by eight to give bytes. Halving the bits halves the memory.

Table 1 applies that to a model with 1.7 billion weights, the size of Qwen3 1.7B. Sizes on this site are decimal, so 1 GB is 1,000,000,000 bytes.

Table 1Memory for 1.7 billion weights at common precisions, computed as weights × bits ÷ 8. The fifth row adds the scales that block-wise formats store (section 04).
FormatBits per weight1.7 billion weights
32-bit floating point (FP32)326.8 GB
16-bit floating point (FP16, BF16)163.4 GB
8-bit integer (INT8), 256 values81.7 GB
4-bit integer (INT4), 16 values40.85 GB
INT4 with one 16-bit scale per 32 weights4.5≈0.96 GB
2-bit integer, 4 values20.425 GB

The table counts weights only. A running model also needs memory for its key-value cache, which holds the text it is working on and grows with the length of the conversation, for intermediate results and for the runtime itself. Context windows, explained works through the cache.

These numbers decide what fits. Brello recommends Brello Core for phones with 4 GB of RAM; at 16 bits, the weights of its base model alone would take 3.4 GB of that.

Fewer bits also mean less data to move. A model writes one token at a time, and each token means reading the weights from memory again, so generating text a request at a time is usually limited by memory bandwidth rather than by arithmetic 7. litert-community, which publishes the LiteRT-LM files Brello downloads, measured two versions of Qwen3 1.7B on a Samsung Galaxy S26. On its GPU, the 4-bit file wrote 40.5–40.7 tokens per second against 14.9–18.1 for the 8-bit file; on its CPU the two were close, at 8.0–8.3 and 7.7–7.8 10. These are the publisher’s measurements on one phone, not Brello’s, and they show that the gain depends on the hardware as well as the format.

03

What are INT8, INT4 and mixed precision?

INT8 and INT4 store each weight as an 8-bit or 4-bit integer, which can take 256 or 16 distinct values. Mixed precision uses different bit widths for different parts of the same model.

The integers are mapped back to real numbers with a scale and, in some schemes, an offset. In the scheme Jacob and colleagues used for integer-only inference, a real value r is represented by an integer q through a real-valued scale S and an integer zero-point Z, the integer that stands for exactly zero 2:

r=S(q−Z)
Equation 1Affine quantization: the real value r, the stored integer q, the scale S and the zero-point Z.

Trained weights are spread fairly evenly around zero, so they are often quantized symmetrically, with Z fixed at 0. A 4-bit integer can hold the 16 values from −8 to 7; a symmetric scheme can use −7 to 7, so that the levels sit evenly either side of zero, and set S to the largest weight’s magnitude divided by 7. Figure 1 below uses that scheme. At 8 bits the divisor is 127, so the steps between levels are about 18 times finer (127 ÷ 7 ≈ 18.1).

Not every part of a model is equally sensitive to rounding, so many formats mix precisions. Dettmers and colleagues found that large language models develop a few outlier features with very large values; their LLM.int8() method handles those dimensions in 16-bit while “more than 99.9% of values are multiplied in 8-bit” 3. Lin and colleagues found that protecting about 1% of the most important weights, identified from the activations they meet rather than from the weights themselves, greatly reduces quantization error 6. The same idea appears in files made for phones: litert-community says its build of Gemma 4 E2B, the file Brello Vision runs, “uses a mixture of 2bit, 4bit and 8 bit weights” 11.

04

How does block-wise quantization work?

Block-wise quantization splits the weights into small groups, such as 32 consecutive weights, and gives each group its own scale. An unusually large weight then coarsens the levels of its own block and no other.

With one scale for a whole matrix, the step size has to stretch to cover the largest weight anywhere in it, and with only 16 levels, ordinary weights end up crowded into the few levels nearest zero. Smaller blocks keep the step matched to the values around it. In more than 35,000 experiments on inference at 3 to 8 bits, Dettmers and Zettlemoyer found small block sizes to be one of only two changes that improved the trade-off between bits and accuracy; the other was the choice of data type 4.

The scales are not free. One 16-bit scale per 32 weights adds 16 ÷ 32 = 0.5 bits to every weight, so a 4-bit block format averages 4.5 bits per weight. Smaller blocks follow the weights more closely but spend more memory on scales.

Figure 1 follows one block of 32 weights through symmetric 4-bit quantization, using the block size of the file Brello Core runs, and then applies the same arithmetic to a whole model.

Rounding a block of 32 weights to 4 bits, and what it does to memory A chart of 32 illustrative weights in one block, drawn as lollipops around zero. The largest, 0.0581, sets a step of 0.0083 (0.0581 divided by 7), which spaces 15 levels from minus 7 to plus 7 steps. Each weight is rounded to its nearest level and stored as a 4-bit code; the other 31 weights use only codes from minus 3 to plus 3, and the largest rounding error is 0.0041. Two bars compare the memory for the block: 512 bits at 16 bits per weight, against 144 bits for 32 four-bit codes plus one 16-bit scale, or 4.5 bits per weight. For the 1.7 billion weights of Qwen3 1.7B, the same arithmetic gives 3.40 GB against about 0.96 GB; the 4-bit, block-32 file Brello Core runs is 977 MB. One block of 32 weights 16-bit value Rounded to a level 0.06 0.03 0 −0.03 −0.06 +7 0 −7 Step = largest |w| ÷ 7 = 0.0581 ÷ 7 = 0.0083 Largest |w| = 0.0581 Code 0+10−1−20+2+1+20+10−3+2+1+1−3−3−2−1+70+1−1+1+1−1+3+1+2−1−1 Largest rounding error: 0.0041, within half a step (0.00415). The other 31 weights use only codes −3 to +3. Memory for this block Memory for 1.7 billion weights 16-bit 512 bits 3.40 GB 4-bit 128 + 16 = 144 bits ≈0.96 GB 4.5 bits per weight · 3.6 times less memory Brello Core’s file: 977 MB 53.1 million blocks of 32 One block of 32 weights 16-bit value Rounded to a level +7 0 −7 Step = 0.0581 ÷ 7 = 0.0083 Largest |w| = 0.0581 4-bit codes, in order 0+10−1−20+2+1+20+10−3+2+1+1−3−3−2−1+70+1−1+1+1−1+3+1+2−1−1 Memory for this block Memory, 1.7 billion weights 16-bit 512 bits 3.40 GB 4-bit 128 + 16 = 144 bits ≈0.96 GB 4.5 bits per weight, 3.6 times smaller Brello Core’s file: 977 MB
  1. A block holds 32 weights, each stored in 16 bits. The values are illustrative but typical in size: small numbers spread around zero. Together they take 512 bits.
  2. The largest weight sets the step. Its size, 0.0581, divided by 7 gives a step of 0.0083, and the levels sit at whole multiples of the step, from −7 to +7.
  3. Each weight snaps to its nearest level and is stored as a 4-bit code. At run time, code × step rebuilds an approximation. No weight moves more than half a step; here the largest error is 0.0041.
  4. The block now takes 144 bits instead of 512. That is 32 codes of 4 bits plus one 16-bit scale: 4.5 bits per weight, 3.6 times less memory.
  5. Across 1.7 billion weights, 3.40 GB becomes about 0.96 GB. Brello Core’s 4-bit, block-32 file of Qwen3 1.7B is 977 MB. The arithmetic counts only weights and scales, so it lands close to the file, not on it.
Figure 1Symmetric 4-bit quantization of one block of 32 weights with a 16-bit scale per block, then the same arithmetic for Qwen3 1.7B’s 1.7 billion parameters. The weights are illustrative; the step, codes, errors and sizes are computed from them. Sizes are decimal (1 GB = 1,000,000,000 bytes).
Show data
Weight16-bit value4-bit codeRebuilt valueError
1−0.004100.0000−0.0041
20.0082+10.0083−0.0001
3−0.003600.0000−0.0036
4−0.0050−1−0.00830.0033
5−0.0149−2−0.01660.0017
6−0.003400.0000−0.0034
70.0178+20.01660.0012
80.0068+10.0083−0.0015
90.0166+20.01660.0000
100.004000.00000.0040
110.0063+10.0083−0.0020
120.003000.00000.0030
13−0.0267−3−0.0249−0.0018
140.0137+20.0166−0.0029
150.0081+10.0083−0.0002
160.0080+10.0083−0.0003
17−0.0271−3−0.0249−0.0022
18−0.0279−3−0.0249−0.0030
19−0.0142−2−0.01660.0024
20−0.0075−1−0.00830.0008
210.0581+70.05810.0000
22−0.000700.0000−0.0007
230.0083+10.00830.0000
24−0.0103−1−0.0083−0.0020
250.0049+10.0083−0.0034
260.0063+10.0083−0.0020
27−0.0106−1−0.0083−0.0023
280.0275+30.02490.0026
290.0089+10.00830.0006
300.0192+20.01660.0026
31−0.0099−1−0.0083−0.0016
32−0.0118−1−0.0083−0.0035

Step = 0.0581 ÷ 7 = 0.0083. Memory per block: 32 × 16 = 512 bits at 16 bits; 32 × 4 + 16 = 144 bits at 4 bits with a 16-bit scale. For 1.7 billion weights: 3.40 GB and about 0.96 GB (1.7 × 10⁹ × 4.5 ÷ 8 bytes).

Two things in the figure hold in general. Rounding error is bounded: no weight moves by more than half a step, here 0.0041 against a half-step of 0.00415. And the largest value in a block decides the step for everything else in it. One weight of 0.0581 sets a step of 0.0083, and the other 31 weights use only the seven codes from −3 to +3. That is the outlier problem in miniature, and it is why the block, not the whole matrix, is the unit that gets a scale.

05

Are weights and activations quantized the same way?

Usually not. Weights are fixed once a model is trained, so they can be quantized ahead of time; activations, the intermediate values computed from each input, change with every token and are often kept at higher precision.

Weight-only quantization shrinks the file and the memory the weights occupy; as the model runs, the stored integers are combined with activations held at higher precision. Brello Core’s file is described by its publisher as a “dynamic INT4 variant (block-32 weights, FP32 activations)”: weights stored as 4-bit integers in blocks of 32, and activations kept as 32-bit floating-point numbers 10.

“Dynamic” refers to how activations are handled. In LiteRT’s documentation, dynamic-range quantization “statically quantizes only the weights from floating point to integer at conversion time”; operators that support it then quantize activations on the fly, from the range of values they actually see, and the outputs “are still stored using floating point” 8.

Full integer quantization goes further and makes, in LiteRT’s words, “all model math” integer 8. That is the approach Jacob and colleagues developed for hardware with fast integer arithmetic. It has to fix a range for every activation in advance, and they trained their models with quantization simulated during training so that accuracy held up 2. Activations are harder to quantize in large language models because of the outlier features described above, a few channels with values far larger than the rest 3. Keeping activations in floating point, as Brello Core’s file does, sidesteps that problem.

06

What does quantization cost in accuracy?

Some accuracy, by an amount that depends on the bit width, the method and the model. Published results show little or no loss at 8 bits and small losses at 4 bits with careful methods, and below 4 bits the trade-off usually gets worse.

Three results frame the range. LLM.int8() ran inference at 8 bits in models of up to 175 billion parameters “without any performance degradation” 3. GPTQ, which adjusts the weights not yet rounded to make up for each one it rounds, reduced GPT models with 175 billion parameters to 3 or 4 bits per weight “with negligible accuracy degradation relative to the uncompressed baseline” 5. And across their 35,000 experiments, Dettmers and Zettlemoyer found 4-bit precision “almost universally optimal for total model bits and zero-shot accuracy”: for a fixed memory budget, a larger model at 4 bits tended to beat a smaller one at higher precision 4.

Those results come from particular models, methods and benchmarks, mostly far larger than a phone model, and they do not transfer automatically to a model with 1.7 billion parameters or to a particular file. Rounding to the nearest level, as in Figure 1, is the simplest method; GPTQ and AWQ choose their roundings with the help of a small sample of data, and both report lower error than plain rounding 5 6.

We have not measured how much accuracy the files Brello 1.0 runs lose against their 16-bit originals, so this page gives no figure for them. Like any on-device model, they can be wrong for many reasons; AI hallucinations, explained covers the most visible one.

07

Which quantized files does Brello 1.0 run?

Brello 1.0 runs three LiteRT-LM files that their publisher, litert-community on Hugging Face, quantized when converting the models: 977 MB for Brello Core, 2.59 GB for Brello Vision and 3.66 GB for Brello Pro. Brello downloads them as published, without an account.

Table 2The model files Brello 1.0 downloads. File names and sizes are the app’s; the descriptions are quoted from each file’s Hugging Face page, accessed 5 October 2026.
Brello CoreBrello VisionBrello Pro
Base modelQwen3 1.7B (Alibaba)Gemma 4 E2B (Google)Gemma 4 E4B (Google)
FileQwen3-1.7B_dynamic_wi4b32_afp32.litertlmgemma-4-E2B-it.litertlmgemma-4-E4B-it.litertlm
Download977 MB (977,184,032 bytes)2.59 GB (2,588,147,712 bytes)3.66 GB (3,659,530,240 bytes)
Publisher’s description“dynamic INT4 variant (block-32 weights, FP32 activations)” 10“uses a mixture of 2bit, 4bit and 8 bit weights” 11“a text decoder with 2.24 GB of weights and 0.67 GB of embedding parameters” 12
Recommended RAM4 GB6 GB12 GB

Brello Core’s file shows the arithmetic at work. Its 1.7 billion weights at 4.5 bits come to about 0.96 GB, close to the 977,184,032-byte file. litert-community also publishes an INT8 build of the same model, recipe dynamic_wi8_afp32, of 2,056,729,520 bytes, 2.1 times the size 10. Moving from 8-bit to 4-bit weights roughly halved the download.

The Gemma 4 files are organised differently. Google’s E2B and E4B models use per-layer embeddings, which give “each decoder layer its own small embedding for every token”, so the model card counts 2.3 billion effective parameters for E2B (5.1 billion with embeddings) and 4.5 billion for E4B (8 billion with embeddings) 13. litert-community says the E2B file’s “weight footprint in memory can be as low as 0.8 GB while the runtime uses memory mapping to support the 1.12GB of embedding parameters” 11. Memory mapping lets the operating system read parts of a file from storage as they are needed, so the embedding tables do not all have to sit in RAM at once.

For how these files reach a phone and run there, see how a local language model runs on Android, and for how Brello matches a model to a phone’s memory, see our research note on fitting a model to the phone.

References

Reviewed . Model cards, model pages and file listings were checked on that date.

  1. Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W. and Keutzer, K. (2021). “A Survey of Quantization Methods for Efficient Neural Network Inference.” arXiv preprint. arxiv.org/abs/2103.13630
  2. Jacob, B. et al. (2018). “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference.” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018). arxiv.org/abs/1712.05877
  3. Dettmers, T., Lewis, M., Belkada, Y. and Zettlemoyer, L. (2022). “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.” Advances in Neural Information Processing Systems 35 (NeurIPS 2022). arxiv.org/abs/2208.07339
  4. Dettmers, T. and Zettlemoyer, L. (2023). “The case for 4-bit precision: k-bit Inference Scaling Laws.” Proceedings of the 40th International Conference on Machine Learning (ICML 2023). arxiv.org/abs/2212.09720
  5. Frantar, E., Ashkboos, S., Hoefler, T. and Alistarh, D. (2023). “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.” International Conference on Learning Representations (ICLR 2023). arxiv.org/abs/2210.17323
  6. Lin, J. et al. (2024). “AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration.” Proceedings of Machine Learning and Systems 6 (MLSys 2024). arxiv.org/abs/2306.00978
  7. Pope, R. et al. (2023). “Efficiently Scaling Transformer Inference.” Proceedings of Machine Learning and Systems 5 (MLSys 2023). arxiv.org/abs/2211.05102
  8. Google AI Edge. “Post-training quantization.” LiteRT documentation. Accessed 5 October 2026. developers.google.com/edge/litert/models/post_training_quantization
  9. Qwen. “Qwen3-1.7B.” Model card and configuration file, Hugging Face. Accessed 5 October 2026. huggingface.co/Qwen/Qwen3-1.7B
  10. litert-community. “Qwen3-1.7B.” Model page and file listing, Hugging Face. Accessed 5 October 2026. huggingface.co/litert-community/Qwen3-1.7B
  11. litert-community. “gemma-4-E2B-it-litert-lm.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/litert-community/gemma-4-E2B-it-litert-lm
  12. litert-community. “gemma-4-E4B-it-litert-lm.” Model page, Hugging Face. Accessed 5 October 2026. huggingface.co/litert-community/gemma-4-E4B-it-litert-lm
  13. Google DeepMind. “Gemma 4 model card.” Google AI for Developers. Accessed 5 October 2026. ai.google.dev/gemma/docs/core/model_card_4

Version history

  1. 1.0First published.