How it works

BM25, explained

The classic ranking function of keyword search, term by term: the formula, what k1 and b control, an interactive example, and where matching words falls short.

Brello Research11 min readVersion 1.0

Summary

BM25 is the classic ranking function of lexical search. It scores a passage against a query by adding, for each query word, the word’s rarity across the collection multiplied by a term-frequency factor that saturates with repetition and is corrected for the passage’s length. Two parameters set its behaviour: k1 controls saturation and b controls length normalisation. This explainer derives each part of the formula, scores an example by hand and interactively, compares BM25 with TF-IDF and embedding retrieval, and sets out where matching words falls short. Brello 1.0 uses BM25, with k1 = 1.2 and b = 0.75, to choose which web passages its on-device model reads.

  • BM25 adds one contribution per query word: the word’s inverse document frequency multiplied by a term-frequency factor that saturates with repetition and is corrected for passage length.
  • At k1 = 1.2, the first occurrence of a word in an average-length passage is worth 1.000, the second adds 0.375, and no number of repetitions lifts the factor past k1 + 1 = 2.2.
  • At b = 0.75, a single occurrence counts 1.257 in a passage half the average length and 0.710 in a passage twice the average length.
  • BM25 matches exact words only, so a relevant passage written in other words, or in other word forms without stemming, can score close to zero.
  • Brello 1.0 ranks web passages of about 420–700 characters with BM25 (k1 = 1.2, b = 0.75) and keeps at most two per source, within 3,400 characters for its Gemma 4 models and 2,600 for Qwen3.
Contents8 sections

01

What is BM25?

BM25 is a ranking function that scores how well a passage of text matches a search query. It adds up evidence from each word of the query: how often the word appears in the passage, how rare it is across the collection, and how long the passage is. The highest score ranks first.

The name stands for Best Matching 25. It comes from the Okapi retrieval system at City University London, where Stephen Robertson, Steve Walker and colleagues tested a series of weighting functions; BM25 first appeared in their experiments for TREC-3, the third Text REtrieval Conference, in 1994 2. Its theory is the probabilistic relevance framework that Robertson, Karen Spärck Jones and others developed from the 1970s onwards, which Robertson and Hugo Zaragoza set out in full in 2009 1.

More than three decades later, BM25 is still the reference point for lexical search, the family of methods that match the words of a query against the words of a document. It needs no training data and no neural network, and every part of a score can be traced to a count. New retrieval models are routinely measured against it, and on the BEIR benchmark it proved a more robust baseline across unfamiliar collections than many trained retrievers 7.

Brello 1.0 uses BM25 for one job: choosing which passages of the web pages it has just read are passed to the language model on the phone (section 06).

02

The BM25 formula, term by term

BM25 gives a passage one contribution for each query word and adds them up. Each contribution is the word’s rarity, its inverse document frequency (IDF), multiplied by a term-frequency factor that rises with each repetition, levels off, and is scaled down for longer passages. Equations 1 to 3 state it in full.

score(D,Q) = ∑q∈Q IDF(q) · TF(q,D) (1)
IDF(q) = ln (1+ N−n(q)+0.5 n(q)+0.5 ) (2)
TF(q,D) = f·(k1+1) f+k1· (1−b+b·|D|avgdl) (3)
Equations 1–3BM25. Q is the query and D the passage; f is the number of times the word q occurs in D; |D| is the passage’s length in words and avgdl the average length across the collection; N is the number of passages and n(q) the number that contain q. Equation 2 is the non-negative IDF used by Elasticsearch and other engines built on Apache Lucene; the classic form omits the 1 inside the logarithm 1 5.

Inverse document frequency: how rare the word is

IDF rewards words that pick out few passages. A word found in every passage tells the ranker almost nothing, so its weight falls close to zero; a word found in only a few separates them sharply. The idea of weighting a search term by its rarity goes back to Karen Spärck Jones in 1972 3.

The 1 inside the logarithm of Equation 2 is a practical refinement. Without it, the weight is ln((N − n + 0.5) ÷ (n + 0.5)), the form given by Robertson and Zaragoza 1, which turns negative for any word found in more than half the passages, so matching a common word would lower a score. Elasticsearch adds the 1 and keeps every weight positive 5. The examples on this page use that form.

The term-frequency factor: how often the word appears

The factor in Equation 3 grows with f, the number of times the word occurs in the passage, but with diminishing returns. Because f appears in both the numerator and the denominator, the factor climbs steeply for the first occurrence and then flattens towards a ceiling of k1 + 1. Table 1 gives its values at k1 = 1.2 and b = 0.75, the settings Brello 1.0 uses: in a passage of average length, the first occurrence is worth 1.000, the second adds 0.375 and the third adds 0.196, and ten occurrences reach 1.964 of a possible 2.2.

Length normalisation: how long the passage is

The ratio |D| ÷ avgdl compares the passage’s length with the average. A long passage has more chances to contain any word by accident, so BM25 enlarges the denominator for passages longer than average and shrinks it for shorter ones; b sets how strongly. At b = 0.75, a single occurrence counts 1.257 in a passage half the average length and 0.710 in one twice the average length.

Table 1The term-frequency factor of Equation 3 at k1 = 1.2 and b = 0.75. Every column approaches k1 + 1 = 2.2; shorter passages approach it faster. Computed from the formula.
Occurrences (f)Half the average lengthAverage lengthTwice the average length
11.2571.0000.710
21.6001.3751.073
31.7601.5711.294
51.9131.7741.549
102.0471.9641.818
Ceiling, k1 + 12.2002.2002.200

03

What k1 and b control

k1 sets how quickly repeated occurrences of a word stop adding to a passage’s score, and b sets how much a passage’s length counts against it. Both are chosen by whoever runs the search; neither is learned from the query.

At k1 = 0 the factor is exactly 1 for any word that is present, so BM25 only asks which of the query’s words a passage contains. As k1 grows, each repetition counts for more, and for very large values the factor grows almost in proportion to the count, as raw term frequency does in classic TF-IDF. b runs from 0, where length is ignored, to 1, where a passage’s matches are scaled fully by its length relative to the average. The middle ground reflects two reasons a passage can be long: it may say the same thing in more words, which deserves to be normalised away, or it may cover more ground, which does not.

Typical settings sit in a narrow band. Manning, Raghavan and Schütze report that experiments support k1 between 1.2 and 2 and b = 0.75 4, and Elasticsearch uses k1 = 1.2 and b = 0.75 unless it is configured otherwise 5. Table 2 summarises the extremes.

Table 2What each parameter does at the ends of its range. Derived from Equation 3.
SettingEffect on the score
k1 = 0Only presence counts. Each query word in the passage adds its IDF once, and length has no effect.
Large k1Each repetition counts almost in full, as raw term frequency does.
b = 0No length normalisation. A long passage keeps every match it collects.
b = 1Full normalisation. Matches are scaled by the passage’s length relative to the average.

Figure 1 puts both parameters in your hands. It scores six short passages against the query ‘why is the sky blue’, the search text Brello 1.0 writes for the question “Why is the sky blue? Keep it short.”, and draws each word’s contribution as a coloured segment. Play the steps, or drag the sliders and watch the ranking re-sort.

Querywhy is the sky blue

1.20
0.75
  • whyIDF 0.4424 of 6
  • isIDF 0.0746 of 6
  • theIDF 0.0746 of 6
  • skyIDF 0.6933 of 6
  • blueIDF 0.6933 of 6
  • 4Sunset26 words1.43
  • 6Word forms26 words0.19
  • 2Long article84 words2.16
  • 5Synonyms22 words0.20
  • 1Short explainer23 words2.61
  • 3Paint shop27 words2.01

Arithmetic for

Synonyms

WordCountIDFFactorScore
why00.4420.0000.000
is10.0741.1760.087
the20.0741.5320.114
sky00.6930.0000.000
blue00.6930.0000.000
BM25 score0.201

Each word’s score is IDF × factor (Equations 2 and 3), with an average passage length of 34.67 words. Point at or tap a passage to see its arithmetic.

  1. Each query word is weighted by its rarity. ‘the’ and ‘is’ appear in all six passages, so their IDF is only 0.074. ‘sky’ and ‘blue’ appear in three each and weigh 0.693.
  2. A passage’s score is the sum of its words’ contributions. At k1 = 1.2 and b = 0.75, the values Brello 1.0 uses, the 23-word explainer ranks first with 2.61.
  3. Raising k1 to 3 lets repetition count for more. The paint shop page, which says ‘blue’ seven times, overtakes the long article. At 1.2, each repeat adds less than the one before.
  4. Setting b to 0 removes the length correction. The 84-word article mentions ‘sky’ and ‘blue’ three times each, and now outscores the 23-word explainer by 2.88 to 2.31.
  5. Two relevant passages score almost nothing. One says ‘azure heavens’, the other ‘skies’ and ‘bluest’: BM25 counts exact words only. Drag the sliders to try other settings.
Figure 1BM25 scores for six passages and the query ‘why is the sky blue’, computed in your browser with Equations 1 to 3. The passages are illustrative and shorter than the 420–700 characters Brello 1.0 ranks, and words are matched exactly, without stemming. “Show data” lists every number.
Show data
BM25 scores at k1 = 1.2 and b = 0.75, by query word, rounded to three decimals.
RankPassageWordswhyistheskyblueScore
1Short explainer230.5120.1260.1130.8041.0532.607
2Long article840.2790.0890.1240.8350.8352.162
3Paint shop270.4860.1090.0810.0001.3342.010
4Sunset260.4920.0830.0830.7720.0001.429
5Synonyms220.0000.0870.1140.0000.0000.201
6Word forms260.0000.0830.1100.0000.0000.192

IDF by word (Equation 2, N = 6): ‘why’ 0.442, in 4 of 6 passages; ‘is’ 0.074, in 6 of 6 passages; ‘the’ 0.074, in 6 of 6 passages; ‘sky’ 0.693, in 3 of 6 passages; ‘blue’ 0.693, in 3 of 6 passages. Average passage length 34.67 words (208 in all). Text is lower-cased and split at every character that is not a letter or a digit.

  1. Short explainer (23 words; why 1, is 3, the 2, sky 1, blue 2): “Why is the sky blue? Sunlight is scattered by the molecules in air, and blue light is scattered far more strongly than red.”
  2. Long article (84 words; why 1, is 3, the 8, sky 3, blue 3): “In the nineteenth century, John Tyndall and Lord Rayleigh studied why the sky looks blue. Rayleigh showed that very small particles scatter short wavelengths far more than long ones, so blue light from the Sun is spread across the whole sky. Violet light is scattered even more, but sunlight contains less of it and our eyes are less sensitive to it, so the sky appears blue rather than violet. Near the horizon the colour is paler, because the light has passed through more air.”
  3. Paint shop (27 words; why 1, is 2, the 1, sky 0, blue 7): “Why choose blue? Blue paint, blue tiles, blue throws and blue lamps: the blue range is made for every room, and delivery is free on blue orders.”
  4. Sunset (26 words; why 1, is 1, the 1, sky 1, blue 0): “Why is the sky red at sunset? Low sunlight crosses far more air, so most of its short wavelengths are scattered away before it reaches you.”
  5. Synonyms (22 words; why 0, is 1, the 2, sky 0, blue 0): “Air molecules scatter the short wavelengths of sunlight much more than long ones, which is what makes the daytime heavens look azure.”
  6. Word forms (26 words; why 0, is 1, the 2, sky 0, blue 0): “Clear skies look bluest high overhead and fade to a paler shade near the horizon, where the light that reaches us is scattered more than once.”

04

BM25, TF-IDF and embeddings

TF-IDF and BM25 both score a passage by the words it shares with the query; BM25 adds saturation and an adjustable length correction. Embedding-based retrieval works differently: a neural network turns the query and each passage into vectors, and passages are ranked by how close their vectors are, which lets them match on meaning rather than spelling.

Classic TF-IDF multiplies a word’s count by its IDF, often damping the count with a logarithm and normalising long documents by the length of their vectors 4. BM25 grew out of a probabilistic model of relevance rather than vector geometry, and its two parameters make saturation and length explicit 1. Both are sparse methods: a passage is represented by the words it contains, so an inverted index can find every passage that contains a query word without reading the rest.

Dense retrieval, such as Dense Passage Retrieval 6, trains encoders so that questions and the passages that answer them land close together in vector space. It can find a passage that shares no words with the question. The costs are a trained model, an index of vectors and the computation to embed every passage, and quality can drop on collections unlike the training data: in BEIR, BM25 was a robust zero-shot baseline that dense retrievers often failed to beat outside their training domain 7. Hybrid systems combine the two, using sparse scores for exact names, numbers and rare terms and dense scores for paraphrase. Table 3 compares the three.

Table 3Three ways to score passages against a query. General properties; individual systems vary.
PropertyTF-IDFBM25Dense embeddings
ComparesWordsWordsLearned vectors
Training neededNoNoYes, an encoder model
Repeated wordsCounted in full, or log-dampedSaturate towards k1 + 1No explicit counts
Long documentsVector-length normalisationAdjustable with bDepends on the encoder and how text is split
Different words, same meaningMissedMissedOften matched
Exact names, codes and numbersMatchedMatchedCan be blurred

05

A worked example

Scoring one passage by hand shows how the parts combine. Take the short explainer from Figure 1, “Why is the sky blue? Sunlight is scattered by the molecules in air, and blue light is scattered far more strongly than red.”, against the query ‘why is the sky blue’ with k1 = 1.2 and b = 0.75. Its five word contributions add up to 2.61.

Start with the collection. There are N = 6 passages with an average length of 34.67 words (208 words in all), and the explainer has 23. Its length term is 1 − 0.75 + 0.75 × 23 ÷ 34.67 = 0.7476, so the denominator of Equation 3 adds k1 × 0.7476 = 0.8971 to each word’s count. Table 4 then applies Equations 2 and 3 to each query word.

Table 4BM25 for the short explainer, word by word. Values are rounded; the total is computed before rounding.
WordCount (f)Passages containing itIDFFactorContribution
why14 of 60.4421.1600.512
is36 of 60.0741.6940.126
the26 of 60.0741.5190.113
sky13 of 60.6931.1600.804
blue23 of 60.6931.5191.053
BM25 score2.607

Three things stand out. ‘sky’ and ‘blue’, found in only three of the six passages, supply 1.86 of its 2.61 points. ‘is’ and ‘the’, found in every passage, add about 0.24 between them, so BM25 discounts them without a list of stop words: rarity does the work. And the second ‘blue’ is worth less than the first, because the factor for two occurrences is 1.519, not twice 1.160.

Figure 1 runs the same arithmetic for every passage each time a slider moves, and its “Show data” table lists the results.

06

BM25 inside an on-device answer engine

Brello 1.0 uses BM25 to decide which parts of the web pages it has just read are given to the language model on the phone. BM25 suits the job because it needs no trained model, no server and, at the scale of a few pages, no index built in advance: the passages are scored on the phone, moments after they are fetched.

The coverage bonus counters a weakness visible in Figure 1: plain BM25 can rank a passage highly for repeating one rare word, as the paint shop page does with ‘blue’, and rewarding passages that match more of the question favours the ones that address all of it. The lead-paragraph bonus favours the opening of a page, which often states its subject. The two-per-source limit stops one long page from filling the budget, so an answer can draw on several sources.

The budget caps how much of the model’s 4,096-token context window the passages can take; the window also has to hold the instructions, up to six earlier messages and the answer (Context windows, explained). BM25 is the ranking stage of a retrieval-augmented generation pipeline. Retrieval-augmented generation, explained covers the other stages, and the research note ‘Answering from the open web, without a server’ describes the whole system.

07

The limits of lexical ranking

BM25 compares spellings, not meanings, so it misses relevant passages written in other words and can favour irrelevant ones that repeat the right words. It also cannot tell whether a passage is true.

  • Vocabulary mismatch. A passage that says ‘azure heavens’ scores nothing for ‘sky’ or ‘blue’, however well it answers the question. Query rewriting and dense retrieval are the usual remedies (section 04).
  • Word forms. Without stemming, which reduces words to a shared root, ‘skies’ and ‘bluest’ do not match ‘sky’ and ‘blue’. Stemming helps but can merge unrelated words, such as ‘universe’ and ‘university’.
  • Word order and negation. BM25 treats a passage as a bag of words: ‘sky blue’ and ‘blue sky’ score the same, and ‘not blue’ still matches ‘blue’.
  • Repetition. Saturation limits what repetition can earn, but a page stuffed with a query’s words can still outrank a better one, as the paint shop page does at k1 = 3.
  • Small collections. IDF is only as reliable as the passages it is measured on. Over a few pages, a word can look rare by chance: ‘why’ weighs 0.442 in Figure 1 only because two of the six passages lack it, and it would weigh far less across a large collection.
  • Relevance is not accuracy. A passage can match every word of a question and still be wrong, out of date or written to mislead, including text crafted to give instructions to the model that reads it. Ranking decides what the model sees, not whether it is true (Why AI makes things up; Prompt injection, explained).

Brello 1.0 adds checks around BM25 rather than replacing it. Before ranking, it moves on to the next search provider unless at least a third of the top results mention the question’s key terms, which is how it rejects junk such as dictionary pages for ‘why’ when the question was ‘why is the sky blue’. After ranking, the model is told: “If the results do not answer the question, say so briefly and answer from general knowledge.” Neither check can confirm that a source is correct, and the on-device model, far smaller than frontier cloud models, can still be wrong.

Lexical ranking’s strengths mirror these limits. It matches exact names, numbers and rare terms that embedding models can blur, its scores can be explained word by word, as in section 05, and it needs no model of its own, which matters on a phone that is already running a language model.

References

Reviewed . Web sources were accessed on that date.

  1. Robertson, S. and Zaragoza, H. (2009). “The Probabilistic Relevance Framework: BM25 and Beyond.” Foundations and Trends in Information Retrieval, 3(4), 333–389. doi.org/10.1561/1500000019
  2. Robertson, S. E., Walker, S., Jones, S., Hancock-Beaulieu, M. M. and Gatford, M. (1995). “Okapi at TREC-3.” In Overview of the Third Text REtrieval Conference (TREC-3), NIST Special Publication 500-226, p. 109. trec.nist.gov/pubs/trec3/t3_proceedings.html
  3. Spärck Jones, K. (1972). “A statistical interpretation of term specificity and its application in retrieval.” Journal of Documentation, 28(1), 11–21. doi.org/10.1108/eb026526
  4. Manning, C. D., Raghavan, P. and Schütze, H. (2008). Introduction to Information Retrieval, section 11.4.3, “Okapi BM25: a non-binary model”. Cambridge University Press. nlp.stanford.edu/IR-book
  5. Connelly, S. (2018). “Practical BM25 – Part 2: The BM25 Algorithm and its Variables.” Elastic blog, 19 April 2018. elastic.co/blog/practical-bm25-part-2-the-bm25-algorithm-and-its-variables. Accessed 5 October 2026.
  6. Karpukhin, V. et al. (2020). “Dense Passage Retrieval for Open-Domain Question Answering.” Proceedings of EMNLP 2020. arxiv.org/abs/2004.04906
  7. Thakur, N. et al. (2021). “BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models.” NeurIPS 2021 Datasets and Benchmarks Track. arxiv.org/abs/2104.08663

Version history

  1. 1.0First published.