Skip to content

Glossary Information gain

Information gain

Definition

Information gain measures how much new information a document adds compared with the documents a user has already seen on the same topic; the concept comes from a Google patent family and is not confirmed as part of the ranking system.

On this page 5
  1. What information gain means
  2. Where the concept comes from and where the evidence ends
  3. Why it matters
  4. Good practice
  5. Common mistakes
In brief

A concept taken from a Google patent family that describes how much new information a document adds compared with what the user has already read on the same topic.

What information gain means

The term comes from information theory, where it describes how much a new piece of data reduces uncertainty. Applied to web documents, the idea is easy to state: if someone has already read three pages on a topic, the fourth adds something only if it contains material the earlier ones did not have.

The reference point is what matters. Information gain does not describe absolute originality but a difference relative to a set of documents already seen. A text can be very well written and add zero new information if it repeats what the reader has just found elsewhere. That same text would be highly valuable to someone arriving at the topic with no prior reading.

In industry usage the term stands in for original contribution: data nobody else publishes, a test run by the person who signs the piece, an interview, a calculation that does not appear at competitors. That translation is convenient to work with, although it simplifies the technical definition, which depends on each person's reading history rather than on a fixed comparison between pages.

Where the concept comes from and where the evidence ends

The source is a Google LLC patent family titled «Contextual estimation of link information gain», with a priority date of 18 October 2018. The first grant, US11354342B2, dates from June 2022; a later continuation, US12013887B2, was granted in June 2024 and remains in force. The text defines the score as one «indicative of additional information that is included in a given document beyond information contained in other documents that were already presented to the user».

The described procedure feeds semantic representations of the documents already seen and of the candidate documents into a machine learning model, which returns a score. The patent places the mechanism mainly in conversational assistants, where the system decides which document to present next, and adds that search results may be ranked in part according to those scores.

This is where the evidence ends. A patent proves that someone registered an idea, not that the idea runs in a product. Google files many procedures that are never deployed, and it does not publish the list of those it actually uses. In its public documentation for creators, the company never uses the expression «information gain». What it does write is a self-assessment question: «Does the content provide original information, reporting, research, or analysis?», alongside another asking whether a text that draws on other sources «avoids simply copying or rewriting those sources, and instead provides substantial additional value and originality». That language resembles the idea in the patent, but it confirms neither a score nor a system by that name.

Why it matters

Even though the score is unconfirmed, the criterion behind it helps decide what gets published. Putting out the fifteenth version of the same listicle consumes budget and rarely moves a domain, because the reader found that information before arriving. The useful question is not whether the text reads well, but what it contains that the top ten results do not.

The topic has gained weight with generative systems. A model summarising several sources has reason to cite the one that supplies a data point missing from the rest, because that is what justifies an extra reference. A text that repeats the consensus is absorbed into the summary without any need for a link.

There is an economic reading too. Recycled content is cheap to produce and its return has fallen. Running your own survey, a measured comparison or a documented case costs more and creates material that others cite, with links and mentions as a side effect. Information gain sums up that difference in two words, which is why it sits comfortably in a meeting, as long as nobody presents it as a metric Google publishes.

Good practice

  • Read the top ten results before writing and note what each one claims; the gap that remains is the starting point for your text.
  • Include at least one original element per article: a measurement, a screenshot of a real process, a figure from your own analytics account or the answer of a source you asked.
  • Date your original data and document the method and sample size, so others can cite it without suspicion.
  • Put the new contribution in the upper part of the page, because readers and summarising systems process the first paragraphs first.
  • Update old texts by adding new material rather than rewriting the same paragraphs in different words.
  • When a topic is already well covered by others and you have nothing to add, link the source and spend the effort on different content.

Common mistakes

  • Presenting information gain as a confirmed ranking factor. Google has not acknowledged any system by that name in its public documentation.
  • Confusing it with length. Adding a thousand words of filler lowers information density instead of raising it.
  • Trusting tools that promise to «measure information gain» without disclosing which corpus they compare against; the figure they return is the vendor's own model.
  • Chasing novelty at the cost of accuracy, publishing striking claims that no source supports.
  • Applying the criterion to transactional pages, where the user expects price, availability and delivery terms rather than an unseen angle.
Manuel Riveiro Rodriguez CEO & Digital Strategist

A technical audit covers this and everything else in one pass.

Request an audit

Frequently asked

Is information gain a Google ranking factor?

There is no confirmation. The concept appears in a Google patent family granted between 2022 and 2024, and a patent records an idea without proving it is used. The public documentation for creators never uses the term, although it does ask whether content provides original information and analysis.

How can information gain be measured?

No official measure exists. In practice you compare your own draft with the results already holding the top positions and count how many claims, figures or examples appear in none of them. It is a manual count, checkable by anyone and enough to decide whether publishing is worthwhile.

Does a longer text provide more information gain?

There is no relationship. Length measures characters, not contribution. A six-hundred-word article with an original test adds more than a four-thousand-word one summarising what is already published. Filler also makes it harder for the reader to find the new part.

Is the concept useful for AI-generated answers?

It is useful as an editorial criterion. A system writing a summary from several pages has reason to cite the one supplying a data point missing from the rest. No provider publicly documents a score of this kind inside its systems, so treat it as a working hypothesis.

How does it differ from duplicate content?

Duplicate content describes near-identical texts, often within one domain, and raises a technical problem of URL selection. Information gain describes how much a text adds compared with what third parties have already published, even when there is no literal overlap between them.

Sources

  1. Google LLC patent «Contextual estimation of link information gain», priority 18 October 2018 and granted in June 2022, where the information gain score is defined.
  2. A continuation in the same family, granted in June 2024 and in force, describing the ranking of documents by their information gain score.
  3. A second continuation of the patent family, useful for showing that Google kept the filing active for years without ever announcing a rollout.
  4. Google documentation for creators, with the self-assessment questions on original information and substantial added value, and with no mention of the term information gain.