Information Gain as a Ranking Signal: What It Means in Practice
Google holds an information gain patent, but never confirmed it as a live signal. What is documented, what is inferred, and how to add real gain.

Information gain is the amount of genuinely new information a page adds beyond what a reader has already seen on the same topic. Google holds a patent describing the concept — Contextual Estimation of Link Information Gain, filed in 2018 — but has never confirmed it as a named, live ranking system, and issued no guidance connecting it to any 2026 update. What is defensible is narrower and more useful than the industry’s version: content that only restates what already ranks has been losing ground for years, and the most reliable way to earn a position now is to put something on the page that is not already in the index.
This article separates the documented from the inferred, then gets to the part that actually matters — how to manufacture information gain when you are not sitting on proprietary data.
What is documented, and what is not
| Claim | Status |
|---|---|
| Google holds a patent referencing information gain | Documented. Contextual Estimation of Link Information Gain, filed 2018 |
| Google asks whether content “provides substantial value when compared to other pages in search results” | Documented. A real question from Google’s published self-assessment list |
| Information gain is a named live ranking system | Not confirmed. Google has never said this |
| The March 2026 core update made it the dominant signal | Not confirmed. Google issued no new guidance with that update at all |
| Specific percentages (“AI content lost 60–80%”) | Vendor estimates. No published methodology, sample or controls |
A patent is not a shipped feature. Google files thousands it never deploys, and the existence of one tells you an idea was considered, not that it is running. Anyone presenting the patent as proof that a system is live is overreaching.
The honest position: information gain is an excellent description of what recent core updates appear to reward, and it maps onto guidance Google has published since 2022. It is not a confirmed named signal. That distinction matters if you put it in a client report. The March 2026 core update piece covers how far the claims outran the evidence.
Why it describes something real anyway
Set the patent aside. Two things are independently true and neither is disputed.
Google assesses quality partly at site level. Google has confirmed this. A large volume of pages that add nothing does not merely fail individually; it weighs on the site as a whole. Publishing a competent restatement of the top ten is no longer neutral — it is a small liability.
Generative engines synthesise from the existing corpus. An AI answer is assembled from what is already indexed. Content that only recombines existing material is, by definition, already covered by the synthesis. The only thing an engine cannot generate from the corpus is something that is not in the corpus yet. That is information gain, and it is why it matters as much for AI-answer citation as for rankings.
Both point the same way regardless of what any patent is doing.
The six kinds of information gain
“Add original insight” is useless advice. These are the six forms that actually work, roughly in descending order of defensibility.
1. Proprietary data
Numbers only you have. Aggregate anonymised client results, your own platform data, a survey you ran. Highest value, hardest to fake, and impossible for a competitor to copy without doing the work.
2. Documented first-hand testing
You used the tool, ran the migration, tried the tactic — and you recorded what happened including what failed. Screenshots of your own screen, not stock imagery. This is what Google’s product review guidance has asked for explicitly, and it is the most accessible form for a small agency.
3. An original framework
A named, reusable way of thinking about the problem: a decision tree, a scoring rubric, a sequence with thresholds. If it is genuinely useful, other people cite it by name, which compounds.
4. Named expert attribution
A quote from a real, identifiable person with relevant standing — including yourself, if you have the credentials to back it. Peer-reviewed GEO research (Aggarwal et al., KDD 2024) found that adding quotations from named sources, alongside statistics and citations to authoritative sources, raised a page’s visibility in generative answers by up to around 40%.
5. Synthesis nobody else has done
Pulling scattered primary sources into one place with the contradictions resolved. Lower value than original data, but real — and often the most practical option for a topic you cannot test directly.
6. A genuinely better explanation
The most underrated. If everyone explains something badly and you explain it clearly, that is gain. It is the one form available on literally any topic.
How to check whether a draft actually has any
Run this before publishing. It takes ten minutes and kills a surprising number of drafts.
- Open the current top ten for your target query. Not the keyword — the live results.
- List every substantive claim in your draft. One line each.
- Mark each claim that already appears in at least one ranking page. Be honest; “I said it differently” does not count.
- Look at what is unmarked. That is your entire information gain.
- If nothing is unmarked, do not publish. You have written a paraphrase of the SERP. Either find something to add or spend the time on a topic where you have something.
The uncomfortable version of this test: if a reader who has already read the top three results learns nothing new from your page, why would Google rank it above them, and why would an AI engine cite it instead of them?
What information gain is not
- Not length. Adding 1,000 words of restatement reduces gain per word. There is no word count that creates it.
- Not a contrarian take for its own sake. Being wrong differently is not gain. A contrarian position needs to be defensible, and you should say what would change your mind.
- Not novelty of phrasing. Saying a common thing in an unusual way adds nothing.
- Not more subheadings. Structure helps retrieval; it does not create substance.
- Not invented statistics. The fastest way to destroy the credibility this is meant to build. If a claim needs a source and there is not one, cut the claim.
Building it into how you work
Information gain is a production problem more than a writing problem. If your process is “read the top ten, write a better version of them,” it cannot produce gain by construction — the inputs are the existing corpus.
Three changes that fix it structurally. Decide the gain before commissioning, at brief stage, so the writer knows what the page is for. Keep a running log of anything unusual you observe in client work — the crawl finding that surprised you, the fix that did not work. That log is your proprietary data source, and most agencies throw it away. Name your authors and give them real credentials, because attribution is what converts a claim into evidence; anonymous content cannot answer Google’s “who” question and an author page that signals genuine expertise is the mechanism.
The through-line across every confirmed 2026 update points the same direction, which what the 2026 updates have in common sets out in full: evidence over volume.

Where information gain actually comes from
- Proprietary data — Strongest. Numbers only you hold.
- First-hand testing — Strong. Documented, including what failed.
- Original framework — Strong. A named, reusable way to think.
- Named expert input — Good. A real person with standing.
- Novel synthesis — Moderate. Primary sources nobody combined.
- A clearer explanation — Underrated. Available on any topic.
Frequently asked questions
What is information gain in SEO?
Information gain is the amount of genuinely new information a page adds beyond what already exists in the top-ranking results for that query. It comes from information theory, where it measures how much uncertainty is reduced by new data. In content terms it means original data, first-hand testing, an original framework, named expert input, novel synthesis, or a materially better explanation.
Is information gain a confirmed Google ranking factor?
No. Google holds a patent titled Contextual Estimation of Link Information Gain, filed in 2018, but has never confirmed it as a named live ranking system and issued no guidance connecting it to any 2026 update. A patent shows an idea was considered, not that it shipped. What is documented is Google’s published self-assessment question asking whether content provides substantial value compared to other pages in search results.
How do I add information gain to my content?
Use one of six approaches: proprietary data only you hold, documented first-hand testing including what failed, an original reusable framework, quotes from named experts with relevant standing, synthesis of scattered primary sources nobody has combined, or a genuinely clearer explanation of something everyone explains badly. The last is available on any topic and is the most underrated.
How can I test whether my article has information gain?
Open the current top ten results for your target query, list every substantive claim in your draft, and mark each one that already appears in at least one ranking page. What remains unmarked is your entire information gain. If nothing remains, you have written a paraphrase of the search results, and publishing it adds a small liability rather than an asset.
Does information gain help with AI search visibility too?
Yes, and arguably more directly. Generative engines assemble answers from content already in the index, so material that only recombines existing sources is already covered by the synthesis. Research published at KDD 2024 found that adding statistics, quotations from named sources and citations to authoritative sources raised visibility in generative answers by up to around 40%.
Sources
- Creating helpful, reliable, people-first content — Google Search Central
- GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024)
- Google's Information Gain Patent — Search Engine Journal
Want this done on your site?
Every PalV’s DM engagement starts with a free audit of your actual website — a 12-point
crawl covering what is blocking indexation, on-page gaps against your primary keywords, speed
findings, and the three to five fixes worth making first. Delivered in two working days. No
payment details, and the findings are yours whether you hire us or not.