← Back to blog
Mara Ellison

AI Rerankers for SEO: How They Work and What We Tested

See a reproducible AI reranker test where a false passage outranked the correct answer. Learn what relevance scores can and cannot tell SEO teams.

CitationsStrategyData
Summarize withChatGPTClaudePerplexityGrokMarkdown
On this page

An AI reranker takes a set of candidate documents and orders them by relevance to a query. It can help a search application choose which passages to send to a language model. For SEO teams, this offers a useful way to examine whether a page actually answers a specific question.

For this article, we ran a small local test with three questions and four fictional passages. The model preferred a false answer for one question, despite assigning it a high relevance score. The complete inputs, scores, and reproduction code appear below.

That result illustrates a practical limitation: use a relevance test to examine how a model orders passages, and verify the facts through a separate process. It does not reveal how Google or a public answer engine would rank the same content.

What is an AI reranker?

A reranker evaluates candidates that have already been collected. A retrieval system might find a broad selection of potentially useful documents. The reranker then applies another relevance assessment to that selection.

The Sentence Transformers retrieve-and-rerank documentation describes a common approach: an initial retriever finds candidates, and a cross-encoder examines query-and-document pairs to order them. Evaluating pairs is more computationally expensive than comparing precomputed document embeddings, which helps explain the two stages.

That is one documented architecture. Reranking systems differ, and the same architecture should not be assumed for every commercial search product.

An example service is Cohere Rerank. It accepts a query and a list of documents and returns a relevance ordering. Its API reference documents the inputs and returned relevance scores.

For a content team, the useful question is simple: given the buyer's actual question, does this passage provide the information needed to answer it?

Retrieval, reranking, and generation do different jobs

A simplified application could follow this sequence:

User question
    ↓
Retrieve candidate documents
    ↓
Rerank candidates for relevance
    ↓
Select context within the application's limits
    ↓
Generate an answer, potentially with citations

This diagram describes a possible system, not a disclosed map of a particular public answer engine.

StageQuestion the stage addressesWhat a publisher can investigate
RetrievalWhich documents might help?Whether content is accessible and addresses the query's subject
RerankingWhich available candidates appear most relevant?Whether passages satisfy the question's specific requirements
Context selectionWhich information should fit into the answer's working context?Whether important facts remain understandable when read as a passage
GenerationWhat answer should the system produce from its available information?Whether the final response represents the source accurately
Citation presentationWhich sources are shown to the user?Which URLs appear, and whether they support the associated claims

A relevant page can fail to appear for several reasons. It may not be available to the retriever. It may compete with a more suitable source. Its useful information may be omitted from the selected context. The generated answer may use other information.

The absence of a citation does not identify which stage caused the omission. A publisher generally cannot observe all those stages in a public engine.

What rerankers mean for SEO

The practical implication is to make the relationship between the question and the evidence clear.

A page about payroll exports can repeat “payroll integration” throughout and still fail to answer whether a particular product supports an automatic connection. Another page can use fewer matching phrases and explain the supported transfer method, required fields, and limitations.

That distinction matters to a reader regardless of which model evaluates the page.

Start with real buyer prompts, then identify the conditions within each question. These might include:

  • The named product or category.
  • The task the reader needs to complete.
  • A required capability.
  • A restriction such as region, subscription plan, or transfer method.
  • The evidence needed to decide whether the answer applies.

Treat these conditions as requirements to address where relevant. Do not insert them into every paragraph or invent capabilities to produce a closer semantic match.

Google's generative AI optimization guide rejects the idea that publishers need special AI wording or tiny content chunks. Write complete answers for readers; a third-party model score is not a Google optimization target.

Our small test: a relevant answer can still be false

We evaluated cross-encoder/ms-marco-TinyBERT-L2-v2, a publicly available passage-ranking model, on September 12, 2026. We fixed the following inputs before running it and retained all twelve query–passage scores.

Declared fictional facts: HarborRoster exports approved hours as CSV on the Teams plan. An administrator imports the file into CedarPay; there is no automatic synchronization. Managers approve shift changes before the schedule updates.

The model received each query and candidate passage, without a separate copy of that fact sheet. Passage C deliberately contradicts the declared facts. The test asks how a relevance model handles the candidates; it does not ask the model to verify them against an authoritative source.

IDExact candidate passage
AHarborRoster streamlines payroll workflows with powerful integrations for growing teams.
BHarborRoster exports approved hours as a CSV file on the Teams plan. An administrator downloads the file and imports it into CedarPay. It does not synchronize automatically.
CHarborRoster automatically synchronizes approved hours with CedarPay on the Basic plan. No CSV download or manual import is required.
DIn HarborRoster, managers review requested shift changes and approve or reject them before the schedule updates.

We used these exact questions:

  • Q1: Does HarborRoster automatically synchronize approved hours with CedarPay?
  • Q2: Which HarborRoster plan includes CSV export for CedarPay?
  • Q3: How do managers approve shift changes in HarborRoster?

Observed scores

These are raw model logits, rounded to three decimals. Higher means the model scored that passage as more relevant within this test. They are not probabilities, factual-confidence scores, or predicted citation rates.

QuestionA: broad claimB: correct export detailsC: false automatic connectionD: shift approvalsHighest-scored passage
Q1−7.1569.47210.123−3.353C, which contradicts the declared facts
Q2−5.1178.5377.410−5.670B
Q3−4.810−1.633−2.7809.069D

For Q1, the model scored the false passage above the correct explanation. The topic match did not establish truth. For Q2 and Q3, the highest-scored passages answered the question consistently with the declared facts.

This small result supports a narrow conclusion: this model's relevance ordering was insufficient to verify factual correctness in this candidate set. It does not estimate the model's overall error rate, compare current reranking vendors, or measure Google, ChatGPT, or AI citations.

Reproduce the test

We used Python 3.10.4, PyTorch 2.5.1, and Transformers 4.57.6, on CPU in evaluation mode with float32 weights and one thread. The pinned model revision is in the code. We set a 512-token limit; the longest encoded pair contained 57 tokens, so this test did not truncate a candidate.

Install the compatible library versions in an isolated environment, then run the following code. The first run downloads the public model. The code includes every query and candidate, so you can inspect the full output rather than relying on the rounded table.

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "cross-encoder/ms-marco-TinyBERT-L2-v2"
revision = "81d1926f67cb8eee2c2be17ca9f793c7c3bd20cc"

passages = [
    "HarborRoster streamlines payroll workflows with powerful integrations for growing teams.",
    "HarborRoster exports approved hours as a CSV file on the Teams plan. An administrator downloads the file and imports it into CedarPay. It does not synchronize automatically.",
    "HarborRoster automatically synchronizes approved hours with CedarPay on the Basic plan. No CSV download or manual import is required.",
    "In HarborRoster, managers review requested shift changes and approve or reject them before the schedule updates.",
]
queries = [
    "Does HarborRoster automatically synchronize approved hours with CedarPay?",
    "Which HarborRoster plan includes CSV export for CedarPay?",
    "How do managers approve shift changes in HarborRoster?",
]

tokenizer = AutoTokenizer.from_pretrained(
    model_id, revision=revision, trust_remote_code=False
)
model = AutoModelForSequenceClassification.from_pretrained(
    model_id, revision=revision, trust_remote_code=False,
    use_safetensors=True
).to("cpu").eval()
torch.set_num_threads(1)

for query in queries:
    inputs = tokenizer(
        [query] * len(passages), passages,
        padding=True, truncation=True, max_length=512,
        return_tensors="pt"
    )
    with torch.no_grad():
        scores = model(**inputs).logits.squeeze(-1).tolist()
    ranked = sorted(zip("ABCD", scores), key=lambda item: item[1], reverse=True)
    print(query, ranked)

The model is compact and publicly reproducible, which makes it suitable for this teaching example. We did not select it as a proxy for a public search engine or tune the passages after seeing the results.

Apply the result to your content review

Keep the subject, claim, and conditions together. A buyer asking about automatic synchronization needs the transfer method and its limitation in the answer, even when that limitation makes the product a poor fit.

Use a direct section opening such as: “HarborRoster requires a manual CSV transfer in this example.” Follow it with the procedure and documented alternatives. Repeating “automatic payroll integration” cannot make an unsupported capability true.

When reviewing a real page, link the capability to current documentation and resolve any conflicting description elsewhere on the site. This is useful factual maintenance regardless of whether a relevance score improves.

How to design a larger reranker evaluation

The example above is too small to evaluate production quality. For a broader test, choose questions from actual buyer research and preserve a separate set for final evaluation.

  1. Define correctness before scoring. Record the facts and conditions a satisfactory answer must address. Keep partial, unrelated, and contradictory passages distinct.
  2. Freeze the candidate set. Save exact texts, source URLs, capture dates, and preprocessing. A different shortlist can change the outcome before reranking begins.
  3. Save the configuration. Record the model version, input limit, settings, and complete scores. Check whether any passage was truncated.
  4. Review the results against the facts. Count useful, correct answers within the number of candidates your application uses. Inspect failures instead of optimizing only for a higher score.
  5. Evaluate revised content on unseen questions. Repeated edits against a few familiar questions can fit the test without improving the page for other readers.

A developer can also compare retrieval approaches. Lexical retrieval matches terms; dense retrieval compares learned representations; a cross-encoder assesses query–passage pairs together. The Sentence Transformers documentation provides a concrete retrieve-and-rerank implementation.

For a public answer engine, you usually cannot see the complete candidate set. An absent citation therefore cannot tell you whether retrieval, context selection, or another step caused the omission.

Measure published outcomes separately

After publishing a useful revision, evaluate the actual channels where readers encounter the content.

For organic search, watch relevant query impressions, clicks, and landing-page engagement. For AI answers, track a stable prompt set across repeated runs and comparable windows. Review the cited sources and whether the answer represents your product correctly.

Citlyze's prompt tracking can support observation of answer visibility. Use citation-rate measurement to define what counts as an appearance and which runs belong in the denominator.

Those observations describe outcomes. They do not expose a hidden reranker score or establish which internal stage selected a source.

Maintain a change log with the page revision, publication date, affected prompt group, and other changes that could influence the comparison. A visibility increase following a revision is worth investigating; timing alone does not demonstrate causation.

Rerankers are not embedding models, and neither is a target

An embedding model produces representations that can be compared for similarity search. A reranker reassesses candidate relevance for a specific query. Some documented pipelines use embeddings for initial retrieval and a cross-encoder for reranking, but that is not the only possible arrangement, and a public answer engine does not publish which one it uses.

You can evaluate content against a specified model, especially if you operate the application that uses it. Optimize for correctly answering representative questions, and keep a separate evaluation set. A publisher cannot assume the same result transfers to unrelated public engines.

The same caution applies to format. Not every answer should be a short paragraph: some questions need a sentence, others require a procedure, a comparison, or qualifications. Keep the answer and its important conditions together, then use enough space to make the explanation complete. Structured data should accurately describe eligible content; it does not establish that a passage will be retrieved, reranked highly, included in context, or cited.

What to change first

Find an important page that receives a specific buyer question but answers it vaguely. Verify the facts, write a direct answer with its restrictions, link the supporting documentation, and measure the published outcome. A useful correction gives readers something concrete even when no visibility improvement follows.