Reranking

Reorder knowledge search results with a reranker on Knowledge, including candidate pool sizing, MMR diversity, recency boosting, and migration from vector database rerankers.

Pass a reranker to Knowledge. Every search fetches a wider candidate pool from the vector database, reranks it, and returns the top max_results documents.

from agno.knowledge.knowledge import Knowledge
from agno.knowledge.reranker.cohere import CohereReranker
from agno.vectordb.pgvector import PgVector, SearchType

knowledge = Knowledge(
    vector_db=PgVector(
        table_name="docs",
        db_url="postgresql+psycopg://ai:ai@localhost:5532/ai",
        search_type=SearchType.hybrid,
    ),
    reranker=CohereReranker(model="rerank-v3.5"),
    max_results=5,
)

knowledge.insert(url="https://docs.agno.com/introduction")

results = knowledge.search("What is Agno?")

Install the dependencies, run PostgreSQL with pgvector, and set the OpenAI key for PgVector's default embedder and the Cohere key for the reranker:

uv pip install -U agno cohere openai pgvector psycopg sqlalchemy
export OPENAI_API_KEY=your_openai_api_key_here
export CO_API_KEY=your_cohere_api_key_here

How It Works

  1. Knowledge asks the vector database for reranker.search_limit(max_results) candidates instead of max_results.
  2. The reranker scores and reorders the candidates.
  3. Knowledge trims the reordered list to max_results.

The same steps run for search() and asearch(), and for agents that search through the search_knowledge_base tool. Async searches call reranker.arerank(), which runs a sync reranker in a worker thread so it does not block the event loop.

Size the Candidate Pool

Reranking can only surface documents that the vector database returned. Two fields on every reranker control how many candidates Knowledge fetches:

FieldDefaultEffect
candidate_multiplier3Candidates fetched per requested result. Set 1 to reorder without widening.
max_candidates100Ceiling on the widened fetch. It never reduces the fetch below max_results.

With the defaults, a search for 5 results fetches 15 candidates, and a search for 50 results fetches 100.

reranker = CohereReranker(
    model="rerank-v3.5",
    candidate_multiplier=5,
    max_candidates=50,
)

A wider pool gives the reranker more documents to promote. It also sends more documents to the reranking model on every search, which increases latency and, for hosted rerankers, cost.

Leave top_n unset on a reranker attached to Knowledge. Knowledge already trims to max_results, and a smaller top_n returns fewer documents than the search requested.

Available Rerankers

RerankerImportRuns
CohereRerankeragno.knowledge.reranker.cohereCohere Rerank API
SentenceTransformerRerankeragno.knowledge.reranker.sentence_transformerLocal cross-encoder model
InfinityRerankeragno.knowledge.reranker.infinitySelf-hosted Infinity server
AwsBedrockReranker, CohereBedrockReranker, AmazonRerankeragno.knowledge.reranker.aws_bedrockAWS Bedrock rerank models
MMRRerankeragno.knowledge.reranker.mmrLocal diversity selection. See Diversify Results With MMR.
RecencyRerankeragno.knowledge.reranker.recencyLocal blend of search score and document age. See Boost Recent Documents.

To write your own, subclass agno.knowledge.reranker.base.Reranker and implement rerank(query, documents, limit=None).

Diversify Results With MMR

Vector search often returns several chunks that say the same thing. MMRReranker applies Maximal Marginal Relevance: it picks the most relevant candidate first, then scores each remaining candidate by its relevance to the query minus its similarity to the closest document already picked. lambda_mult weights the two terms.

from agno.knowledge.reranker.mmr import MMRReranker

knowledge = Knowledge(
    vector_db=PgVector(table_name="docs", db_url=db_url),
    reranker=MMRReranker(lambda_mult=0.5),
)
FieldDefaultEffect
lambda_mult0.51.0 ranks by relevance alone. 0.0 ranks by difference alone.
candidate_multiplier5MMR needs a wider pool than the other rerankers to find different documents.
top_nNoneLeave unset on Knowledge.

MMR compares the embeddings of the candidates, so it needs two things from the vector database:

RequirementVector databases
Returns embeddings on search resultsPgVector, Qdrant vector and hybrid search, Chroma, LanceDB, and Elasticsearch. Pinecone returns them when built with return_vectors=True.
Does not return embeddingsMilvus, MongoDB, Redis, Valkey, and Qdrant keyword search
Attaches an embedder to search resultsRequired to embed the query with the model the documents were indexed with. Upstash hosted embeddings do not expose one.

When a requirement is missing, MMR raises a ValueError instead of returning unreranked results.

Use the returned order. Each document's reranking_score holds its MMR score at the moment it was picked. These scores are not in descending order and can be negative, so sorting by them discards the diversity ordering.

Boost Recent Documents

Vector search has no notion of time, so an outdated document ranks as well as the revision that replaced it. RecencyReranker blends each document's search score with an exponential decay on its age:

score = (1 - weight) * relevance + weight * 0.5 ** (age_days / half_life_days)

An older document still wins when it is clearly more relevant.

from agno.knowledge.reranker.recency import RecencyReranker

knowledge = Knowledge(
    vector_db=PgVector(table_name="docs", db_url=db_url, return_updated_at=True),
    reranker=RecencyReranker(half_life_days=30, weight=0.3),
)
FieldDefaultEffect
weight0.30.0 ranks by relevance alone. 1.0 ranks by age alone.
half_life_days30Age at which the recency term drops to half.
timestamp_key"updated_at"Metadata key read for each document's timestamp.

The reranker reads a document's timestamp from these sources, in order:

  1. The updated_at key in the document metadata. Set it when inserting content, as an ISO 8601 string, a datetime, or epoch seconds or milliseconds.
  2. The time the vector database stored the row. PgVector reports it when built with return_updated_at=True. Re-inserting content under the same name replaces the row and resets its age.

return_updated_at is off by default because the timestamp is added to document metadata, which reaches the model's context. Documents without a timestamp receive no recency term. If no result has a timestamp, the reranker logs a warning and keeps the relevance order.

Relevance comes from the score the vector database reports. Scores outside 0 to 1, such as raw BM25 scores, are divided by the highest score in the pool. If the database reports no scores, the reranker uses rank position instead.

Migrate From a Vector Database Reranker

Setting reranker on a vector database is deprecated and will be removed in a future release. The parameter still works, but Agno logs a deprecation warning when it is set. Move the reranker to Knowledge:

# Before
knowledge = Knowledge(
    vector_db=PgVector(
        table_name="docs",
        db_url=db_url,
        reranker=CohereReranker(model="rerank-v3.5"),
    ),
)

# After
knowledge = Knowledge(
    vector_db=PgVector(table_name="docs", db_url=db_url),
    reranker=CohereReranker(model="rerank-v3.5"),
)
Vector database rerankerKnowledge reranker
StatusDeprecatedSupported
Vector databasesAdapters that implement the parameterEvery vector database
Candidate poolReranks the max_results the database returnsFetches a wider pool sized by the reranker

If both are set, Knowledge applies only its own reranker, skips the vector database reranker for its searches, and logs a warning. Calls made directly on the vector database, such as vector_db.search(), still use the vector database reranker.

Next Steps