Knowledge-Level Reranking

Rerank a widened candidate pool from Qdrant with Cohere by setting the reranker on Knowledge.

CohereReranker is set on Knowledge, so every search fetches a wider candidate pool from Qdrant and Cohere reorders it before the top results are returned.

knowledge_level_reranking.py
"""
Knowledge-Level Reranking
=========================
A reranker set on Knowledge runs after the vector db returns results, rather than
inside the vector db itself. Two differences follow from that:

1. It works with any vector db, so the same reranker moves between backends.
2. It widens the fetch, so the reranker chooses from a real pool rather than only
   reordering what the vector db already returned. candidate_multiplier (capped by
   max_candidates) is set on the reranker itself.

The widened fetch is what makes ordering strategies possible: a reranker can only
surface a document that was retrieved in the first place.

See also: 03_reranking.py for vector db level reranking.
"""

import asyncio

from agno.agent import Agent
from agno.knowledge.knowledge import Knowledge
from agno.knowledge.reranker.cohere import CohereReranker
from agno.models.openai import OpenAIResponses
from agno.vectordb.qdrant import Qdrant

# ---------------------------------------------------------------------------
# Setup
# ---------------------------------------------------------------------------

qdrant_url = "http://localhost:6333"

knowledge = Knowledge(
    vector_db=Qdrant(collection="knowledge_reranking_demo", url=qdrant_url),
    reranker=CohereReranker(
        # Candidates fetched per requested result, so Cohere can rescue a document that
        # plain search ranked outside max_results. Costs that many times the API calls,
        # so lower it to 1 to only reorder what the vector db already returned.
        candidate_multiplier=3,
        # Ceiling on the widened fetch, once the multiplier is above 1.
        max_candidates=100,
    ),
)

agent = Agent(
    model=OpenAIResponses(id="gpt-5.6-luna"),
    knowledge=knowledge,
    search_knowledge=True,
    markdown=True,
)


async def main():
    await knowledge.ainsert(
        url="https://agno-public.s3.amazonaws.com/recipes/ThaiRecipes.pdf"
    )

    # Retrieves 25 candidates, reranks them, returns the top 5.
    results = await knowledge.asearch("What are some Thai curry dishes?", max_results=5)
    print("Reranked results:")
    for document in results:
        print(f"  {document.name}")

    await agent.aprint_response("What are some Thai curry dishes?", stream=True)


if __name__ == "__main__":
    asyncio.run(main())

What Happens

With candidate_multiplier=3, knowledge.asearch(..., max_results=5) fetches 15 candidates from Qdrant, sends them to Cohere in one rerank request, and returns the top 5. The source comment says 25 candidates, which would require candidate_multiplier=5. Set candidate_multiplier=1 to reorder only the documents Qdrant would have returned.

The agent searches the same knowledge base through search_knowledge_base, so its results are reranked too. See Reranking for how the pool is sized.

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno cohere openai pypdf qdrant-client

Export your API keys

export CO_API_KEY="your_co_api_key_here"
export OPENAI_API_KEY="your_openai_api_key_here"

Run Qdrant

docker run -d --name qdrant -p 6333:6333 qdrant/qdrant:latest

Run the example

Save the code above as knowledge_level_reranking.py, then run:

python knowledge_level_reranking.py

Full source: cookbook/07_knowledge/02_building_blocks/07_knowledge_level_reranking.py