Query Transformation

Rewrite the search query before it reaches the vector database with a query transformer on Knowledge, including HyDE hypothetical-answer search, model selection, and search-type behavior.

Pass a query transformer to Knowledge. Every search rewrites the query before it reaches the vector database. HyDE searches with an LLM-written hypothetical answer instead of the question.

from agno.agent import Agent
from agno.knowledge.knowledge import Knowledge
from agno.knowledge.query_transformer import HyDE
from agno.models.openai import OpenAIResponses
from agno.vectordb.pgvector import PgVector

knowledge = Knowledge(
    vector_db=PgVector(
        table_name="docs",
        db_url="postgresql+psycopg://ai:ai@localhost:5532/ai",
    ),
    query_transformer=HyDE(),
)

knowledge.insert(
    name="q3-review",
    text_content="Enterprise renewals slipped last quarter as three large accounts delayed signing.",
)

agent = Agent(model=OpenAIResponses(id="gpt-5.4"), knowledge=knowledge)
agent.print_response("Why did revenue drop?")

Install the dependencies, run PostgreSQL with pgvector, and set the OpenAI key used by the embedder and the model:

uv pip install -U agno openai pgvector psycopg sqlalchemy
export OPENAI_API_KEY=your_openai_api_key_here

How HyDE Works

Questions and the passages that answer them rarely share wording. "Why did revenue drop?" embeds some distance from "Enterprise renewals slipped last quarter". HyDE (Hypothetical Document Embeddings) asks an LLM to write a short passage that answers the question, then searches with that passage. The passage resembles the documents being searched for, so its embedding lands closer to real answers.

  1. Knowledge calls query_transformer.transform() (or atransform() for async searches) with the query.
  2. HyDE sends its prompt to a model and trims the reply to max_characters, cutting at a sentence or word break.
  3. The vector database searches with the generated passage.
  4. If Knowledge has a reranker, it scores the results against the original question.

The generated passage never reaches the user, and it does not need to be correct. It only has to look like the kind of document being searched for.

The same steps run for search() and asearch(), for agents and teams that search through the search_knowledge_base tool, and for add_knowledge_to_context=True.

Choose the Model

HyDE writes the passage with its own model. It does not use the model of the agent or team that triggered the search. When model is unset, HyDE sets it to OpenAIResponses(id="gpt-5.4") on construction, the same default that Agent and Team use. Constructing HyDE() without a model raises an ImportError if the openai package is not installed.

Set model to write passages with a smaller, faster model than the one answering the question:

HyDE(model=OpenAIResponses(id="gpt-5.6-luna"))

When an agent or team searches through search_knowledge_base or with add_knowledge_to_context=True, the generation call is counted in the run's metrics.

Configure HyDE

FieldDefaultEffect
modelOpenAIResponses(id="gpt-5.4")Model that writes the hypothetical answer. See Choose the Model.
promptBuilt-in promptReplaces the built-in prompt. Must contain {query}, or HyDE raises a validation error.
include_queryFalseSearch with the question followed by the passage instead of the passage alone.
max_characters2000Ceiling on the passage length. Must be greater than 0.

A custom prompt replaces the default instruction entirely, including its instruction to answer without hedging. Write that instruction into your own prompt:

HyDE(
    prompt=(
        "Write a two-sentence excerpt from an internal finance report that answers "
        "the question. Do not hedge. Reply with the excerpt only.\n\nQuestion: {query}"
    ),
)

Search Types

The transform interacts with each search type differently:

SearchBehavior
SearchType.vectorSearches with the transformed query.
SearchType.hybridBoth the vector and keyword halves receive the transformed query. Set include_query=True so the keyword half still matches the words the user typed.
SearchType.keywordSkips the transform and searches with the original query. Keyword search requires every query term in a document, so a generated passage would match nothing.
Published pagesSearches with the original query and adds the transformed query as an extra phrasing. Results for each phrasing are fused.

Knowledge logs a warning when HyDE has include_query=False and the vector database uses hybrid or keyword search, or a page store is configured.

from agno.vectordb.search import SearchType

knowledge = Knowledge(
    vector_db=PgVector(
        table_name="docs",
        db_url="postgresql+psycopg://ai:ai@localhost:5532/ai",
        search_type=SearchType.hybrid,
    ),
    query_transformer=HyDE(include_query=True),
)

Failure and Cost

HyDE makes one LLM call per search, which adds latency and cost to every search.

If the call fails, HyDE logs a warning and the search uses the original query. An empty reply also falls back to the original query. If a transformer raises, Knowledge logs the error and searches with the original query. A provider outage degrades results without breaking search.

Write a Query Transformer

Subclass QueryTransformer and implement transform(). Return the query to search with. Returning query unchanged is valid.

from typing import Any, Optional

from agno.knowledge.query_transformer import QueryTransformer


class ExpandAcronyms(QueryTransformer):
    def transform(self, query: str, run_response: Optional[Any] = None) -> str:
        return query.replace("PTO", "paid time off")


knowledge = Knowledge(vector_db=vector_db, query_transformer=ExpandAcronyms())

A transformer that calls an LLM configures its own model as a field, as HyDE does. Pass run_response to the model call so it is counted in the run's metrics. The default atransform() runs transform() in a worker thread so it does not block the event loop. Override atransform() to call a model asynchronously.

Next Steps