Contents Database

Track and manage the content you've added to your knowledge base.

Contents Database is an optional component that tracks what you've added to your knowledge base. While the vector database stores embeddings for search, Contents Database stores metadata about each piece of content: what it is, when you added it, and its processing status.

from agno.knowledge.knowledge import Knowledge
from agno.db.postgres import PostgresDb
from agno.vectordb.pgvector import PgVector

knowledge = Knowledge(
    vector_db=PgVector(table_name="vectors", db_url=db_url),
    content_db=PostgresDb(db_url=db_url),  # Enables content tracking
)

content_db is the preferred constructor spelling; contents_db remains a supported read/write alias. The management methods below apply to ordinary inserted content. When page_store is configured, use published-page synchronization to change coordinated pages and vectors instead.

Why Use Contents DB

Without Contents DB, you can search your knowledge base but can't see what's in it or manage individual pieces of content.

With Contents DB, you get:

  • Visibility: See all content that's been added, track processing status, view metadata
  • Management: Delete specific content and automatically clean up associated vectors
  • Updates: Edit names, descriptions, and metadata without rebuilding the knowledge base
  • Filtering: Use agentic filtering to filter search results by metadata

Contents DB is required for agentic filtering and the AgentOS Knowledge UI.

Setup

Agno supports multiple database backends:

from agno.db.postgres import PostgresDb

contents_db = PostgresDb(
    db_url="postgresql+psycopg://user:pass@localhost:5432/db",
    knowledge_table="knowledge_contents"  # Optional custom table name
)

Common backends include PostgreSQL, SQLite, MySQL, MongoDB, Redis, Valkey, DynamoDB, and Firestore. See database providers for the complete list.

Managing Content

Add Content with Metadata

knowledge.insert(
    name="Product Manual",
    path="docs/manual.pdf",
    metadata={"department": "engineering", "version": "2.1"}
)

List Content

contents, total_count = knowledge.get_content(
    limit=20,
    page=1,
    sort_by="created_at",
    sort_order="desc"
)

for content in contents:
    print(content.name, content.status, content.created_at)

Get Content by ID

content = knowledge.get_content_by_id(content_id)

print(content.name)         # Content name
print(content.description)  # Description
print(content.metadata)     # Custom metadata
print(content.file_type)    # File type (.pdf, .txt, etc.)
print(content.size)         # File size in bytes
print(content.status)       # Processing status
print(content.created_at)   # When it was added
print(content.updated_at)   # Last modification

Content Status

Every content row records how ingestion ended. status is a ContentStatus value from agno.knowledge.content:

StatusMeaning
processingIngestion is still running.
completedEvery chunk was embedded and written to the vector database.
partialSome chunks were embedded and are searchable. The rest failed and cannot be retrieved.
failedNo chunks are retrievable.

For partial and failed, status_message explains what happened. A failed embedding names the embedder, the failure reason, the HTTP status, the provider's message with credentials redacted, and the next step:

Embedding failed for "handbook.pdf" (0 of 12 chunks embedded). Embedder: OpenAI text-embedding-3-small.
Reason: rate_limit (HTTP 429). Provider said: Rate limit reached for text-embedding-3-small.
Retrying did not succeed after 4 attempts. The embedding provider rate-limited this request; wait for
the limit to reset, or lower the embedder batch size. Re-ingest /docs/handbook.pdf once the cause is resolved.

A partial message reports the shortfall, such as 7 of 10 chunks were embedded; 3 failed and are not retrievable.

from agno.knowledge.content import ContentStatus

contents, _ = knowledge.get_content()
for content in contents:
    if content.status in (ContentStatus.PARTIAL, ContentStatus.FAILED):
        print(content.name, content.status.value, content.status_message)

Vector stores that embed server-side (LlamaIndex, LangChain, LightRAG) do not expose per-chunk embeddings, so their content is marked completed once the write succeeds.

Retry Failed Embeddings

Ingestion does not retry embedding failures by default. Set max_embedding_retries to retry the vector database write before any status is recorded:

knowledge = Knowledge(
    vector_db=PgVector(table_name="vectors", db_url=db_url),
    content_db=PostgresDb(db_url=db_url),
    max_embedding_retries=3,       # Up to 3 extra attempts
    embedding_retry_backoff=1.0,   # Wait 1s, then 2s, then 4s
)

A retry re-embeds the whole document, so a late failure in a large file bills every chunk again. Authentication failures and oversized chunks fail on the first attempt, because the same request fails every time.

Re-ingest Incomplete Content

Insert the same content again after fixing the cause. Content whose stored status is partial or failed is re-ingested even with skip_if_exists=True:

knowledge.insert(
    name="Product Manual",
    path="docs/manual.pdf",
    skip_if_exists=True,  # Skips completed content, re-ingests partial or failed content
)

For vector databases without upsert, Agno embeds the new chunks first, then deletes the chunks left by the earlier attempt, so a second failure does not remove what was already searchable. Uploaded file bytes are not stored, so re-ingesting an upload means supplying the file again.

Delete Content

Deleting content automatically:

  1. Removes the content metadata from Contents DB
  2. Deletes associated vectors from the vector database
  3. Maintains consistency between both databases
# Delete specific content
knowledge.remove_content_by_id(content_id)

# Delete all content
knowledge.remove_all_content()

Filter by Metadata

# Get available filter keys
valid_filters = knowledge.get_valid_filters()

# Search with filters
results = knowledge.search(
    query="technical documentation",
    filters={"department": "engineering"}
)

Schema

Contents DB stores the following fields for each piece of content:

FieldTypeDescription
idstrUnique identifier
namestrContent name
descriptionstrContent description
metadatadictCustom metadata
typestrContent type
sizeintFile size in bytes
linked_tostrID of linked content
access_countintNumber of times accessed
statusstrProcessing status: processing, completed, partial, or failed
status_messagestrFailure or partial-ingestion details
created_atintCreated timestamp
updated_atintUpdated timestamp
external_idstrExternal ID for integrations like LightRAG

AgentOS Integration

Contents DB is required for the AgentOS Knowledge UI. With it, the web interface provides:

  • Content Browser: View all uploaded content with metadata
  • Upload Interface: Add new content through the web UI
  • Status Monitoring: Processing status and error details
  • Metadata Editor: Update content metadata through forms
  • Search and Filtering: Find content by metadata attributes
  • Bulk Operations: Manage multiple content items at once
from agno.os import AgentOS
from agno.agent import Agent

knowledge = Knowledge(
    vector_db=PgVector(table_name="vectors", db_url=db_url),
    content_db=PostgresDb(db_url=db_url),
)

agent = Agent(name="Knowledge Agent", knowledge=knowledge)

agent_os = AgentOS(
    id="knowledge-demo",
    agents=[agent],
)

app = agent_os.get_app()

See AgentOS Knowledge Management for more details.

Next Steps

Content deletion returns a boolean. Check the result of remove_content_by_id() or remove_all_content() before treating the operation as complete; failed deletion can preserve affected content rows for a retry.