Data labeling and classification

Classify data, extract records, build preference datasets, and review labels with agents.

ML and data teams use agents to turn large collections of unstructured inputs into datasets for training, evaluation, search, and automation. Agno applies the same Pydantic schema pattern across text, images, audio, video, and PDFs, with workflows for parallel labeling, review, and conditional adjudication.

output_schema requests structured output and Agno attempts to parse it. A failed or unparseable run can leave content as text. Check RunStatus.completed and the expected Pydantic type before reading fields, indexing, or writing downstream. These examples stop on failure; an application can instead send the original input to a review queue. Schema validation checks the shape and constraints, so verify extracted facts against the source separately.

from agno.run.base import RunStatus

from typing import Literal

from agno.agent import Agent
from pydantic import BaseModel, Field


class Classification(BaseModel):
    label: Literal["positive", "negative", "neutral"] = Field(
        ..., description="The assigned sentiment label"
    )


agent = Agent(
    model="google:gemini-3.5-flash",
    instructions="You classify product reviews by sentiment.",
    output_schema=Classification,
)

result_run = agent.run("Broken on arrival, total waste of money.")
if result_run.status != RunStatus.completed or not isinstance(result_run.content, Classification):
    raise RuntimeError("No validated Classification; send the input to review before continuing")
result = result_run.content
# Classification(label='negative')

Agent.run() attempts to validate the response against the Pydantic model; the check above rejects failed runs and unparsed output. Change the schema and instructions to extract fields, label spans, score responses, or rank preferences.

What you can build

OutcomeInputPattern
Classify feedback, support requests, or documentsText or filesClassification
Extract contacts, line items, action items, or attributesAny supported modalityData extraction
Detect entities and PII spansTextClassification and span labeling
Search an image library in natural languageImagesImage Search
Build pairwise preference datasetsPrompt and two responsesPreference data
Score generated responses against a rubricPrompt and responseLLM as judge
Review and adjudicate important labelsAny labeling taskQuality pipeline
Label images, audio, video, and PDFsMedia or filesMultimodal inputs

The data labeling cookbook contains 18 labeling patterns and a complete Image Search application. The patterns cover text, image, audio, video, and document inputs.

Run the examples

Create an environment and install the Google provider:

uv venv .venv --python 3.12
source .venv/bin/activate
uv pip install "agno[google]"
export GOOGLE_API_KEY="..."

The quality review workflow also uses Anthropic:

uv pip install "agno[anthropic,sqlite]"
export ANTHROPIC_API_KEY="..."

Developer Resources