Image Input Bytes

Send raw image bytes to a Vertex AI Claude agent and search the web for context.

Pass raw image bytes with Image(content=...).

Code

from pathlib import Path

from agno.agent import Agent
from agno.media import Image
from agno.models.vertexai.claude import Claude
from agno.tools.websearch import WebSearchTools
from agno.utils.media import download_image

agent = Agent(
    model=Claude(id="claude-sonnet-4-6"),
    tools=[WebSearchTools()],
    markdown=True,
)

image_path = Path(__file__).parent.joinpath("sample.jpg")

download_image(
    url="https://upload.wikimedia.org/wikipedia/commons/0/0c/GoldenGateBridge-001.jpg",
    output_path=str(image_path),
)

# Read the image file content as bytes
image_bytes = image_path.read_bytes()

agent.print_response(
    "Tell me about this image and give me the latest news about it.",
    images=[
        Image(content=image_bytes),
    ],
    stream=True,
)

Usage

Complete the Vertex project, model enablement, permissions, and CLI prerequisites first. Use a region that serves Sonnet 4.6.

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Set your environment variables

export CLOUD_ML_REGION="us-east5"
export GOOGLE_CLOUD_PROJECT=xxx

Authenticate your CLI session

gcloud auth application-default login
You don't need to authenticate your CLI every time.

Install dependencies

uv pip install -U 'anthropic[vertex]' ddgs agno

Run Agent

Save the code above as image_input_bytes.py, then run:

python image_input_bytes.py