Agent with Image Input

Send raw image bytes to a Llama agent alongside a web search tool.

Code

from pathlib import Path

from agno.agent import Agent
from agno.media import Image
from agno.models.meta import Llama
from agno.tools.websearch import WebSearchTools
from agno.utils.media import download_image

agent = Agent(
    model=Llama(id="Llama-4-Maverick-17B-128E-Instruct-FP8"),
    tools=[WebSearchTools()],
    markdown=True,
)

image_path = Path(__file__).parent.joinpath("sample.jpg")

download_image(
    url="https://upload.wikimedia.org/wikipedia/commons/0/0c/GoldenGateBridge-001.jpg",
    output_path=str(image_path),
)

# Read the image file content as bytes
image_bytes = image_path.read_bytes()

agent.print_response(
    "Tell me about this image and give me the latest news about it.",
    images=[
        Image(content=image_bytes),
    ],
    stream=True,
)

These examples require access to the Meta Llama API, an API key, and the chosen model enabled for your account. Check the model catalog while signed in, and replace the example ID if your account offers a different compatible model. Installing llama-api-client alone does not grant API access.

Usage

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Set your LLAMA API key

export LLAMA_API_KEY=YOUR_API_KEY

Install dependencies

uv pip install llama-api-client ddgs agno

Run Agent

Save the code above as image_input_bytes.py, then run:

python image_input_bytes.py