Agent with Image Input
Send raw image bytes to a Llama agent alongside a web search tool.
Code
from pathlib import Path
from agno.agent import Agent
from agno.media import Image
from agno.models.meta import Llama
from agno.tools.websearch import WebSearchTools
from agno.utils.media import download_image
agent = Agent(
model=Llama(id="Llama-4-Maverick-17B-128E-Instruct-FP8"),
tools=[WebSearchTools()],
markdown=True,
)
image_path = Path(__file__).parent.joinpath("sample.jpg")
download_image(
url="https://upload.wikimedia.org/wikipedia/commons/0/0c/GoldenGateBridge-001.jpg",
output_path=str(image_path),
)
# Read the image file content as bytes
image_bytes = image_path.read_bytes()
agent.print_response(
"Tell me about this image and give me the latest news about it.",
images=[
Image(content=image_bytes),
],
stream=True,
)These examples require access to the Meta Llama API, an API key, and the chosen model enabled for your account. Check the model catalog while signed in, and replace the example ID if your account offers a different compatible model. Installing llama-api-client alone does not grant API access.
Usage
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateSet your LLAMA API key
export LLAMA_API_KEY=YOUR_API_KEYInstall dependencies
uv pip install llama-api-client ddgs agnoRun Agent
Save the code above as image_input_bytes.py, then run:
python image_input_bytes.py