Audio to Text Transcription
Transcribe a multi-speaker audio recording into a speaker-labeled conversation using Gemini.
Create an agent that transcribes audio conversations and identifies the different speakers.
Code
import requests
from agno.agent import Agent
from agno.media import Audio
from agno.models.google import Gemini
agent = Agent(
model=Gemini(id="gemini-3.5-flash"),
markdown=True,
)
url = "https://agno-public.s3.us-east-1.amazonaws.com/demo_data/QA-01.mp3"
response = requests.get(url)
audio_content = response.content
# Give a transcript of this audio conversation. Use speaker A, speaker B to identify speakers.
agent.print_response(
"Give a transcript of this audio conversation. Use speaker A, speaker B to identify speakers.",
audio=[Audio(content=audio_content)],
stream=True,
)Usage
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateInstall dependencies
uv pip install -U agno google-genai requestsExport your GOOGLE API key
export GOOGLE_API_KEY="your_google_api_key_here"Run Agent
python audio_to_text.py