Skip to content
Changelog

MCP tools now return audio

September 16, 20261 min read

When an MCP tool returns audio, Agno's MCPTools decodes it and adds it to run.audio. Say your agent calls a text-to-speech server. Until now, MCPTools discarded the audio and told the model the content type was unsupported.

from pathlib import Path
 
from agno.agent import Agent
from agno.tools.mcp import MCPTools
 
async with MCPTools(transport="streamable-http", url="https://tts.example.com/mcp") as tts:
    agent = Agent(tools=[tts])
    run = await agent.arun("Read the release summary aloud")
 
    for audio in run.audio or []:
        Path(f"summary.{audio.format}").write_bytes(audio.content)

Each audio item keeps its MIME type, and MCPTools derives the format from it, so audio/mpeg becomes mp3 and audio/wav becomes wav. When MCPTools can't decode the audio, it reports that to the model and keeps the rest of the tool's result.

Learn more about MCP tools in the documentation.

Frequently asked questions

Connect the MCP server with MCPTools and run the agent. MCPTools decodes the audio the tool returns and attaches it to the run output, so you can read it from run.audio.

MCPTools keeps the audio's MIME type and reads the format from it. So audio/mpeg comes out as mp3, and audio/wav as wav.

MCPTools reports the audio it can't decode to the model and keeps the rest of the tool's result.

Shipped around the same time