I/O Streams
I/O streams carry user input to BidiAgent and deliver its output to your application. An input stream reads text, images, or audio from a source such as a keyboard, microphone, or WebSocket. An output stream handles the agent’s events to play audio, display text, or update your interface.
Connect these streams with agent.run(), which starts the agent and its streams. The inputs and outputs lists can each contain multiple streams. The agent reads each input independently and sends every event to all outputs; each output chooses which events to handle.
A single run can span multiple model responses. This example connects an input and output stream and uses a 30-second timeout to end the conversation:
import asyncio
from strands.bidi.agent import BidiAgentfrom strands.bidi.types import InputStream, OutputStream
async def run_conversation( input_stream: InputStream, output_stream: OutputStream,) -> None: agent = BidiAgent() try: await asyncio.wait_for( agent.run(inputs=[input_stream], outputs=[output_stream]), timeout=30, ) except asyncio.TimeoutError: passWhen the timeout expires, asyncio.wait_for() cancels the run. It waits for the agent and its streams to stop, then raises TimeoutError.
Your application can also cancel the task running run() when a user presses a stop button or a client disconnects. To let users end the conversation through a tool, see End the conversation by voice.
The following sections show how to use Strands’ built-in AudioIO and ConsoleIO for audio and terminal interactions.
Audio I/O
Section titled “Audio I/O”Use AudioIO to capture microphone audio and play the assistant’s speech through your speakers. It displays speech transcripts and tool call names in the terminal and handles barge-in automatically by stopping interrupted playback.
Install PortAudio, then install the SDK with audio I/O support:
pip install "strands-agents[bidi,bidi-io,bidi-pyaudio]"Create an AudioIO instance and pass its input and output streams to run():
import asyncio
from strands.bidi.agent import BidiAgentfrom strands.bidi.io import AudioIO
async def main() -> None: agent = BidiAgent() audio_io = AudioIO()
try: await asyncio.wait_for( agent.run(inputs=[audio_io.input()], outputs=[audio_io.output()]), timeout=30, ) except asyncio.TimeoutError: pass
asyncio.run(main())Configure voice and audio settings on the model; AudioIO uses that configuration for capture and playback. For device settings and other I/O options, see the AudioIO API reference.
Audio processing
Section titled “Audio processing”Use a headset to limit echo, or enable microphone processing for speaker playback. Processing applies echo cancellation, noise suppression, and automatic gain control.
Install the additional processing dependency:
pip install "strands-agents[bidi-aec]"Then set audio_processor=True on your AudioIO instance:
from strands.bidi.io import AudioIO
audio_io = AudioIO(audio_processor=True)Use input and output streams from the same AudioIO instance for echo cancellation. Processing requires mono microphone audio; echo cancellation also requires mono playback.
To customize processing, such as disabling echo cancellation when using a headset, pass an AudioProcessorConfig as audio_processor.
Console I/O
Section titled “Console I/O”Use ConsoleIO to type messages and display text, reasoning, speech transcripts, and tool call names.
This example uses OpenAI Realtime with text output. Install the terminal and model dependencies:
pip install "strands-agents[bidi-openai,bidi-io]"Set OPENAI_API_KEY to your API key, then run:
import asyncio
from strands.bidi.agent import BidiAgentfrom strands.bidi.io import ConsoleIOfrom strands.bidi.models import OpenAIRealtimeModel
async def main() -> None: model = OpenAIRealtimeModel( model_id="gpt-realtime-2.1", transcription_model_id=None, params={"output_modalities": ["text"]}, ) agent = BidiAgent(model=model) console_io = ConsoleIO(placeholder="Type a message…")
try: await asyncio.wait_for( agent.run(inputs=[console_io.input()], outputs=[console_io.output()]), timeout=30, ) except asyncio.TimeoutError: pass
asyncio.run(main())Press Enter to send a message. ConsoleIO displays the text and transcripts the model emits.
Use placeholder to customize the input hint. The show_text, show_reasoning, show_transcript, and show_tools options all default to True; set an option to False to hide that content. See ConsoleIOConfig for the configuration reference.
Type and speak
Section titled “Type and speak”To type and speak in the same conversation, share a ConsoleIO with AudioIO and register both input streams:
import asyncio
from strands.bidi.agent import BidiAgentfrom strands.bidi.io import AudioIO, ConsoleIO
async def main() -> None: agent = BidiAgent() console_io = ConsoleIO(placeholder="Type or speak…") audio_io = AudioIO(console=console_io)
try: await asyncio.wait_for( agent.run( inputs=[console_io.input(), audio_io.input()], outputs=[audio_io.output()], ), timeout=30, ) except asyncio.TimeoutError: pass
asyncio.run(main())audio_io.output() handles both playback and the shared console display, so register it as the only output for those destinations.
Custom I/O
Section titled “Custom I/O”Connect your own input sources and output destinations to run() with asynchronous callables:
- An input callable takes no arguments and returns one
BidiAgentInputper call. - An output callable receives one
BidiOutputEventand chooses which events to handle.
The following FastAPI example exchanges text messages with a WebSocket client, using OpenAI Realtime configured for text output.
Install the server and model dependencies:
pip install "strands-agents[bidi-openai]" fastapi "uvicorn[standard]"Set OPENAI_API_KEY to your API key, then save the following as server.py:
from fastapi import FastAPI, WebSocket, WebSocketDisconnect
from strands.bidi.agent import BidiAgentfrom strands.bidi.models import OpenAIRealtimeModelfrom strands.bidi.types import BidiOutputEvent, BidiTextBlockEvent
app = FastAPI()
@app.websocket("/conversation")async def conversation(websocket: WebSocket) -> None: await websocket.accept() model = OpenAIRealtimeModel( model_id="gpt-realtime-2.1", transcription_model_id=None, params={"output_modalities": ["text"]}, ) agent = BidiAgent(model=model)
async def send_text(event: BidiOutputEvent) -> None: if isinstance(event, BidiTextBlockEvent): await websocket.send_text(event.text)
try: await agent.run( inputs=[websocket.receive_text], outputs=[send_text], ) except WebSocketDisconnect: passStart the server with uvicorn server:app --reload, then connect a WebSocket client to ws://localhost:8000/conversation. Send a message such as Say hello. to receive a completed text response. Disconnecting the client ends the run and cleans up the agent.
To display text as it arrives or play audio, handle the corresponding stream events in your output callable. The agent waits for all outputs to finish each event, so queue audio for playback and return promptly. Clear queued audio on barge-in.
For streams that need setup or cleanup, subclass InputStream or OutputStream and implement __call__(). Override start(agent) to prepare resources after the agent starts and stop() to release them.
If startup or processing fails, run() cancels the remaining tasks, stops the streams and agent, and propagates the error. Make stop() safe to call after partial setup.