Skip to main content
This page gets the Python SDK working with two minimal examples: first listening (join a channel and record each person’s voice to a wav file), then speaking (push a chunk of PCM into the channel so others can hear it). Prerequisites:
A token is bound to one session, so every Channel.join needs a freshly issued token. Reusing the same token is rejected by the server with 1032 (session is not online).

Listening: record each person’s voice to wav

Key points:
  • auto_subscribe_audio=True automatically subscribes to the audio of everyone in the channel (including people who join later)
  • frame.uid identifies the speaker; frame.pcm is S16LE interleaved PCM, and frame.to_numpy() gives you an int16 array directly
  • When the remote side is silent it sends no packets; the SDK fills in silence frames of equal length (frame.is_silence is True), so the recorded duration matches the real duration
  • Exiting async with leaves the channel automatically; async for ends naturally when, for example, the channel is destroyed or you are removed from the channel

Speaking: push PCM into the channel

Key points:
  • You don’t pace it yourself: TTS generates much faster than real time, so writing several seconds of audio in one write is normal usage; the SDK sends it out at real-time pace
  • When the buffer exceeds max_buffer_seconds (30 seconds by default), write waits, which naturally provides backpressure
  • When the user barges in, call tts.clear() to discard whatever hasn’t played yet immediately; see Voice AI agent guide (Chinese)

Listening for channel events

When you need to know who joined and who published what, subclass ChannelHandler and override the methods you need:
Methods can be regular functions or async functions. For the full list, see Event callbacks (Chinese).

Next steps