- You have run
pip install srtc; see Integration - Your server can already issue channel join tokens; see Server API · Get a channel join token
- Join the same channel from any other client, such as Web or an app, to talk and listen
A token is bound to one session, so every
Channel.join needs a freshly issued token. Reusing the same token is rejected by the server with 1032 (session is not online).Listening: record each person’s voice to wav
auto_subscribe_audio=Trueautomatically subscribes to the audio of everyone in the channel (including people who join later)frame.uididentifies the speaker;frame.pcmis S16LE interleaved PCM, andframe.to_numpy()gives you anint16array directly- When the remote side is silent it sends no packets; the SDK fills in silence frames of equal length (
frame.is_silenceisTrue), so the recorded duration matches the real duration - Exiting
async withleaves the channel automatically;async forends naturally when, for example, the channel is destroyed or you are removed from the channel
Speaking: push PCM into the channel
- You don’t pace it yourself: TTS generates much faster than real time, so writing several seconds of audio in one
writeis normal usage; the SDK sends it out at real-time pace - When the buffer exceeds
max_buffer_seconds(30 seconds by default),writewaits, which naturally provides backpressure - When the user barges in, call
tts.clear()to discard whatever hasn’t played yet immediately; see Voice AI agent guide (Chinese)
Listening for channel events
When you need to know who joined and who published what, subclassChannelHandler and override the methods you need:
async functions. For the full list, see Event callbacks (Chinese).
Next steps
- Voice AI agent guide (Chinese): sentence segmentation, barge-in interruption, latency
- Integrating pipecat (Chinese)
- API reference (Chinese)