Skip to main content

AudioFormat

PCM format. All PCM sent and received by the SDK is S16LE, interleaved.

AudioTrack

The local audio track returned by Channel.publish_audio.

write

Writes PCM (in the audio_format specified when publishing) of any length. The SDK splits it into 20 ms frames and sends them at real-time pace.
  • bytes: S16LE interleaved PCM; the byte count must be a multiple of 2 × number of channels, otherwise ValueError is raised
  • np.ndarray: converted to int16; for stereo, flatten it in interleaved order
When the audio not yet sent exceeds max_buffer_seconds, write waits until there’s room in the buffer before returning (backpressure).
When idle, the SDK keeps sending silence frames, keeping RTP timestamps in sync with real time. So however long the pause between two sentences, the next sentence isn’t dropped by the remote side as late data.

clear

Immediately discards all audio not yet sent. Call it when the user barges in; see Voice AI agent guide.

wait_for_playout

Waits until all written audio has been sent.

buffered_seconds

Duration of audio not yet sent (seconds). Greater than 0 means the agent is still “speaking”.

AudioFrame

A decoded frame of remote audio. On the same track, two adjacent frames satisfy next.pts == prev.pts + prev.samples, so the timeline is continuous.

VideoFrame

A frame of remote video.
A decoded image is provided for every frame. Vision models usually don’t need such a high frame rate, so sample frames as needed (for example, one frame per second).