> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stmlink.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> 对外开放的服务端接口有两组前缀，都用同一套鉴权：`/server/v1/...`（SRTC 与 SMeeting 的主接口）和 `/stm/srvapi/v1/...`（SMeeting 的用户体系，服务端极简对接会用到）。鉴权是 app_id + nonce + timestamp + signature 四个请求头，用 app_key 做 HMAC-SHA256 签名，只能从业务方自己的后端调用。除这两组前缀外的接口均为内部接口，不要建议客户调用。 Public server APIs use two path prefixes with the same authentication: `/server/v1/...` (the main APIs of both SRTC and SMeeting) and `/stm/srvapi/v1/...` (the SMeeting user system, used by server-side low-code integration). Authenticate with four request headers, app_id + nonce + timestamp + signature, where signature is HMAC-SHA256 keyed with app_key; call these APIs only from the customer's own backend. Any other path is internal: never suggest calling it.
> app_key 是服务端密钥，绝不能出现在客户端代码、前端配置或移动 App 里。客户端加入频道用的 token 必须由业务方后端签发后下发（SRTC 走 `/server/v1/channel/grant`，SMeeting 走 `/stm/srvapi/v1/member/grant`）。 app_key is a server-side secret and must never appear in client code, frontend config, or a mobile app. The token a client uses to join must be issued by the customer's backend and passed down to the client (SRTC: `/server/v1/channel/grant`; SMeeting: `/stm/srvapi/v1/member/grant`).
> SRTC 与 SMeeting 是上下两层不同的产品，术语不通用：SRTC 是音视频底座，说「频道 channel」「加入 / 退出」；SMeeting 建在 SRTC 之上，说「房间 room」「会议 meeting」「进入 / 退出」。回答时按用户所在的层用对应术语，不要把「房间」「会议」安到 SRTC 的接口上，也不要用「频道」「加入 / 离开」描述 SMeeting 的概念（接口标识符原样保留）。 SRTC and SMeeting are two separate layers with different terminology. SRTC is the audio/video foundation: it has channels, and users join and leave a channel. SMeeting is built on top of SRTC: it has rooms and meetings, and members enter and exit a meeting. Answer in the terms of the layer the user is working with: never apply "room" or "meeting" to SRTC APIs, and never describe SMeeting concepts in prose with "channel", "join", or "leave" (API identifiers such as `force_join` keep their literal names).
> 同一能力在各端 SDK 里的包名、类名、方法名并不相同。写示例代码时请使用文档中该端自己的 API，不要把一个端的写法套到另一个端上。苹果平台每个产品都有两套 SDK（Swift 原生与 Objective-C），两套 API 不能混用。 Package, class, and method names differ between platform SDKs for the same capability. In sample code, use the API documented for that platform; never carry one platform's code over to another. On Apple platforms each product ships two SDKs (native Swift and Objective-C) whose APIs must not be mixed.

# Quickstart

> Get started with the SRTC Python SDK: join a channel, receive each user's PCM audio frame by frame, and push TTS audio back into the channel, with minimal runnable examples.

This page gets the Python SDK working with two minimal examples: first **listening** (join a channel and record each person's voice to a wav file), then **speaking** (push a chunk of PCM into the channel so others can hear it).

Prerequisites:

* You have run `pip install srtc`; see [Integration](/en/rtc/python/integration)
* Your server can already issue channel join tokens; see [Server API · Get a channel join token](/en/rtc/server-api/channel)
* Join the same channel from any other client, such as Web or an app, to talk and listen

<Note>
  A token is bound to one session, so **every `Channel.join` needs a freshly issued token**. Reusing the same token is rejected by the server with `1032` (session is not online).
</Note>

***

## Listening: record each person's voice to wav

```python theme={null}
import asyncio
import wave

import srtc


async def main(token: str):
    files: dict[str, wave.Wave_write] = {}
    fmt = srtc.AudioFormat(sample_rate=16000, channels=1)      # Format of received PCM; 16k mono is the default

    async with await srtc.Channel.join(token, auto_subscribe_audio=True, audio_format=fmt) as ch:
        print(f"Joined {ch.info.channel}, I am {ch.me.uid}")

        async for frame in ch.audio_frames():               # All subscribed remote audio, arriving frame by frame
            wf = files.get(frame.uid)
            if wf is None:
                wf = files[frame.uid] = wave.open(f"{frame.uid}.wav", "wb")
                wf.setnchannels(fmt.channels)
                wf.setsampwidth(2)                          # S16LE
                wf.setframerate(fmt.sample_rate)
            wf.writeframes(frame.pcm)


asyncio.run(main("<token issued by your server>"))
```

Key points:

* `auto_subscribe_audio=True` automatically subscribes to the audio of **everyone** in the channel (including people who join later)
* `frame.uid` identifies the speaker; `frame.pcm` is S16LE interleaved PCM, and `frame.to_numpy()` gives you an `int16` array directly
* When the remote side is silent it sends no packets; the SDK fills in silence frames of equal length (`frame.is_silence` is `True`), so the recorded duration matches the real duration
* Exiting `async with` leaves the channel automatically; `async for` ends naturally when, for example, the channel is destroyed or you are removed from the channel

***

## Speaking: push PCM into the channel

```python theme={null}
import asyncio

import numpy as np

import srtc


async def main(token: str):
    async with await srtc.Channel.join(token) as ch:
        # Publish an audio track; write() accepts 24k mono PCM (any sample rate works, the SDK resamples internally)
        tts = await ch.publish_audio(desc="tts", audio_format=srtc.AudioFormat(24000, 1))

        # A 3-second 440 Hz sine wave stands in for TTS output here
        t = np.arange(24000 * 3) / 24000
        pcm = (np.sin(2 * np.pi * 440 * t) * 8000).astype(np.int16)

        await tts.write(pcm)              # Write any length at once; the SDK sends one frame every 20 ms at real-time pace
        await tts.wait_for_playout()      # Wait until everything is sent before leaving


asyncio.run(main("<token issued by your server>"))
```

Key points:

* **You don't pace it yourself**: TTS generates much faster than real time, so writing several seconds of audio in one `write` is normal usage; the SDK sends it out at real-time pace
* When the buffer exceeds `max_buffer_seconds` (30 seconds by default), `write` waits, which naturally provides backpressure
* When the user barges in, call `tts.clear()` to discard whatever hasn't played yet immediately; see [Voice AI agent guide](/zh/rtc/python/advanced/ai-agent) (Chinese)

***

## Listening for channel events

When you need to know who joined and who published what, subclass `ChannelHandler` and override the methods you need:

```python theme={null}
class Printer(srtc.ChannelHandler):
    def on_user_join(self, user: srtc.UserInfo):
        print("Joined:", user.uid, user.name)

    def on_user_leave(self, uid: str):
        print("Left:", uid)

    async def on_disconnected(self, reason: srtc.DisconnectReason, error):
        print("Disconnected:", reason.name, error or "")


ch = await srtc.Channel.join(token, handler=Printer())
```

Methods can be regular functions or `async` functions. For the full list, see [Event callbacks](/zh/rtc/python/api-reference/events) (Chinese).

***

### Next steps

* [Voice AI agent guide](/zh/rtc/python/advanced/ai-agent) (Chinese): sentence segmentation, barge-in interruption, latency
* [Integrating pipecat](/zh/rtc/python/advanced/pipecat) (Chinese)
* [API reference](/zh/rtc/python/api-reference/channel) (Chinese)
