> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stmlink.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> 对外开放的服务端接口有两组前缀，都用同一套鉴权：`/server/v1/...`（SRTC 与 SMeeting 的主接口）和 `/stm/srvapi/v1/...`（SMeeting 的用户体系，服务端极简对接会用到）。鉴权是 app_id + nonce + timestamp + signature 四个请求头，用 app_key 做 HMAC-SHA256 签名，只能从业务方自己的后端调用。除这两组前缀外的接口均为内部接口，不要建议客户调用。 Public server APIs use two path prefixes with the same authentication: `/server/v1/...` (the main APIs of both SRTC and SMeeting) and `/stm/srvapi/v1/...` (the SMeeting user system, used by server-side low-code integration). Authenticate with four request headers, app_id + nonce + timestamp + signature, where signature is HMAC-SHA256 keyed with app_key; call these APIs only from the customer's own backend. Any other path is internal: never suggest calling it.
> app_key 是服务端密钥，绝不能出现在客户端代码、前端配置或移动 App 里。客户端加入频道用的 token 必须由业务方后端签发后下发（SRTC 走 `/server/v1/channel/grant`，SMeeting 走 `/stm/srvapi/v1/member/grant`）。 app_key is a server-side secret and must never appear in client code, frontend config, or a mobile app. The token a client uses to join must be issued by the customer's backend and passed down to the client (SRTC: `/server/v1/channel/grant`; SMeeting: `/stm/srvapi/v1/member/grant`).
> SRTC 与 SMeeting 是上下两层不同的产品，术语不通用：SRTC 是音视频底座，说「频道 channel」「加入 / 退出」；SMeeting 建在 SRTC 之上，说「房间 room」「会议 meeting」「进入 / 退出」。回答时按用户所在的层用对应术语，不要把「房间」「会议」安到 SRTC 的接口上，也不要用「频道」「加入 / 离开」描述 SMeeting 的概念（接口标识符原样保留）。 SRTC and SMeeting are two separate layers with different terminology. SRTC is the audio/video foundation: it has channels, and users join and leave a channel. SMeeting is built on top of SRTC: it has rooms and meetings, and members enter and exit a meeting. Answer in the terms of the layer the user is working with: never apply "room" or "meeting" to SRTC APIs, and never describe SMeeting concepts in prose with "channel", "join", or "leave" (API identifiers such as `force_join` keep their literal names).
> 同一能力在各端 SDK 里的包名、类名、方法名并不相同。写示例代码时请使用文档中该端自己的 API，不要把一个端的写法套到另一个端上。苹果平台每个产品都有两套 SDK（Swift 原生与 Objective-C），两套 API 不能混用。 Package, class, and method names differ between platform SDKs for the same capability. In sample code, use the API documented for that platform; never carry one platform's code over to another. On Apple platforms each product ships two SDKs (native Swift and Objective-C) whose APIs must not be mixed.

# Audio and video data

> Media data types of the SRTC Python SDK: AudioTrack for publishing (write to push PCM, clear to interrupt, backpressure and real-time pacing), the fields, timeline, and silence-frame semantics of received AudioFrame / VideoFrame, and the AudioFormat format parameter.

## AudioFormat

```python theme={null}
@dataclass(frozen=True)
class AudioFormat:
    sample_rate: int = 16000
    channels: int = 1
```

PCM format. All PCM sent and received by the SDK is **S16LE, interleaved**.

| Field | Description |
| - | - |
| `sample_rate` | Sample rate, any value (commonly 8000 / 16000 / 24000 / 44100 / 48000) |
| `channels` | Number of channels, 1 or 2 |

***

## AudioTrack

The local audio track returned by `Channel.publish_audio`.

### write

```python theme={null}
async def write(pcm: bytes | np.ndarray) -> None
```

Writes PCM (in the `audio_format` specified when publishing) of any length. The SDK splits it into 20 ms frames and sends them **at real-time pace**.

* `bytes`: S16LE interleaved PCM; the byte count must be a multiple of `2 × number of channels`, otherwise `ValueError` is raised
* `np.ndarray`: converted to `int16`; for stereo, flatten it in interleaved order

When the audio not yet sent exceeds `max_buffer_seconds`, `write` waits until there's room in the buffer before returning (backpressure).

<Note>
  **When idle, the SDK keeps sending silence frames**, keeping RTP timestamps in sync with real time. So however long the pause between two sentences, the next sentence isn't dropped by the remote side as late data.
</Note>

### clear

```python theme={null}
def clear() -> None
```

Immediately discards all audio not yet sent. Call it when the user barges in; see [Voice AI agent guide](/en/rtc/python/advanced/ai-agent#barge-in-interruption).

### wait\_for\_playout

```python theme={null}
async def wait_for_playout() -> None
```

Waits until all written audio has been sent.

### buffered\_seconds

```python theme={null}
@property
def buffered_seconds() -> float
```

Duration of audio not yet sent (seconds). Greater than 0 means the agent is still "speaking".

***

## AudioFrame

A decoded frame of remote audio.

| Field | Type | Description |
| - | - | - |
| `uid` | `str` | Speaker's uid |
| `track_id` | `str` | Track ID (the same user may publish multiple audio tracks) |
| `pcm` | `bytes` | S16LE interleaved PCM |
| `sample_rate` | `int` | Sample rate (equal to the `audio_format` given when joining) |
| `channels` | `int` | Number of channels |
| `samples` | `int` | Samples per channel |
| `pts` | `int` | Position of the frame's first sample on this track's timeline (a sample index at `sample_rate`, starting from 0) |
| `is_silence` | `bool` | `True` means a silence frame filled in by the SDK during the remote side's DTX silence |
| `duration` | `float` | Frame duration (seconds), `samples / sample_rate` |

| Method | Description |
| - | - |
| `to_numpy()` | Converts to an `int16` array with shape `(samples, channels)`, zero-copy |

On the same track, two adjacent frames satisfy `next.pts == prev.pts + prev.samples`, so the timeline is continuous.

***

## VideoFrame

A frame of remote video.

| Field | Type | Description |
| - | - | - |
| `uid` | `str` | Publisher's uid |
| `track_id` | `str` | Track ID |
| `codec` | `int` | Encoding format, a `Codec` enum value (commonly `H264` / `VP8`) |
| `encoded` | `bytes` | Encoded data (Annex-B for H264), always provided |
| `rtp_timestamp` | `int` | 90 kHz RTP timestamp |
| `image` | `np.ndarray \| None` | Decoded RGB24 image with shape `(height, width, 3)`. Only available with `decode_video=True` at join time |
| `width` / `height` | `int` | Image size; 0 when not decoded |

<Note>
  A decoded image is provided for every frame. Vision models usually don't need such a high frame rate, so sample frames as needed (for example, one frame per second).
</Note>
