> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stmlink.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> 对外开放的服务端接口有两组前缀，都用同一套鉴权：`/server/v1/...`（SRTC 与 SMeeting 的主接口）和 `/stm/srvapi/v1/...`（SMeeting 的用户体系，服务端极简对接会用到）。鉴权是 app_id + nonce + timestamp + signature 四个请求头，用 app_key 做 HMAC-SHA256 签名，只能从业务方自己的后端调用。除这两组前缀外的接口均为内部接口，不要建议客户调用。 Public server APIs use two path prefixes with the same authentication: `/server/v1/...` (the main APIs of both SRTC and SMeeting) and `/stm/srvapi/v1/...` (the SMeeting user system, used by server-side low-code integration). Authenticate with four request headers, app_id + nonce + timestamp + signature, where signature is HMAC-SHA256 keyed with app_key; call these APIs only from the customer's own backend. Any other path is internal: never suggest calling it.
> app_key 是服务端密钥，绝不能出现在客户端代码、前端配置或移动 App 里。客户端加入频道用的 token 必须由业务方后端签发后下发（SRTC 走 `/server/v1/channel/grant`，SMeeting 走 `/stm/srvapi/v1/member/grant`）。 app_key is a server-side secret and must never appear in client code, frontend config, or a mobile app. The token a client uses to join must be issued by the customer's backend and passed down to the client (SRTC: `/server/v1/channel/grant`; SMeeting: `/stm/srvapi/v1/member/grant`).
> SRTC 与 SMeeting 是上下两层不同的产品，术语不通用：SRTC 是音视频底座，说「频道 channel」「加入 / 退出」；SMeeting 建在 SRTC 之上，说「房间 room」「会议 meeting」「进入 / 退出」。回答时按用户所在的层用对应术语，不要把「房间」「会议」安到 SRTC 的接口上，也不要用「频道」「加入 / 离开」描述 SMeeting 的概念（接口标识符原样保留）。 SRTC and SMeeting are two separate layers with different terminology. SRTC is the audio/video foundation: it has channels, and users join and leave a channel. SMeeting is built on top of SRTC: it has rooms and meetings, and members enter and exit a meeting. Answer in the terms of the layer the user is working with: never apply "room" or "meeting" to SRTC APIs, and never describe SMeeting concepts in prose with "channel", "join", or "leave" (API identifiers such as `force_join` keep their literal names).
> 同一能力在各端 SDK 里的包名、类名、方法名并不相同。写示例代码时请使用文档中该端自己的 API，不要把一个端的写法套到另一个端上。苹果平台每个产品都有两套 SDK（Swift 原生与 Objective-C），两套 API 不能混用。 Package, class, and method names differ between platform SDKs for the same capability. In sample code, use the API documented for that platform; never carry one platform's code over to another. On Apple platforms each product ships two SDKs (native Swift and Objective-C) whose APIs must not be mixed.

# Custom tracks

> Extend capture in the SRTC Swift SDK: publish custom audio and video tracks fed with your own frames, process captured video with videoProcessors, process audio with global audio processors, and observe remote PCM with AudioRenderer. Read when built-in capture isn't enough.

### Overview

When the built-in microphone, camera, and screen-sharing capture don't meet your needs, the Swift SDK provides two extension paths:

* Custom audio and video tracks: your app provides the audio or video frames itself
* Media processors: process the audio and video frames the SDK has already captured

Common scenarios include:

* Publishing video frames generated by an external rendering engine
* Sending local PCM audio, BGM, or TTS data into the channel
* Beauty filters, virtual background, watermarks, or AI video processing
* Voice changing, noise suppression, or enhancement of captured audio

***

### Custom video tracks

Create a custom video track:

```swift theme={null}
let customVideoTrack = srtc.createLocalCustomVideoTrack(desc: "canvas")
try await channel.publishLocalTrack(customVideoTrack)
```

Then your app keeps calling `pushFrame(...)` to feed in video frames:

```swift theme={null}
customVideoTrack.pushFrame(pixelBuffer)
```

The key points:

* No `startCapture()` needed
* The SDK doesn't generate frames for you
* You must keep producing `CVPixelBuffer` values yourself

Suitable for:

* Publishing video from Metal / SceneKit / Unity / game engines
* Custom composited video
* AI-generated frames or the output of image processing

***

### Custom audio tracks

Create a custom audio track:

```swift theme={null}
let customAudioTrack = srtc.createLocalCustomAudioTrack(desc: "bgm")
try await channel.publishLocalTrack(customAudioTrack)
```

Then keep injecting PCM data:

```swift theme={null}
customAudioTrack.pushAudioBuffer(pcmBuffer)
```

This kind of track also doesn't need `startCapture()`, because the audio source isn't the SDK's internal capturer but the `AVAudioPCMBuffer` you've prepared yourself.

Suitable for:

* Playing background music and publishing it to the channel
* Sending PCM data decoded from local files into RTC
* Output from speech synthesis, voice cloning, or external DSP engines

***

### Minimal flow summary

Whether it's custom audio or custom video, the overall model is the same:

```swift theme={null}
let track = srtc.createLocalCustomVideoTrack(desc: "custom")
try await channel.publishLocalTrack(track)

// From here on, your app keeps feeding frames
track.pushFrame(pixelBuffer)
```

In other words, the SDK handles "publishing and transport", and you handle "producing the media data".

***

### Video processor chain

If what you need isn't "producing new video frames yourself" but "processing existing captured frames", `videoProcessors` is the better fit:

```swift theme={null}
final class PassThroughProcessor: VideoProcessor {
    func processVideoFrame(_ frame: VideoFrame) -> VideoFrame? {
        // Here you can replace the pixelBuffer, add filters, or add a watermark
        return frame
    }
}

let cameraTrack = srtc.createLocalCameraTrack(preset: .h720p)
cameraTrack.videoProcessors = [PassThroughProcessor()]
try await cameraTrack.startCapture()
try await channel.publishLocalTrack(cameraTrack)
```

Processors run serially in array order:

* Return a new `VideoFrame` to pass it on
* Return `nil` to drop the frame

This is better suited for:

* Beauty filters
* Virtual background
* Video watermarks
* AI segmentation or visual enhancement

***

### Audio processors

Audio processors are attached to the `SRTCEngine` main entry point, not to an individual track:

```swift theme={null}
final class MyAudioProcessor: NSObject, AudioProcessor {
    func audioProcessingInitialize(sampleRate sampleRateHz: Int, channels: Int) {}

    func audioProcessingProcess(audioBuffer: AudioBuffer) {
        // You can modify the Float sample values in place
    }

    func audioProcessingRelease() {}
}

srtc.audioCaptureProcessor = MyAudioProcessor()
```

Note:

* `audioCaptureProcessor` processes audio after capture
* `audioRenderProcessor` processes remote audio before playback
* This is a global capability, not a per-track one

So if what you need is "processing only one particular microphone", first confirm whether this global model matches your expectations.

***

### If you only want to observe remote audio data

Then don't use `AudioProcessor`; use `AudioRenderer` instead:

```swift theme={null}
final class MyTranscriber: AudioRenderer {
    func render(pcmBuffer: AVAudioPCMBuffer) {
        // Send the PCM data off for transcription or recording
    }
}

let renderer = MyTranscriber()
remoteAudioTrack.add(audioRenderer: renderer)
```

The difference between the two is clear:

* `AudioProcessor`: can modify audio
* `AudioRenderer`: only observes, doesn't modify

***

### Recommendations

If your goal is to:

* Generate audio and video data yourself and publish it: use a custom track
* Process local video the SDK has already captured: use `videoProcessors`
* Process captured or playback audio: use `audioCaptureProcessor` / `audioRenderProcessor`
* Read remote PCM for analysis or transcription: use `AudioRenderer`

Don't conflate these mechanisms. They solve four different kinds of problems.
