Overview
When the built-in microphone, camera, and screen-sharing capture don’t meet your needs, the Swift SDK provides two extension paths:- Custom audio and video tracks: your app provides the audio or video frames itself
- Media processors: process the audio and video frames the SDK has already captured
- Publishing video frames generated by an external rendering engine
- Sending local PCM audio, BGM, or TTS data into the channel
- Beauty filters, virtual background, watermarks, or AI video processing
- Voice changing, noise suppression, or enhancement of captured audio
Custom video tracks
Create a custom video track:pushFrame(...) to feed in video frames:
- No
startCapture()needed - The SDK doesn’t generate frames for you
- You must keep producing
CVPixelBuffervalues yourself
- Publishing video from Metal / SceneKit / Unity / game engines
- Custom composited video
- AI-generated frames or the output of image processing
Custom audio tracks
Create a custom audio track:startCapture(), because the audio source isn’t the SDK’s internal capturer but the AVAudioPCMBuffer you’ve prepared yourself.
Suitable for:
- Playing background music and publishing it to the channel
- Sending PCM data decoded from local files into RTC
- Output from speech synthesis, voice cloning, or external DSP engines
Minimal flow summary
Whether it’s custom audio or custom video, the overall model is the same:Video processor chain
If what you need isn’t “producing new video frames yourself” but “processing existing captured frames”,videoProcessors is the better fit:
- Return a new
VideoFrameto pass it on - Return
nilto drop the frame
- Beauty filters
- Virtual background
- Video watermarks
- AI segmentation or visual enhancement
Audio processors
Audio processors are attached to theSRTCEngine main entry point, not to an individual track:
audioCaptureProcessorprocesses audio after captureaudioRenderProcessorprocesses remote audio before playback- This is a global capability, not a per-track one
If you only want to observe remote audio data
Then don’t useAudioProcessor; use AudioRenderer instead:
AudioProcessor: can modify audioAudioRenderer: only observes, doesn’t modify
Recommendations
If your goal is to:- Generate audio and video data yourself and publish it: use a custom track
- Process local video the SDK has already captured: use
videoProcessors - Process captured or playback audio: use
audioCaptureProcessor/audioRenderProcessor - Read remote PCM for analysis or transcription: use
AudioRenderer