Skip to main content

Overview

When the built-in microphone, camera, and screen-sharing capture don’t meet your needs, the Swift SDK provides two extension paths:
  • Custom audio and video tracks: your app provides the audio or video frames itself
  • Media processors: process the audio and video frames the SDK has already captured
Common scenarios include:
  • Publishing video frames generated by an external rendering engine
  • Sending local PCM audio, BGM, or TTS data into the channel
  • Beauty filters, virtual background, watermarks, or AI video processing
  • Voice changing, noise suppression, or enhancement of captured audio

Custom video tracks

Create a custom video track:
Then your app keeps calling pushFrame(...) to feed in video frames:
The key points:
  • No startCapture() needed
  • The SDK doesn’t generate frames for you
  • You must keep producing CVPixelBuffer values yourself
Suitable for:
  • Publishing video from Metal / SceneKit / Unity / game engines
  • Custom composited video
  • AI-generated frames or the output of image processing

Custom audio tracks

Create a custom audio track:
Then keep injecting PCM data:
This kind of track also doesn’t need startCapture(), because the audio source isn’t the SDK’s internal capturer but the AVAudioPCMBuffer you’ve prepared yourself. Suitable for:
  • Playing background music and publishing it to the channel
  • Sending PCM data decoded from local files into RTC
  • Output from speech synthesis, voice cloning, or external DSP engines

Minimal flow summary

Whether it’s custom audio or custom video, the overall model is the same:
In other words, the SDK handles “publishing and transport”, and you handle “producing the media data”.

Video processor chain

If what you need isn’t “producing new video frames yourself” but “processing existing captured frames”, videoProcessors is the better fit:
Processors run serially in array order:
  • Return a new VideoFrame to pass it on
  • Return nil to drop the frame
This is better suited for:
  • Beauty filters
  • Virtual background
  • Video watermarks
  • AI segmentation or visual enhancement

Audio processors

Audio processors are attached to the SRTCEngine main entry point, not to an individual track:
Note:
  • audioCaptureProcessor processes audio after capture
  • audioRenderProcessor processes remote audio before playback
  • This is a global capability, not a per-track one
So if what you need is “processing only one particular microphone”, first confirm whether this global model matches your expectations.

If you only want to observe remote audio data

Then don’t use AudioProcessor; use AudioRenderer instead:
The difference between the two is clear:
  • AudioProcessor: can modify audio
  • AudioRenderer: only observes, doesn’t modify

Recommendations

If your goal is to:
  • Generate audio and video data yourself and publish it: use a custom track
  • Process local video the SDK has already captured: use videoProcessors
  • Process captured or playback audio: use audioCaptureProcessor / audioRenderProcessor
  • Read remote PCM for analysis or transcription: use AudioRenderer
Don’t conflate these mechanisms. They solve four different kinds of problems.