Virtual background and beauty filters act on the same shared camera capture pipeline, so settings apply to all channel instances at once.When both are on, the order is fixed: beauty filter first, then virtual background. Person segmentation takes the beautified image as input, so its edges match the final video.
Step 1: Install the virtual background component
We recommend installing it before you need virtual background, for example when the call screen opens. Passnil for modelPath to use the SDK’s built-in person segmentation model.
Step 2: Set the background effect
Background blur and background replacement are mutually exclusive; the later call wins. Both APIs are also remembered if called before installing, and take effect automatically once installation completes, so you don’t need to worry about their order relative toinstallVirtualBackground:.
Step 3: Turn virtual background on or off
After installing, it’s off by default and must be turned on explicitly. When off, it’s a zero-overhead pass-through and runs no inference.RTCEngineErrorConflict. Turning it off clears the inter-frame state, so the next time it’s turned on it converges again from the first frame and doesn’t flash a stale mask.
Step 4: Keep frame rate on low-end devices (optional)
By default, person segmentation runs on every frame. On low-end devices you can increase the inference interval and reuse masks to gain frame rate; then turn on mask sync as needed to eliminate trailing artifacts.When
inferenceInterval is 1, turning setVirtualBackgroundMaskSync: on or off makes no difference—it only takes effect after you increase the inference interval.Performance cost
Starting with3.1.1, person segmentation always runs CPU inference and no longer uses CoreML. The model is only 256×256, and its operators can’t be fully handled by CoreML, so the cost of shuttling data between the CPU and CoreML on every frame exceeds the compute saved—in testing, it’s a net slowdown.
Measured on iPhone XS Max / iOS 18.7.9 (720p@25fps capture, average over 300 frames, inferenceInterval at the default 1):
On
3.1.0, 53 ms per frame already exceeds the 40 ms budget for 25 fps and drags the encoder into dropping frames; on 3.1.1, running segmentation on every frame by default is enough for 25 fps on this device, so you usually don’t need to increase inferenceInterval. For lower-end devices, we still recommend testing as in the previous section before deciding.