Overview
When you need extra audio or video processing between capture and publishing (such as AI noise suppression, beauty filters, or voice changing), the SDK provides a unified processor mechanism. From first principles, a processor is essentially “take oneMediaStreamTrack in, output one processed MediaStreamTrack”. The SDK only needs to provide an attachment point on local tracks, while plugins implement the processing logic; the two are decoupled through the TrackProcessor interface.
Processors are distributed as separate npm packages, such as the RNN noise suppression plugin @seastart/srtc-plugin-rnnoise.
Attaching and detaching
Local audio/video tracks (LocalMicTrack, LocalCameraTrack, custom tracks, etc.) all provide two methods:
replaceTrack without renegotiation, and remote users notice nothing.
Chaining multiple processors
setProcessor accepts an array and chains multiple processors in order (source → P1 → P2 → … → publish). A common combination: noise suppression first, then voice changing.
ProcessorPipeline to combine the processors into one; audio chains share the same AudioContext, reducing the cost of multi-stage processing. You can also use ProcessorPipeline directly:
Writing a custom processor
Implement theTrackProcessor interface to plug in:
ProcessorOptions fields:
When writing an audio processor, reuseoptions.audioContextif it exists (don’tcloseit yourself); only close it indestroyif you created the context yourself.
Lifecycle and notes
- You must
startCaptureto have a track before callingsetProcessor; otherwise it throws. setProcessoris idempotent: calling it again detaches the existing processor before attaching the new one.- Automatic reattachment: after switching devices (
changeDeviceId) or callingstartCaptureagain, the SDK automatically rebuilds the processor with the new source track, so you don’t need to callsetProcessoragain. - Processors mostly rely on
AudioContext/WASM, need HTTPS (or localhost), and may need a user gesture before they can run.