Overall model
The Swift SDK’s object model can be summarized as:SRTCEngine: the main SDK entry point, responsible for joining channels and creating local tracksChannel: an actual channel connection, responsible for publishing, subscribing, the user list, and event dispatchTrack: the abstraction of an audio or video streamChannelDelegate/TrackDelegate: entry points for event callbacksSRTCVideoView/VideoView/SRTCVideoRenderer: the video rendering layer
- Connection state
- User state
- Media track state
SRTCEngine and Channel
SRTCEngine
SRTCEngine works more like a “factory + entry point for joining”. You typically use it for two things:
joinChannel(token:options:)createLocalMicTrack(...)/createLocalCameraTrack(...)/createLocalScreenTrack(...)
Channel
After you join a channel, Channel is what actually holds the channel’s state:
- Publish local tracks
- Subscribe to remote tracks
- Get channel info and user info
- Listen for users joining and leaving, track changes, disconnects and reconnects, and custom messages
joinChannel can be called multiple times—one engine can join multiple channels at the same time, and each channel’s publishing, subscriptions, users, and events
are independent of each other; srtc.channels is the list of currently live channels. Tracks belong to the engine, not to a channel: the same capture track
can be published to multiple channels, capture happens only once, and cleanup is handled by the last releaser. Publishing audio to multiple channels at the same time has one
hard constraint (the set of audio sources must be the same in every channel)—see Multi-channel (Chinese).
Track system
Local tracks
Remote tracks
Audio mix mode
A key implementation detail of the current Swift SDK: local audio is sent in mix mode by default. This means:- The microphone, custom audio, and screen audio first go into the internal
AudioMixer - What the PeerConnection actually sends is a single internal
audio_mixaudio track - Your code can still control creation, muting, and stopping capture of each audio source separately
- It reduces the complexity of publishing multiple audio tracks concurrently
- Custom audio, the microphone, and screen audio share one sending model
- It stays consistent with the existing rtc-js architecture
Video capture and publishing
A video track’s lifecycle usually has three steps:- You can preview locally first, then decide whether to publish
- You can control hardware capture and network sending independently
- When a permission failure or device switch failure occurs, the state is easier to pinpoint
Video rendering
The SDK provides three common rendering entry points:SRTCVideoView
A SwiftUI component, best for rendering directly in a view:
VideoView
For UIKit / AppKit; use bind(track:) to manage binding.
SRTCVideoRenderer
A lower-level rendering view, bound by the track itself calling addRenderer(...) / removeRenderer(...).
Device management
DeviceManager.shared handles device enumeration and device change monitoring:
- Enumerate cameras:
cameras() - Enumerate microphones / speakers on macOS:
microphones(),speakers() - Enumerate audio routes on iOS:
audioRoutes() - Monitor device hot-plugging and audio session interruptions:
DeviceManagerDelegate
DeviceManager rather than writing your own platform-specific handling.
Event model
Channel-level events go throughChannelDelegate:
- Join succeeded
- Reconnect / disconnect
- User joined, left, or updated
- Remote track added, updated, or removed
- Custom messages
TrackDelegate:
- Track info changes
- Mute / unmute
- Capture ended
- Underlying WebRTC track binding completed
Further reading
- Mute vs. unpublish (Chinese)
- Device management (Chinese)
- Screen sharing (Chinese)
- Multi-channel (Chinese)
- Custom tracks (Chinese)