Skip to main content
Your backend receives requests from your client, makes business decisions, and then calls the SRTC server API with app_key. This page is a reference for building your own backend, which calls the SRTC server API.

Why this layer is required

The SRTC server API has the highest authority across the board. Every request must be HMAC-signed with app_key, and app_key must never appear on the client —if it leaks, others can use your app’s identity to join any channel and remove any user. So only your backend can do the following: The client only gets a token to join the channel and sends and receives streams; everything else goes through your backend.

One endpoint per capability

Usually each of your business endpoints maps to one capability, and your backend code decides which SRTC server API to call:
Each endpoint should validate its parameters and check permissions.

Where to check audio and video permissions

For turning the camera or microphone on or off, sharing, and so on, wrap them as backend “endpoints”—/open-video, /open-audio, /start-share. They only check permissions and don’t need to call any SRTC server API. In effect, before your client turns on the camera, it first asks your backend “is this user allowed to turn on video right now?” (whether the call has started, whether the host has disabled their video, whether the concurrency limit has been exceeded). If not, the backend returns an error code, and the client shows the user a message accordingly.

props conventions

SRTC channels and users each carry a props extension field. You define its content; the SRTC server only stores and broadcasts it. For example: Channel props—shared state for the whole channel
User props—state of a single user
Change props with Update user info or Update channel info. The server broadcasts a change event to everyone in the channel, and clients refresh their UI accordingly. This is the easiest way to sync state such as “who is sharing”, without building a messaging path of your own.

Custom message conventions

When you need to send a one-time notification (rather than persistent state), use Send a custom message. You define the message body structure entirely; you can follow a two-part action + content format:
Use props for state, custom messages for actions: props describe “what things look like now”, and users who join later can read them too; custom messages describe “what just happened”, and only users online at the time receive them.

IM messages vs. in-channel custom messages

Both can “send a message to someone”, so they’re easy to confuse. The dividing line is whether the recipient is in the channel. In short: use custom messages for things inside the channel, and IM for bringing people into the channel.

When only IM works

The key is that the other party hasn’t joined the channel yet, so channel messages can’t reach them at all. These are the two most typical scenarios.

Calling: bringing someone into the channel

  1. The caller asks your backend to start a call. The backend generates a channel name and notifies the callee over IM:
  1. The callee’s device rings when it receives this. Answering, declining, or being busy each sends an IM back to the caller:
  1. After answering, each side gets a token and joins the channel. From then on, interactions (mute, sharing, chat) should switch to in-channel custom messages—the users are already in the channel, so there’s no need to go through IM.
Call timeouts and caller cancellation work the same way, each with its own message (call_timeout / call_cancel); you choose the command names.

Having a terminal join the channel silently

A controlled terminal has no interactive UI. Your backend sends a command to have it join the channel and publish on its own:
When the terminal receives monitor_start, it gets a token, joins the specified channel, and publishes its camera stream, with no user action needed. It isn’t in any channel before it receives this message, so IM is the only way.

Others

Cross-channel notifications (messages to people who aren’t in this channel), system announcements, and so on—any scenario where “the recipient isn’t in this channel” belongs to IM.

Three capabilities unique to IM

  • Delivery by device: one uid may be online on several devices at once (phone + PC). rsids delivers to a specific device, while ruids sends to all of that user’s online devices. Channel messages don’t have this dimension
  • Query online devices: Get users’ online devices, so you can check whether the other party is online before calling
  • Force a device offline: Force an IM device offline, used when a later login replaces an earlier one
Devices coming online and going offline also trigger im_connect / im_disconnect callbacks to your backend, which you can use to track online status.
For messages that must not be lost, such as calls and invitations, remember to set important to true—they are resent after a reconnect. Ordinary state sync doesn’t need it; the cost is slightly higher latency. Both endpoints have this field.

Tracks

Each track a client publishes carries a desc description, which the server and other clients use to tell what it’s for: When subscribing, pick the one you need by desc—for example, a 3×3 monitoring grid subscribes only to camera_big.