Skip to main content
This page covers how cloud recording and live streaming work as a whole. For the parameters and response structure of each endpoint, see Cloud recording and live streaming.
This page covers recording the whole channel: the entire channel is mixed into one stream. The output is not necessarily a single file—long recordings are split into segments by duration; see “One recording produces multiple files” below. If what you need is audio split into tracks by speaker and into segments by utterance, that is a different set of endpoints: see Per-speaker voice recording. The differences between the two are in the comparison table below.

One task, four capabilities

task_type is a bit mask. Each of the four independent capabilities takes one bit; add the values of the ones you want at the same time: So 3 = video recording + stream mixing, 9 = video recording + live stream, and 15 = all four.
Don’t treat 3 as a separate “hybrid mode” type. Earlier docs listed the values as “1 recording, 2 mixing, 3 hybrid”, which was wrong—it left out audio recording and live stream, and misrepresented how the values combine.

One recording produces multiple files

Recording is not “one recording, one file”. Two situations split it into multiple files:
  • The recording is longer than the segment limit (1 hour by default): when transcoding after the task ends, it is split into rolling segments by duration, with continuous time between segments
  • The underlying recording stops automatically partway through because there has been no audio or video for 30 seconds, and is then restarted: a new segment starts, with a time gap between segments
So the data model has two levels: one recording = one task (task_id), and the task has N recording files (record_id): Each file carries these fields, which are enough to build complete playback:
  • seq starts at 1; sorting by it gives the playback order
  • offset_ms is this segment’s offset (in milliseconds) from the start of the whole recording; use it to build the progress bar for continuous playback
  • began_at / ended_at are wall-clock times, for aligning with the timeline of your own business
  • reason=2 means there is a time gap between this segment and the previous one (recording was interrupted and then restarted); continuous playback jumps there
Recording files start transcoding and uploading only after the task ends. There is a wait between stopping the recording and being able to play it, which is longer for long recordings. To tell when a recording is playable, rely on the mcu_record callback (is_last is true on the last one), not on the task status changing to “ended”. See the Callback events guide.

Audio recording is not voice recording

Audio recording (task_type=4) on this page is MCU mixed audio recording: the entire channel is mixed into one audio stream, one file per task. If what you need is audio split into tracks by speaker and into segments by utterance (push-to-talk records, per-person timing, sentence-by-sentence transcription), use Per-speaker voice recording, which is a different set of endpoints: You can run both at the same time; they don’t affect each other.

Key differences between video recording and live streaming

This difference determines the call sequence: a live stream “can be distributed as soon as it starts”, while a recording “has output only after it ends”.

Choosing a layout

layout_data.layout determines the video layout. auto picks a grid automatically based on the number of online users, which works for most cases; specify a layout only when you need a fixed one: Note that N in grids_N is not continuous (there is no 7, 10, or 11). Passing an unsupported value returns an error. If you don’t pass layout_data, the app’s default recording configuration is used—put a common watermark, labels, and layout policy there so you don’t have to pass them every time you start a task.

Pinning users to specific cells

layout_data.div_list pins specific users to specific cells. Cells without an assignment are filled automatically in the order users joined.
  • cells[].idx is the cell index, in the same order as <td> cells in an HTML table (left to right, top to bottom)
  • Leaving uids empty means “the remaining online users rotate through these cells” (global rotation); listing several means those users rotate through these cells (group rotation)
  • polling_dur is the rotation interval in seconds; 0 means no rotation
  • When cells[].bind_share is true, the cell is bound to the channel’s screen sharing stream first
A typical use is “main speaker pinned to the large cell, everyone else rotating through the small cells”: put the main speaker’s uid in the large cell’s cells, and leave uids empty in the small cells’ cells with polling_dur set.

Behavior when no one is in the channel

layout_data.nobody_text determines what happens when no one is in the channel:
  • Empty → recording pauses (recommended, to avoid long stretches of black video)
  • Text set → recording continues and shows the text

Watermark and user name labels

  • watermark.type: 0 default, 1 none, 2 single row, 3 multiple rows. If watermark.text is empty, the task’s title is used as the watermark
  • The position of user name labels is written as a combination of letters: L left, R right, T top, B bottom, which can be combined (LB = bottom left). Empty means no labels

Full sequence

When a channel is destroyed, in-progress tasks stop automatically; you don’t need to call stop first.

Billing note

Video recording, audio recording, and live streaming are server-side capabilities that incur ongoing costs. If you forget to stop a task after starting it, it keeps running until the channel is destroyed—so explicitly call stop in your end-of-call flow rather than relying only on channel destruction.