Making a Training Video: the Word Timeline, Cues, and Actions

A Backbuild Studio training video is a narrated walkthrough that drives your real product on screen. You write the narration, anchor each on-screen action to the word it belongs to, and Studio produces a finished video where the voice and the actions never drift apart. This guide teaches the whole flow: the word timeline you write in, the pause and delivery chips that shape the voice, the cues that pin things to words, the actions that drive the product, recording a screen that sits behind a login, framing a widescreen and a vertical cut from one project, and rendering with a credit estimate. After it you will be able to turn a script and a list of steps into a polished demo, and keep that demo current when the product changes without re-recording it by hand.

The Training Video editor. Callout 1 marks the video preview panel across the top, with aspect-ratio choices of 16:9, 9:16, 1:1, and 4:3 above it. Callout 2 marks the editor-style toggle that switches the layout between a Descript, script-first style and a Camtasia, video-first style. Callout 3 marks the right-hand tools panel with tabs for Actions, Speakers, Captions, Audio, Overlays, Motion, Chapters, and Intro and Outro. The media timeline sits below the preview. The account email address is masked in this image.
The Training Video editor: (1) the preview with its aspect choices, (2) the editor-style toggle, and (3) the tools panel. The word-timeline script sits below, shown next.

The Word Timeline

After this section you will know why the editor looks the way it does. A training video is written in a single surface called the word timeline. Your narration is laid out word by word, with space around each word, because every word is an anchor point. Audio, sound effects, actions, and overlays sit on lanes directly under the word they attach to, the way a caption sits under the syllable it belongs to. There is no separate write mode and arrange mode to switch between; you write the words and attach things to them in the same place.

Above the script is a video preview that shows what the finished frame looks like. You can expand or collapse it and drag it to the size you want, and it remembers your choice, so you can give the preview the whole top of the window while you check framing or shrink it to a strip while you write.

Write the Narration, and Let Auto-Markup Pace It

After this section you will have a narration that reads naturally without hand-tuning every pause. Type your script straight into the timeline. You do not have to mark anything up to get a good result: run auto-markup and Studio annotates the unmarked script for you, adding a pause after sentences and paragraphs and a delivery tag per paragraph, so the voice breathes and phrases the way a person would. Auto-markup fills the gaps; it does not overwrite pauses or tags you placed by hand, so you can let it do the bulk and then refine the moments that matter.

Shape the Voice with Pause and Delivery Chips

After this section you will control timing and tone precisely. Two kinds of chip live in the script, rendered as small anchored controls rather than raw text you have to remember the syntax for.

  • Pause chips insert a silence of a set length at that point in the narration. A stepper sets the duration, from a tenth of a second up to five seconds, so you can hold a beat before a key point or space out a list.
  • Delivery chips tag a single word with an emotion or delivery style, chosen from a picker, so a word or a phrase is read calm, warm, emphatic, or however the moment calls for. This is how you keep a long narration from sounding flat.

Because the chips are anchored controls, not literal characters in your text, they move with the words as you edit and never end up spoken aloud by mistake.

A close crop of the word-timeline script. The narration is laid out word by word. Callout 1 marks a delivery cue chip reading confidently, attached under the first word, which sets the tone of the speech that follows. Callout 2 marks a pause chip reading 0.4s, attached under a word, which holds a short pause there. More pause chips appear at the sentence boundaries further along the script.
The word timeline with markers under the words they anchor to: (1) a delivery cue that sets the tone, and (2) a pause chip that holds a beat. Tap a chip to change its length or tone.

Cues: Pin Anything to a Spoken Word

After this section you will attach on-screen events to the exact moment they should happen. A cue is a named marker on a word. You add one by clicking the word you want, and from then on anything you attach to that cue, an action, a sound effect, a music change, an overlay, happens at the instant that word is spoken. The important property is that cues survive edits: when you rewrite a sentence, Studio re-resolves the cue to its word rather than scattering your timing, so you can keep polishing the script without rebuilding the sequence of events every time.

A cue re-resolves onto its word when you edit the script. In the first line, Open the project and click New, a cue is anchored to the word click and points to an action, click the New button. In the second line the sentence is reworded to Now open your project, then choose New. The same cue re-resolves onto the reworded words and still points to the same unchanged action.
Cues anchor to words and re-resolve when you edit the script, so rewriting a line does not scatter your timing.

Actions: Drive Your Real Product

After this section you will know exactly what Studio can do on screen, and why the result is trustworthy. This is what sets a Studio training video apart from a screen recording or an avatar reading a script. When the video renders, a real browser is driven through the actions you listed, against your live Backbuild product or another site on your organization's Recording sites list (see the next section). The pixels the viewer sees are the real product, not a mock and not a static backdrop. The actions you can script include:

  • Navigate to a screen, and wait for an element or a state before continuing.
  • Click a control, type into a field, drag one element onto another, and select, hover, scroll, focus, or press a key.
  • Highlight the control in use with a rounded spotlight so the viewer can follow the pointer, and assert that something is true on screen.
  • Play a clip or an overlay at a chosen moment.

Because the render actually performs each step and can assert the expected result, a training video doubles as a check that the workflow still works: if a step the video depends on has broken, the render surfaces it rather than quietly recording a broken product. That is the answer to the recurring developer-relations question, can an AI video tool really open my product and click through it: yes, it drives the real interface, and the run is verifiable.

Does it just narrate over a screen recording, or does it actually use my product? It drives your real product. During the render a real browser navigates, clicks, types, and drags through the steps you scripted, against your live app or a site on your organization's Recording sites list, and records the genuine result. It is not an avatar over a static backdrop.

The recurring pain: how do I keep a demo current when the UI changes, without re-recording the whole thing? Because the video is generated by driving the live product, a change in the interface is simply re-captured the next time you render. You do not re-shoot by hand. And a timing or wording edit re-renders cheaply, because Studio only re-does the parts that changed. See Narration, Voices, and Cost for the caching that makes this true.

Choose the Site You Record: Recording Sites

After this section you will know which sites a training video can open, and how to add yours. A training video records one site, the one its production names. That site must be on your organization's Recording sites list. Backbuild's own sites are always on it, so a walkthrough of your Backbuild workspace needs no setup; to record your own product, an owner or admin of the organization adds its address first. The list belongs to the organization, so it is the same from every project's Studio. The editor does not have a field for the production's site yet; it is set over the REST API or by an AI agent (see AI Agents, the API, and Governance).

  1. Open the list. In Studio, choose Recording sites at the top of the home screen. Everyone in the organization can see the list.
  2. Add a site. An owner or admin types a hostname, such as app.example.com, or a wildcard, such as *.example.com, and chooses Add site. The confirmation names the site that can now be recorded.
  3. Remove a site with Remove on its row when it should no longer be recorded. The built-in Backbuild rule is marked Always allowed and cannot be removed.
Studio, Recording sites, with the heading's explanation that a screen recording can only open a site on this list and that Backbuild sites are always allowed. A green confirmation reads that *.example.org can now be recorded. Callout 1 marks the built-in *.backbuild.ai row, marked Always allowed. Callout 2 marks two added entries, *.example.org and app.example.com, each with a Remove button. Callout 3 marks the Hostname field, whose placeholder reads docs.example.com or *.example.com, above the Add site button.
Recording sites: (1) the built-in Backbuild rule, always allowed, (2) the sites your organization added, each removable, and (3) where an owner or admin adds a hostname or a wildcard.

The rules are deliberately strict, because a render opens the site in a real browser:

  • A hostname matches exactly that host. app.example.com allows only app.example.com.
  • A wildcard matches every host below it, not the name itself. *.example.com allows help.example.com and docs.app.example.com, but not example.com; add that separately if you record it.
  • Only public names are accepted. IP addresses, ports, paths, and private or reserved names (such as localhost or names ending in .local or .internal) are refused, and the screen shows the reason.
  • A list holds up to 100 entries, and every addition and removal is recorded in the audit log.

The list is checked whenever a production's site is saved, again when a render is started, and again when the render begins, so removing a site takes effect for every render that has not started yet. During a render the browser stays on the production's site: a step, a redirect, or a pop-up that tries to open another site ends the render instead of recording it. A sign-in that sends you to a different site, such as a separate identity provider, cannot be recorded for that reason. A site that is not on the list is refused with a message naming it and saying that an owner or admin can add it under Studio, Recording sites.

Can a training video record my own product, not just Backbuild? Yes. An owner or admin adds your product's hostname, or a wildcard for its subdomains, to Recording sites, and the production then records that site.

Why was my site refused? Either it is not on the list yet, or it is not a public hostname: IP addresses, ports, and private names cannot be added. A wildcard also does not cover its own apex, so *.example.com does not allow example.com.

The Tools Panel: Music, Effects, Overlays, and Shapes

After this section you will know where the production controls live. A dockable panel on the right holds the tools you reach for while building: add music, add a sound effect, add an action, add an overlay, and import a shape, plus Speakers, intro and outro management, the Secrets panel for linking a credential, and Metadata. Overlays cover titles, lower thirds, text, a caption band, images, shapes and shape groups, and a credits roll, positioned per output format and anchored to an absolute time, a spoken word, or an audio event, with transitions on the way in and out. You can import a scalable vector shape drawn in the document editor; it is cleaned and prepared for the video and can be saved to a shape library shared across the project or the organization.

Frame a Widescreen and a Vertical Cut Together

After this section you will produce a desktop and a phone version from one project. Studio frames the recording through a virtual camera. It outputs a widescreen 16:9 video for a desktop or a projector and, optionally, a vertical 9:16 video at 1080 by 1920 that follows the active part of the screen for a phone. Each format has its own focus keyframes, so the vertical cut zooms to the control in use rather than showing a shrunken whole screen, and a safe area keeps captions from colliding with the edges. You do not lay the video out twice: one project, two framings.

One source recording carries two camera frames. A wide 16:9 frame captures the whole screen for a desktop output. A tall 9:16 frame focuses on a smaller control area for a phone output, with a dashed safe area near the bottom where captions stay clear. Arrows lead to two outputs labelled 16:9 desktop and 9:16 phone.
One project frames both a widescreen and a vertical cut, each with its own focus and its own caption safe area.

Captions

After this section you will ship an accessible video. Because the narration is aligned word by word, Studio can produce captions that match the voice, configured for the production and available embedded in the video and as a sidecar file in the standard SubRip (.srt) and WebVTT (.vtt) formats that video hosts and players read. Captions are driven by the same alignment that keeps the actions in sync, so they are timed to the words rather than guessed.

Video Segments

After this section you will build a longer video out of parts. A production can include video segments: insert a new video project segment, a segment from your library, or a video-content segment, give it placeholder fields you fill in, and navigate into it to edit. Transitions smooth the seams where one segment meets the next, so a multi-part video reads as one piece.

Recording a Screen Behind a Login

After this section you will record a product that requires signing in, without exposing a credential. Many real walkthroughs need a signed-in screen. You link a credential from your Secrets vault to the production and mark the field it fills, for example a login email, a password, or a one-time code. During the render the value is entered into that field with the typing hidden and the field masked, so the credential never appears in the recording, and it is only ever used on the production's own site. The value itself is never shown in the editor, never placed in a link, and is resolved only at the moment the render needs it. This capability is rolling out; until it is enabled for your workspace, record walkthroughs of screens that do not require a sign-in during the recording.

Can it record an app that is behind a login? Yes. You link a credential from your Secrets vault and mark the field it fills. The value is typed in masked, never appears in the recording or in the editor, and is only ever used on the production's own site. It is resolved only at render time. The sign-in page has to be on that same site, since the render never leaves it.

Render and Its Cost

After this section you will render without a surprise bill. A render returns a credit estimate for the work it will do, and an unchanged re-render is free. The finished video is produced in every output format you chose. The composed cloud render is available on the paid plans and draws universal usage credits; if you bring your own text-to-speech key, that part of the work is billed by your provider rather than by Backbuild. Narration, Voices, and Cost explains the estimate and the caching in full.

One step is not in the editor yet. A new production starts as a draft, and the editor's Render audio, Render video, and Full render controls do not start a render for a draft: Render now reports that it could not start the render. Mark the production ready, set the site a training video records, and start the render over the REST API or from an AI agent, as AI Agents, the API, and Governance describes.

The render confirmation's estimate line, shown before you spend anything. Callout 1 circles the credit estimate reading about 38 credits, broken down into a text-to-speech portion and a speech-to-text portion, alongside a Cancel button and a Render now button.
The render confirmation's estimate line (1), split into its speech and alignment portions, as it reads when an estimate is available for the production.

If I change one word or one step, do I have to re-render the whole video? No. Studio caches the voice, the alignment, and the recording, and only re-does what changed. A cue or timing edit re-renders without paying to re-synthesize the voice, and a wording change re-synthesizes only the affected narration. An unchanged re-render costs nothing.

What caption formats do I get, and are they auto-generated? Captions are generated from the word alignment, so they match the voice exactly. They are available embedded in the video and as a separate sidecar file you can edit and export, in the standard SubRip (.srt) and WebVTT (.vtt) formats.

How much does a render cost, and can I see it before I spend? A render returns a credit estimate for its work, and an unchanged re-render is free. In-browser export elsewhere in Studio is free on every plan; the composed cloud render for a training video is on the paid plans and draws credits.

Can I package this as a SCORM or xAPI course for my learning-management system? Not today. Studio exports finished video files and caption tracks, which you host or embed wherever your learners watch; a SCORM or xAPI package for a learning-management system is not a current feature. If your program depends on completion tracking inside an existing system, plan to host the exported video there.

Can a non-video person build one of these? Yes. Write your narration, run auto-markup to pace it, and click a couple of words to add actions. You never have to touch a traditional timeline to build a narrated walkthrough.

Where to Go Next