ADR Studio Suite

User Manual
Document version 1.0 — ADR Studio Suite (Light / Pro) · macOS and Windows
Contents

Chapter 1 — Introduction

What is ADR Studio Suite

ADR Studio Suite is a professional application for managing ADR (Automated Dialogue Replacement — post-production looping/dubbing work). It covers the entire workflow: video import, generating dialogue lines (cues) via automatic transcription or subtitle import, text management and translation, multi-take recording with teleprompter, comping of the best takes, and finally export to whichever format the recipient needs — a CSV for editing, a Word script for the actor, or a complete DAW session for the mixing engineer.

The application runs as a native desktop app on macOS and Windows (Electron), with a single timeline shared across every phase of the work: the cues created during transcription are the same ones used for recording and export, with nothing to re-copy by hand between programs.

Light and Pro editions

ADR Studio Suite is available in two license tiers, aimed at different roles in the pipeline:

AreaLightPro
Video, SRT, EDL, CSV importYesYes
Automatic AI transcriptionYesYes
Cue management (text, actor, timecode)YesYes
Assisted translationYesYes
CSV / SRT / EDL / Word script exportYesYes
Multi-take recording (arm/solo, teleprompter, punch in/out)NoYes
Comping lanes, rating, split takeNoYes
Mixer with original-audio referenceNoYes
AI voice isolation (Demucs) and voice detection (VAD)NoYes
AI spatial tagging (framing/position in scene)NoYes
Full DAW session export (Pro Tools, Logic, Nuendo, Reaper, Fairlight, Pyramix, Audition, AAF, FCPXML)NoYes
Advanced exports (Advanced CSV, Advanced EDL, Excel Report, Multi-Stem ZIP)NoYes

In short: Light covers the entire script preparation and translation work through to text delivery (CSV/Word). Pro adds everything needed inside the recording booth: microphone, takes, comping, and audio delivery ready for the mixing engineer.

System requirements

Chapter 2 — Installation and activation

Installing on macOS

Installing on Windows

First launch: the free trial

On first launch, if no license is present, the app automatically starts in a 30-day free trial mode with every Pro feature unlocked. The remaining-days countdown is visible under Settings → License.

Note: once the 30 days run out without an activated license, the app locks with a full-screen block until a valid license is imported.

Activating a license

Whoever provides ADR Studio Suite (or your organization's purchasing department) will send you a license file with a .lic extension, generated specifically for the machine you're installing on.

License states

StateIconMeaning
License Active🟢A valid license has been imported and verified. Tier (Light/Pro) as purchased.
Free trial🟡No license imported: countdown against the 30-day trial, all Pro features active.
Trial expired🔴The 30-day trial period ended with no license. App locks on launch.
License expired🔴A valid license was present but its expiry date has passed. App locks on launch.
License invalid🔴The .lic file fails verification (signature mismatch, wrong Machine ID, corrupted/tampered file). App locks on launch.

The license signature is cryptographically verified (Ed25519) on every launch: a hand-edited .lic file — for instance, one with the expiry date pushed back — is recognized as invalid and rejected outright, not silently ignored.

Chapter 3 — Interface overview

The main window is divided into four areas, all visible together: top bar, left sidebar, video monitor, and center timeline.

Top bar

Organized into sections, top to bottom:

In the Light edition, the Recording, DAW Export and Export Professional sections aren't shown: the interface only displays what that license can actually use.

Video monitor

Shows the current frame of the imported video, with the dialogue text overlaid when relevant (readable by the actor during the take), plus a mirror of the external-monitor window for the control room, if active.

Timeline and cue table

Below the video monitor, in order: the timeline zoom/toolbar row, the color-coded block timeline (one block per cue, color-coded by actor), the comping lanes (Pro only, when a cue has more than one recorded take), and finally the text/editable cue table, with columns for number, selection, play, start, end, actor, spatial position (Pro), and dialogue text.

In the table, each cue's text color indicates its sync status: green if the text fits the take's duration, red if it's likely too long, blue if it's shorter than needed — an automatic estimate based on average reading speed.

Chapter 4 — Quick guide: the workflow in sequence

This chapter walks through the typical path of a project from start to delivery, step by step. Each step points to the chapter that covers it in depth.

Step 1 — Create or open a project

From the sidebar, New Project to start from scratch, or Load Project to resume saved work. The project is saved to a file that includes the cues, settings, and references to the recorded takes.

Step 2 — Import the video

Sidebar → Import → Video. The app automatically reads the file's duration and frame rate and uses them as the reference for the whole timeline. See Chapter 5.

Step 3 — Generate the cues

Three possible routes, which can also be combined:

See Chapter 6 for details on automatic transcription.

Step 4 — Refine the cues

Text, actor name, and timecode for every cue remain editable at any time, directly in the table, in both license tiers. This is where the dialogue writer and translator step in to fix what was auto-generated, align actor names, and resolve any overlaps. See Chapter 5.

Step 5 — Record the takes (Pro)

Arm the cue you want to record (the Arm button in the table or on the timeline), set pre-roll/post-roll if needed, and press Record. The teleprompter shows the text to the actor during the take. Every repetition creates a new take, numbered in sequence. See Chapter 8.

Step 6 — Comp the takes (Pro)

When a cue has multiple takes, the comping lanes let you listen to them, assign a rating, mark the best one (Best Take), and cut/join portions from different takes into a single composite performance, without editing outside the app. See Chapter 9.

Step 7 — Export

Depending on the recipient: CSV or Word Script for whoever works on the text, or a complete DAW session (with the best takes' audio stems already positioned on the timeline) for the mixing engineer. See Chapter 13.

Chapter 5 — Cue management

The cue is ADR Studio's basic unit of work: it represents a single line of dialogue, with a start timecode, an end timecode, an assigned actor, and the text to be dubbed. The entire application — transcription, translation, recording, export — revolves around cues.

Creating a cue manually

Editing text, actor, and timecode

In the cue table, the Text cell is directly editable: a click places the cursor, and you type as normal. The Actor cell is an autocomplete field: as you start typing, it suggests names already used in the project, and automatically assigns a consistent color to each actor (the same color recurs in the timeline, the comping lanes, and exports).

The Start and End cells are edited by clicking on them: a timecode editor opens in HH:MM:SS:FF format, based on the project's FPS.

Search and filters

The search field above the table (Search by actor name...) filters the cues to show only those for the searched actor — useful in productions with many characters, to focus on one voice actor at a time during a recording session.

Multi-selection and merging (Join)

Selecting multiple cues via the checkboxes in the table's first column (or Select All Cues from the Cue menu) lets you merge them into a single line with Join Selected Cues (Join) — useful when automatic transcription has split a single sentence into multiple segments.

Deletion

The ✕ at the end of a row deletes that single cue; the ✕ in the table header (Delete All Cues) clears the entire project — with a confirmation prompt, since this is a destructive action.

Sync check: text color

ColorMeaning
GreenThe text fits within the assigned take's duration
RedThe text is likely too long for the available time
BlueThe text is shorter than necessary for the available time

The estimate is based on an average reading speed per language and is meant as a quick heads-up, not an exact measurement: it's normal for some cues to stay "red" by stylistic choice (e.g. rushed lines) — the mixing engineer and dialogue writer always have the final word.

Undo and Redo

Every change to text, actor, or timecode is recorded in the project history: Ctrl+Z (Cmd+Z on Mac) undoes, Ctrl+Y redoes, for up to 50 steps back.

Chapter 6 — Automatic AI transcription

ADR Studio Suite includes a local transcription engine (based on WhisperX) that listens to the imported video's audio and automatically generates cues: it detects where each line starts and ends, transcribes the text, and estimates the speaker.

Starting it

"Studio Local" engine

Besides standard transcription, a "Studio Transcription" mode is available (menu Transcription → Studio Transcription), aimed at longer sessions or higher quality requirements, with a dedicated local engine and the option to abort partway through if needed (Abort Studio Transcription).

Automatic cleanup

The generated text goes through post-processing that removes typical automatic-transcription "hallucinations" (anomalous repetitions, text generated during silence) before being turned into cues — but a human re-read is still recommended before moving on.

When you don't need automatic transcription

If you already have an SRT file or an ADR CSV prepared elsewhere, you can import it directly (Sidebar → Import) instead of starting from transcription — useful when the script arrives already prepared by another department.

Chapter 7 — Assisted translation

The Text Translation section in the sidebar lets you batch-translate every cue in the project from a source language to a target one.

Note: automatic translation is a starting point to speed up the dialogue writer/adapter's work — it doesn't replace the lip-sync and rhythm adaptation that dubbing requires.

Chapter 8 — Recording (Pro only)

The recording module turns ADR Studio into a genuine ADR booth workstation: per-character arm/solo, configurable pre-roll and post-roll, teleprompter, punch in/out, markers, and level monitoring.

Arming a cue

The Arm button in the table (or on the timeline) readies a cue for recording. Shift+Click on Arm arms an entire range of consecutive cues, handy for recording a sequence of lines without stopping to arm each one individually.

Solo

The Solo button isolates a cue for listening/playback, muting the others — convenient for focusing on a single line during rehearsal.

Pre-roll and post-roll

Pre-roll is the amount of time (in seconds) played before the cue starts, giving the actor time to settle into the scene's rhythm; post-roll is the time after it ends. These are set per cue or as project-wide defaults in Settings.

Fade in / fade out

Every recording can have a short fade at the start and end (in milliseconds), useful to avoid audio clicks at the beginning/end of a take.

Teleprompter

Displays the current cue's text full-screen (or on the external monitor) during the take, with adjustable font size in Settings — designed to stay readable even on a second screen positioned away from the actor.

Punch in/out

Lets you re-record only a portion of an existing take, without repeating the entire line from scratch.

Markers

Pressing M during playback drops a marker on the timeline, visible in the sidebar's Markers section — useful for flagging a spot to revisit without stopping work.

External monitor

Tools → External Monitor opens a dedicated window to move onto a second screen in the control room or booth, showing the video and overlaid dialogue text, independent of the main window.

Continuous recording

Besides the per-cue mode, a continuous recording mode is available that follows the video without stopping at every line, for sessions with experienced actors who prefer a more natural flow.

Chapter 9 — Take management and comping (Pro only)

Every time an already-armed cue is recorded, a new take is created, numbered in sequence (T1, T2, T3...). When a cue has multiple takes, comping comes into play: choosing, joining, or trimming the best parts of each.

Comping lanes

Below the main timeline, the comping lanes show every take for an actor side by side, with direct drag/trim/selection on the waveform: drag to select the portion you want from one take, then move to another lane for the next line.

Rating and colors

Every take can receive a rating (stars) and an identifying color, so the best takes stand out at a glance without having to re-listen to all of them every time.

Best Take

One take per cue can be marked as the Best Take: that's the one the app uses by default for exports (audio stems, DAW session) unless a different choice is made manually during comping.

Split

A take that's too long, or with a mistake partway through, can be split in two (Split) at the exact chosen point, so the final version can combine the good part of one take with another.

Take Manager (dedicated window)

Sidebar → Recording → Take Manager opens a window with the full list of takes per cue, to rename, delete, re-listen, and approve — including all cues in the project at once (Approve Takes for All Cues).

Take Report

Sidebar → Recording → Take Report generates a text summary of every take recorded in the project: how many per cue, which one was approved — useful as a session log to keep or share with production.

Chapter 10 — Mixer (Pro only)

The mixer controls three independent volumes during listening and recording: the original video's audio, playback of already-recorded takes, and the microphone input.

Chapter 11 — Voice isolation and voice detection (Pro only)

Voice isolation (Demucs)

Separates the voice from the score/background noise in the video's original audio, using a local AI model (Demucs). Useful for a cleaner, "dialogue-only" waveform during sync work, and for comparing the original line's rhythm against the new recording more easily.

VAD — Voice Activity Detection

Automatically detects the segments where speech is present in the audio, as a cross-check against the transcription's own segmentation: if speech recognition got a line's alignment wrong, VAD helps pinpoint where the speech actually starts and ends.

Detection-sensitivity presets are available (Settings → AI section), adjustable based on the audio type (clean dialogue, noisy scene, presence of music).

Chapter 12 — Spatial Awareness AI (Pro only)

The Spatial Awareness module automatically classifies the framing and position of the speaking character for each cue, by analyzing the video frame at that line's timecode — not just the text, since "on/off screen, back to camera, wide shot" are visual facts that the transcribed text alone can't provide.

The 10 categories

CodeMeaning
FCOff Screen
ICOn Screen
DSBack to Camera
SDLying Down
ACCCrouched
MOVIn Motion
PPClose-Up
CLWide Shot
CLLExtreme Wide Shot
SMSilent/Mute

The assigned tag appears in the Pos. column of the cue table, and can be edited by hand at any time if the automatic classification isn't correct.

How it works under the hood

Classification runs on a local vision+language model (Qwen2.5-VL, GGUF format, run through llama-cpp-python), completely offline: no frame ever leaves the computer. Since it's a 10-category classification task with constrained output, not free text generation, it performs well on CPU alone, with no dedicated graphics card required.

Regenerating missing tags

Menu Transcription → Retry Missing Spatial Tags reprocesses only the cues that don't yet have a tag assigned (for example, after manually adding new cues), without having to rerun classification on the whole project.

Requirements

This feature requires both model files (language + vision projector) to be present in the application's Models folder. If they're missing or incompatible, classification will error out — see Chapter 16, "Spatial Awareness isn't working."

Chapter 13 — Exporting

ADR Studio Suite exports directly into whatever format the recipient uses, avoiding manual intermediate steps and possible sync loss.

Exports available in the Light edition

FormatWhat it's for
ADR CSVSpreadsheet-format cue list, for importing into other software or tabular review
SRTStandard subtitle file, compatible with most video editors
Word ScriptA print-ready .docx document, to hand to the actor or dubbing director
EDLEdit Decision List, for editing

Exports available in the Pro edition only

Export to DAW

Generates a session ready to open directly in the chosen DAW, with one track per character and the audio files for each take (normally the Best Take) already positioned at the correct timecodes:

DAW / formatNotes
Pro Tools (AAF)An .aaf session importable via File > Import > Session Data
LogicNative Logic session
NuendoSession with AAF support
Reaper.rpp project, one track per character
FairlightSession with AAF and FCPXML support
PyramixDedicated session
Adobe AuditionDedicated session
FCPXMLFor editing, with absolute media references

Export Professional

Running an export

Chapter 14 — Settings

The Settings panel (Cmd/Ctrl+, or Tools → Settings...) gathers every project- and application-level option.

Project

Audio

Language

Interface language (Italian/English) and default source/target languages for assisted translation.

AI (Pro features only)

The Reset AI Settings button restores this section to its defaults — useful if a custom configuration has stopped working as expected.

Timeline

License

Current status, Machine ID, and the import button — see Chapter 2 for details.

Chapter 15 — Keyboard shortcuts

Key(s)Action
Click on the timelineMove the playhead (seek) to that point
Shift + ClickCreate a new cue at the clicked point
MDrop a marker at the current timecode
Shift + LToggle loop
RStart/stop recording
SpacebarPlay / Pause
J / K / LShuttle: rewind / pause / forward (variable speed on hold)
Ctrl/Cmd + ZUndo
Ctrl/Cmd + YRedo
Ctrl/Cmd + NNew Project
Ctrl/Cmd + OOpen Project
Ctrl/Cmd + SSave Project
Ctrl/Cmd + Shift + EExport to DAW...
Ctrl/Cmd + ,Open Settings
Shift + Click on ArmArm a range of consecutive cues

Chapter 16 — Troubleshooting

This chapter collects the most common issues encountered in day-to-day use and during installation, along with their solutions.

License issues

Full-screen "Trial expired" / "License expired" / "License invalid"

The app locks intentionally in these three states — by design, not by error: it's confirmation that a valid .lic file is needed. Check that you've imported the latest file you received (Settings → License) and that the Machine ID shown matches the one given to whoever generated the license. If the file was modified even slightly (for instance with a text editor), its signature is no longer valid and a new file must be requested.

Audio issues

The microphone doesn't show up among the available devices

Automatic transcription doesn't start or gets stuck

Spatial Awareness isn't working

Error: "llama-cpp-python not available or without Qwen2.5-VL support"

This means the internal Python environment used for local AI is either missing the required module or can't find the model files. It's a known issue especially on manually compiled or custom builds:

Export issues

The app won't launch / crashes on startup

Notes for those building the app from source (developers/IT)

The notes below concern whoever builds the executable from source code, not day-to-day use of the already-installed app.

Appendix A — Light vs Pro summary

FeatureLightPro
Video / SRT / EDL / CSV importYesYes
Automatic AI transcriptionYesYes
Cue editing (text, actor, timecode)YesYes
Assisted translationYesYes
CSV / SRT / EDL / Word Script exportYesYes
Recording (arm/solo, teleprompter, punch in/out)Yes
Comping, rating, split takeYes
Mixer with original-audio referenceYes
Voice isolation (Demucs) and VADYes
Spatial Awareness AIYes
DAW session export (Pro Tools, Logic, Nuendo, Reaper, Fairlight, Pyramix, Audition)Yes
Export Professional (Multi-Stem ZIP, Advanced CSV/EDL, Excel Report)Yes

Appendix B — Glossary

TermMeaning
ADRAutomated Dialogue Replacement — recording dialogue in sync with picture, in a booth, to replace or supplement the original production audio
CueA single line of dialogue, with start/end timecode, actor, and text
TakeA single recording/repetition of a cue
CompingAssembling a cue's final version by choosing/joining the best parts from multiple takes
Best TakeThe take marked as best for a cue, used by default in exports
ArmReadying a cue for recording
Pre-roll / Post-rollListening time before and after the cue respectively, to give rhythmic context
Punch in/outRe-recording only a portion of an existing take
TeleprompterOn-screen display of the line's text during the take
VADVoice Activity Detection — automatic detection of speech segments in an audio signal
StemAn isolated audio file for a single source (e.g. one character), ready for mixing
EDLEdit Decision List — a standard list of editing changes, portable across software
AAFAdvanced Authoring Format — a session exchange format between professional software (e.g. Pro Tools)
TimecodeA time reference in HH:MM:SS:FF format (hours:minutes:seconds:frames)
FPSFrames per second of the video, the basis for every timecode calculation in the project