Audio Post Production Workflow: Every Stage Explained (From Session Handoff to Final Delivery)

Audio post production is where raw recordings become a finished, broadcast-ready soundtrack. It is a sequential process with tight dependencies: each phase builds on the one before it, and errors made early compound downstream. Whether you are editing your first short film, producing a podcast at scale, or overseeing a broadcast delivery, understanding the full workflow is the foundation for doing any single part of it well.

Audio Post Production Workflow: Every Stage Explained (From Session Handoff to Final Delivery)

Audio Post Production Workflow: Every Stage Explained (From Session Handoff to Final Delivery)


What Is Audio Post Production? (Scope and Role)

Audio post production refers to all audio work that occurs after principal photography or primary recording is complete. It is distinct from production sound (recording on set), music production (composing and arranging a score), and live sound (reinforcing a live event). Its purpose is to take raw, unpolished recordings and shape them into a cohesive soundtrack that serves the picture and meets technical delivery standards.

What Is Audio Post Production? (Scope and Role)

What Is Audio Post Production? (Scope and Role)

The process involves multiple specialized roles depending on project scale:

  • Dialogue Editor: Cleans and assembles production dialogue tracks

  • ADR Supervisor / Editor: Manages re-recorded replacement dialogue

  • Sound Designer: Creates and integrates non-dialogue sound elements

  • Foley Artist / Foley Editor: Records and cuts synchronized, performance-based sounds

  • Re-Recording Mixer: Balances all elements in the final mix

  • Mastering Engineer: Prepares the final audio for delivery format and loudness compliance

On smaller productions, one or two people cover all of these roles. On feature films and broadcast series, each role belongs to a different specialist.


The Complete Audio Post Production Workflow at a Glance

Use this sequence as your reference map. Every phase is covered in detail in the sections that follow.

The Complete Audio Post Production Workflow at a Glance

The Complete Audio Post Production Workflow at a Glance

  1. Session Handoff and Project Setup — Receive the OMF/AAF, sync to picture, and organize the session before any editing begins.

  2. Dialogue Editing and Cleanup — Edit and restore production dialogue; schedule and cut ADR where needed.

  3. Sound Design and Foley — Create or source all non-dialogue sound elements; perform and record Foley.

  4. Music Editing and Score Integration — Conform licensed music and original score to the locked picture cut.

  5. Pre-Mix and Stem Building — Sub-mix dialogue, music, and effects into organized stems.

  6. The Final Mix — Balance all stems for intelligibility, emotional impact, and spatial coherence.

  7. Mastering and Loudness Compliance — Normalize to platform loudness specs; prepare deliverable formats.

  8. Deliverables and Quality Control — Export all required files, verify against technical specs, and QC before submission.


Phase 1 — Session Handoff and Project Setup

The handoff phase sets the conditions for everything that follows. A disorganized or premature handoff creates compounding problems through every subsequent phase. Getting this right is not glamorous work, but it is the highest-leverage investment in the entire workflow.

Phase 1 — Session Handoff and Project Setup

Phase 1 — Session Handoff and Project Setup

Picture lock is the mandatory precondition for audio post to begin. Picture lock means the video editor has made a final, approved cut with no further changes to timing, scene order, or content. Any edit made after audio work has started invalidates every timecode-synced element already placed, requiring time-consuming re-cutting across the entire session. Audio post should never begin before picture lock is confirmed in writing.

Once picture lock is confirmed, the audio editor receives session data from the NLE (Non-Linear Editor) as an OMF (Open Media Framework), AAF (Advanced Authoring Format), or XML file. This file contains the track layout, audio clips, and edit decisions from the picture cut. The audio editor imports this into a DAW such as Pro Tools or Reaper, loads the reference video with burned-in timecode, and verifies that audio and picture are synchronized frame-accurately.

Before any editing begins, invest time in session organization: name tracks clearly, sort them logically (dialogue tracks first, then room tone, then additional production audio), and build a consistent folder structure for all media. This investment pays back every hour of editing that follows.

A spotting session typically happens at this stage. The dialogue editor, sound designer, and director review the locked cut together to identify where ADR is needed, where sound design elements should be placed, and what the overall sonic direction should be. Notes from the spotting session become the working brief for every subsequent phase.

What to Check When You Receive the OMF or AAF

Before accepting the session and beginning any work, run through this checklist:

  • File integrity: Confirm the AAF or OMF opens without errors in your DAW

  • Missing media: Resolve all missing or offline audio files before proceeding

  • Sample rate consistency: Verify all audio matches the session sample rate (typically 48 kHz for video work)

  • Embedded vs. referenced audio: Know whether audio is embedded within the OMF or referenced from external files, and confirm all referenced files are accessible in the expected locations

  • Timecode and frame rate: Confirm the session frame rate matches the reference video exactly (23.976, 24, 25, and 29.97 fps are the most common formats)

Do not proceed if any of these checks fail. Return the file to the picture editor for correction before touching any audio.


Phase 2 — Dialogue Editing and Cleanup

For narrative work, dialogue editing is the most labor-intensive phase in the workflow. Its goal is to deliver clean, continuous, intelligible dialogue tracks to the mix stage, drawing entirely from the best available production recordings.

Phase 2 — Dialogue Editing and Cleanup

Phase 2 — Dialogue Editing and Cleanup

Continuity editing involves selecting the best take for each line or section of dialogue, then matching mouth sounds, breath patterns, and room ambience between cuts so the edit is inaudible. A technically clean but rhythmically unnatural edit is as problematic as a noisy one.

Noise reduction uses tools like iZotope RX, the industry standard for dialogue restoration. Common treatments include dialogue isolation, spectral de-noise, de-click, de-reverb, and surgical spectral repair for specific intrusions such as car pass-bys, phone rings, or boom equipment strikes.

Room tone fills address gaps, cuts, and unusable sections that cannot be replaced with ADR. Room tone (ambient background noise recorded on set at the end of a production day) is used to fill these gaps and maintain a consistent acoustic space between cuts. Without room tone, the difference between a cut and a filled gap becomes audible as an unnatural silence.

Breath and handling management includes deciding which breaths should be retained for naturalism and which should be attenuated for clarity. Handling noise from microphone movement, clothing rustle, or boom pole contact is addressed here, not deferred to the mix stage.

When to Use ADR vs. When to Fix It in Post

ADR (Automated Dialogue Replacement) means re-recording an actor’s performance in a controlled studio environment and cutting it back into the session in place of unusable production audio. The decision to schedule ADR involves weighing several factors:

Use ADR when: - Production audio is technically unrecoverable due to severe clipping, structural noise, or pervasive background noise that masks speech - A performance adjustment is required for story or emotional tone - A line must change in post for creative, legal, or clearance reasons

Fix it in post when: - Intelligibility is adequate and the performance is strong - The noise floor can be reduced without audible processing artifacts - The intrusion is periodic or spectrally isolated and treatable with RX

image

The key test is intelligibility and emotional integrity. If you cannot clearly understand the line, or if the actor’s performance is buried under restoration processing, ADR is the correct choice even when it adds schedule and cost.

How Clean Production Audio Reduces Post Workload

The quality of field recordings has a direct and measurable effect on dialogue editing time. Audio arriving with clipping, high noise floors, or inconsistent levels requires multi-pass restoration that adds hours to the edit phase and can degrade audio quality regardless of how capable the tools are.

32-bit float recording eliminates one of the most damaging problems: digital clipping. Because 32-bit float files store audio with far greater dynamic range than 24-bit recordings, gain can be recovered in post even when input levels were set incorrectly on location. The Hollyland LARK MAX 2 features 32-bit Float Internal Recording, which means backup recordings from this wireless system are clipping-proof at the point of capture. A dialogue editor receiving 32-bit float files does not need to spend session time reconstructing clipped transients or working around overloaded recordings. The headroom is already preserved.


Phase 3 — Sound Design and Foley

Sound design and Foley together build the acoustic world of the project beyond what production microphones captured. Both phases are guided by the spotting notes developed during the handoff phase.

Phase 3 — Sound Design and Foley

Phase 3 — Sound Design and Foley

Sound design refers broadly to the creation and integration of all non-dialogue audio elements:

  • Hard effects (HFX): Specific, sync-dependent sounds tied to on-screen actions (a door slam, a gunshot, a keyboard click)

  • Designed effects: Processed, synthetic, or heavily layered sounds created for narrative or stylistic purposes (sci-fi environments, creature voices, surreal audio transitions)

  • Ambience and background beds: Continuous atmospheric textures that define the acoustic space of a scene (city noise, forest, office hum, crowd walla)

Sources for sound design include licensed sound libraries (Boom Library, Soundly, SoundSnap, the Hollywood Edge), original field recordings, and synthesized audio built entirely within the DAW.

Foley is sound performed live by a Foley artist on a specialized recording stage. Unlike effects pulled from a library, Foley is recorded in sync with picture to capture the specific texture and timing of movement: footsteps on different surfaces, clothing movement, props handled by actors, and the physical contact that small on-screen actions require. Foley provides intimacy, realism, and sub-perceptual texture that library effects rarely match in precision.

Common Foley categories include: - Footsteps: Matched to actor, surface type, and pacing - Cloth and body movement: Settles, sits, and clothing texture - Props: Physical handling of scene objects - Specialty: Contact walla, fight sounds, and specific object interactions

Sound design and Foley assets are delivered as edited, trimmed clips organized by category and ready for the pre-mix session.


Phase 4 — Music Editing and Score Integration

Music enters the post workflow in two forms: licensed music (pre-existing tracks cleared for use) and original score (composed specifically for the project). The music editor is responsible for conforming both to the locked cut.

Phase 4 — Music Editing and Score Integration

Phase 4 — Music Editing and Score Integration

During early picture editing, a temp score is typically used to guide editorial and director decisions. This placeholder soundtrack is assembled from existing recordings. Once the picture is locked and a composer is engaged, the temp is replaced with original music. This process, called the temp-to-final score swap, involves cutting final cues to sync precisely to picture, trimming or extending music to match scene lengths, and managing crossfades between cues.

Music supervisor handoffs involve receiving cleared music files alongside licensing paperwork and integrating them into the session at approved durations and timecodes. Even minor deviations from picture timing can cause sync drift that becomes audible in the mix. All cue in-points and out-points should be verified against the locked cut frame-by-frame.

Cue sheets documenting title, composer, publisher, duration, and usage type (background vs. featured) are generated at this stage and passed to distribution for licensing compliance. Common sync issues to watch for: tempo-related drift over several minutes, sample rate mismatches between the composer’s deliveries and the session, and gain staging inconsistencies between cues from different sources.


Phase 5 — Pre-Mix and Stem Building

Before the final mix, all audio elements are organized and sub-mixed into stems: separated sub-mixes grouped by element type. Stems are not a technical formality. They are the structural backbone of a professional mix and a critical deliverable for international distribution.

The three primary stems are:

Stem Name

Tracks Included

Output Use

Dialogue (DX)

Production dialogue, ADR, walla

Primary intelligibility; M&E reference

Music (MX)

Score cues, licensed music, stingers

Separate music control; music-only deliverable

Effects (EFX)

Hard FX, designed FX, Foley, ambience

FX-only deliverable; M&E construction

Beyond the primary three, some projects add sub-stems: a separate Foley stem, a backgrounds stem, or a production sound effects stem for granular revision flexibility.

image

Why stems matter for M&E: An M&E (Music and Effects) track is a stereo or surround mix containing only music and effects, with no dialogue. M&E deliverables are required for international sales because they allow foreign distributors to lay in dubbed dialogue in any language while retaining the full sound design and score. Without properly organized stems, building a clean M&E is either impossible or requires rebuilding the entire session.

During the pre-mix, each stem receives dedicated gain rides, group processing, and level automation so it arrives at the final mix stage as a clean, controlled sub-mix. This approach gives the re-recording mixer manageable groups rather than hundreds of individual tracks to balance simultaneously.


Phase 6 — The Final Mix

The final mix is where all creative and technical elements converge. A re-recording mixer takes the pre-mixed stems and shapes them into the complete, finished soundtrack: balanced, spatially coherent, and emotionally effective in relation to the picture.

Phase 6 — The Final Mix

Phase 6 — The Final Mix

Dialogue intelligibility is the non-negotiable priority in virtually every project type. The mix should never bury a line of dialogue under music or effects without deliberate creative intent. The first pass of any mix session typically establishes dialogue levels, with music and effects balanced relative to speech.

The re-recording mixer draws on several categories of processing:

  • EQ (equalization): Sculpts frequency balance within each stem and between competing elements, such as rolling low-end from the dialogue stem to create space for bass-heavy effects

  • Dynamics processing: Compression, limiting, and expansion to control dynamic range within stems and manage level consistency between scenes

  • Reverb and spatial processing: Places elements in an acoustic space that matches the visual environment

  • Automation: Fine-grained level riding to maintain intelligibility across varying scene conditions

Gain staging throughout the session determines available headroom at every stage: track input, bus, and master output. Consistent gain staging prevents clipping and ensures that when a limiter engages at the mastering stage, it is working on a clean signal rather than correcting accumulated level mismanagement.

The final mix is also tailored to playback format:

  • Theatrical mix: Optimized for large speaker arrays in a calibrated cinema environment, typically 5.1, 7.1, or Dolby Atmos

  • Broadcast mix: Conforms to ATSC A/85 or EBU R128 loudness specifications; designed for television playback in uncontrolled listening environments

  • Streaming mix: Aligned to platform-specific loudness targets; typically –14 LUFS integrated for most video streaming platforms

On large productions, separate mix versions are delivered for each format, all built from the same stem session.

Key Tools for the Final Mix

The final mix relies on several tool categories regardless of DAW platform:

  • EQ: Frequency sculpting across all stems (FabFilter Pro-Q 3, Waves API 560, stock DAW channel EQ)

  • Dynamics (compressor / limiter): Gain control, density, and transient management (Waves SSL G-Master, UAD 1176, Vertigo VSC-2)

  • Reverb and spatial processing: Room placement and acoustic matching (Altiverb, Lexicon 480L emulations, Exponential Audio Nimbus)

  • Loudness metering: Real-time integrated and short-term LUFS monitoring with True Peak display (iZotope Insight 2, Nugen Audio VisLM, Waves WLM Plus)

  • DAW platform: Pro Tools is the industry standard for post production mixing due to its session compatibility, editing precision, and long-standing role in delivery pipelines. Reaper is widely used in indie film and podcast workflows. Logic Pro suits Mac-based productions at smaller scale. Adobe Audition is common in broadcast and video-adjacent environments.


Phase 7 — Mastering and Loudness Compliance

Audio mastering in a post production context serves a different primary purpose than music mastering. Where music mastering focuses on tonality, competitive loudness, and format optimization for streaming, post production mastering is primarily concerned with loudness normalization and technical compliance with the delivery platform’s specifications.

Most streaming platforms and broadcast distributors enforce loudness standards that incoming audio must meet before acceptance. Understanding these targets is not optional for anyone involved in final delivery.

Key loudness concepts:

  • Integrated LUFS: The average loudness across the full duration of the program. This is the primary delivery target.

  • Short-term loudness: Average loudness over a rolling 3-second window, used to monitor dense passages.

  • Momentary loudness: Average over 400 milliseconds, representing the loudest instant in the program.

  • True Peak (dBTP): The highest reconstructed peak including inter-sample peaks. Most platforms set a ceiling of –1 dBTP to prevent distortion during codec encoding.

image

Reference table for common delivery standards:

Platform / Standard

Integrated Target

True Peak Ceiling

Spotify

–14 LUFS

–1 dBTP

YouTube

–14 LUFS

–1 dBTP

Apple Podcasts

–16 LUFS

–1 dBTP

Netflix

–27 LUFS (dialogue-gated)

–2 dBTP

EBU R128 (European broadcast)

–23 LUFS ±1 LU

–1 dBTP

ATSC A/85 (US broadcast)

–24 LKFS ±2 LU

–2 dBTP

Theatrical (Dolby)

85 dB SPL reference / –20 dBFS

Format-dependent

Deliverable format requirements typically specify WAV or Broadcast WAV (.bwav) for broadcast and theatrical, at 48 kHz or 96 kHz and 24-bit depth. Streaming platforms often accept AAC or MP3 for consumer-facing files, but a mastered WAV should always be archived as the production master regardless of delivery format.

Pro Tip: Meeting the True Peak ceiling and meeting the integrated LUFS target are two separate requirements. A file can pass True Peak limiting and still fail loudness compliance if the integrated measurement is off-target. Always meter with a dedicated loudness analyzer, not just a peak meter, before exporting a deliverable.


Phase 8 — Deliverables and Quality Control

The deliverables phase packages the project for handoff. A complete professional deliverable package includes substantially more than a single stereo mix file.

Phase 8 — Deliverables and Quality Control

Phase 8 — Deliverables and Quality Control

Standard deliverable types:

  1. Stereo mix — The primary full-mix stereo file with all elements combined

  2. 5.1 or Atmos mix — Surround or immersive mix where the format is required

  3. Dialogue stem (DX) — Isolated dialogue sub-mix

  4. Music stem (MX) — Isolated music sub-mix

  5. Effects stem (EFX) — Isolated effects sub-mix

  6. M&E (Music and Effects) — Combined music and effects with no dialogue, for international distribution

  7. Textless version — Full mix excluding dialogue over title cards or credits, for localization

  8. Conform report — Documentation of any editorial decisions or changes made during post

Not every project requires all of these. A podcast episode needs a stereo mix and a mastered file. A feature film bound for international distribution needs the full package.

QC (Quality Control) is the final gate before any file leaves the post facility or is submitted to a distributor. QC is not optional; it is the professional standard, and skipping it is the most common cause of distributor rejections, costly re-deliveries, and damaged professional relationships.

QC checklist before delivery:

  • Loudness verification: Confirm integrated LUFS, True Peak, and short-term measurements against the platform specification

  • Sync check: Spot-check audio against reference video at the head, middle, and tail of the program

  • Drop-out scan: Listen through at playback speed for unexpected silences, artifacts, or glitches

  • Clipping audit: Confirm no True Peak exceedances in any deliverable file

  • Format validation: Verify sample rate, bit depth, channel count, and file format against the technical spec sheet

  • Naming convention audit: Confirm every file follows the distributor’s naming requirements exactly (incorrect file names are a surprisingly common rejection reason)

After QC is passed, files are delivered. After delivery is confirmed, archive the full session with all media at original resolution in an organized folder structure. Post production sessions are frequently reopened for versioning, format adaptation, or M&E construction months or years after original delivery.


Audio Post Production Workflow by Project Type

The full eight-phase workflow applies to narrative film and broadcast television. For other project types, the pipeline compresses or shifts based on scope and resources.

Phase

Film / TV

Podcast

Corporate / Online Video

Session Handoff

Full OMF/AAF import; spotting session required

Simple file import; no picture lock required

Video export from NLE; basic track organization

Dialogue Editing

Intensive — multiple takes, ADR, room tone

Edit for clarity, pacing, filler removal

Basic cleanup; noise reduction where needed

Sound Design / Foley

Full — all categories

Not applicable

Minimal; spot FX or one ambient bed

Music Integration

Score and licensed cue editing to picture

Theme and ad break placement only

Licensed background track placement

Pre-Mix / Stems

Full stem structure required

Not typically built

Not typically built

Final Mix

Re-recording mix with format variants

Level balance, ride automation, mono check

Stereo balance; dialogue-forward mix

Mastering

Broadcast and theatrical loudness specs

–16 to –14 LUFS for platform delivery

–14 LUFS for web; platform-specific

Deliverables / QC

Full package: mix, stems, M&E, textless

Stereo MP3 or WAV; show notes sync

Stereo WAV or MP3 to platform spec

Even at the simplest end of the spectrum, the sequence logic holds: get the edit right before the mix, get the mix right before mastering, and verify everything before delivery.


Common Mistakes at Each Stage (and How to Avoid Them)

Phase

Common Mistake

Prevention

Session Handoff

Beginning audio work before picture lock

Confirm picture lock in writing; never accept an AAF without a signed-off locked cut

Session Handoff

Accepting a session with unresolved missing media

Run the five-point acceptance checklist before beginning any editing

Dialogue Editing

Skipping room tone collection

Request room tone from the production sound mixer; if unavailable, reconstruct from gaps in production audio

ADR

Over-scheduling ADR instead of using noise reduction

Evaluate each line individually using intelligibility as the test; exhaust iZotope RX options before committing to ADR sessions

Sound Design

Placing effects without spotting notes

Spot the session with the director before beginning; design to confirmed intention, not assumption

Pre-Mix

Skipping stems and going directly to the final mix

Always build stems; the M&E deliverable requirement alone justifies the time investment

Final Mix

Poor gain staging that causes output clipping

Set consistent gain structure from track to bus to master; meter at every stage of the signal path

Mastering

Confusing True Peak ceiling compliance with integrated LUFS compliance

A file can pass True Peak limiting and still fail loudness spec; always use a dedicated loudness analyzer

Deliverables

Submitting files without QC

Treat QC as mandatory; a single listen-through and a loudness meter check takes less time than a re-delivery


FAQ

What software is used in audio post production?

Pro Tools is the industry standard for professional film, television, and commercial post production due to its session format compatibility and mixing capabilities. Reaper is widely used in indie film and podcast workflows as an affordable and flexible alternative. Adobe Audition is common in broadcast and video-adjacent teams. Logic Pro suits Mac-based smaller productions. For dialogue restoration, iZotope RX is the de facto standard tool across all platforms and project types.

What is the difference between audio mixing and mastering in post production?

Mixing is the process of balancing all audio elements within the session, shaping EQ, dynamics, and spatial placement to create a cohesive soundtrack. Mastering is a separate pass applied to the completed mix file to conform it to platform loudness targets, True Peak ceilings, and format specifications. On larger productions, these are distinct sessions handled by different engineers, and the outputs serve different functions in the delivery package.

How long does audio post production take?

Timelines scale with project scope. A feature film typically requires 6 to 12 weeks of audio post. A short film may take 1 to 2 weeks from handoff to delivery. A podcast episode generally involves 1 to 4 hours of dialogue editing plus 30 to 60 minutes of mixing and mastering. Key variables include the quality of production audio, sound design density, number of deliverable formats required, and the volume of revision passes requested.

What is picture lock, and why does audio post depend on it?

Picture lock is the point at which the video edit is finalized with no further changes to timing, structure, or content. Audio post must begin after picture lock because all timecode-synced work, including ADR cuts, sound design placements, and music cues, is calibrated to exact frame positions. Any picture edit made after audio work has started invalidates that sync, forcing recutting across every affected track in the session.

What are stems in audio post production?

Stems are organized sub-mixes grouped by element type: dialogue (DX), music (MX), and effects (EFX). They allow the re-recording mixer to adjust and balance element groups without touching individual tracks, and they provide the source material for M&E deliverables required in international distribution. Without stems, adapting a mix for versioning, format changes, or dubbed language tracks requires rebuilding the session from its source tracks.


The Workflow as a Connected System

Audio post production is a chain of dependent phases. The quality of work in each stage directly shapes the difficulty and outcome of the next. Production audio arriving well-organized and properly captured compresses the dialogue editing workload. A clean dialogue edit makes the pre-mix faster. Well-built stems make the final mix manageable. An accurately metered master makes delivery smooth.

The most expensive mistakes in audio post almost always trace to shortcuts taken upstream: starting before picture lock, skipping file acceptance checks, ignoring room tone, or building no stems before the final mix. Those shortcuts ripple forward and consistently cost more time to correct than they saved.

Use this guide as your workflow map. For deeper coverage of individual phases, explore the linked articles below on dialogue editing, ADR, mixing for broadcast, and loudness compliance standards.