Back to Blog

Agent workflows

Your Video Agent Made It Louder. Did It Check the Final Mix?

Give video agents an audio delivery target, selected stream and final-file check. Separate loudness normalization from a mix that actually makes the narration understandable.

OpenAgentSkillPublished:

The narration sounds quiet, so a user asks a video agent to make it louder. The agent raises the level and produces another file. Music now feels aggressive, but the voice is still difficult to follow.

The missing question is what “louder” was supposed to fix. A delivery-level target, the balance between speech and music, and the choice of audio track are separate decisions. An export can satisfy one while leaving the others wrong.

For product demos, tutorials and narrated clips, specify an audio handoff that identifies the intended mix and the final file to inspect. Do not accept a successful render as evidence that the audience will hear the right thing.

Methodology: inspect the audio release boundary

We reviewed EBU R 128, FFmpeg's loudnorm filter documentation and the ffprobe reference on September 24, 2026. No audio was rendered, measured or auditioned for this article. The review sequence below is a proposed workflow, not a comparison of tested video skills.

The selection criterion was a recurring handoff problem: a numeric export adjustment is treated as approval of the whole listening experience. This guide is for creators and developers who automate video assembly but still need an accountable release check.

Ask which delivery requirement applies

EBU R 128 recommends average programme loudness of -23 LUFS and uses Loudness Range and Maximum True Peak Level as descriptors. That recommendation provides a concrete reference, not evidence that every social platform requires the same target. EBU R 128 publication.

Our recommendation is to obtain the intended destination's current delivery specification or an explicit production brief. Record the target, units, allowed tolerance, peak constraint and channel requirements that actually apply. If the user has no such specification, label the target as a production choice rather than a platform rule.

Do not copy a number from an unrelated broadcast, podcast or social workflow and call it universal. When producing multiple deliveries, name the variants and preserve their separate requirements.

Identify the track before measuring it

ffprobe can report media-stream information and select streams for inspection. Its structured output helps make the inspected artifact and stream explicit. ffprobe documentation.

Our proposed intake record includes the file identity, audio-stream index, available language and disposition information, channel layout and the reason that stream was selected. Do not assume the first audio stream is the final narration mix merely because it is first.

If the project contains voice, music and effects as separate inputs, identify which combined output the listener should receive. A measurement of a voice stem is not a measurement of the completed soundtrack.

Metadata is an inventory aid, not a listening verdict. A stream can be correctly labeled while containing an outdated take or the wrong balance. Treat missing or conflicting metadata as a reason to inspect, not a reason to guess.

Normalize the intended mix, not an intermediate guess

FFmpeg documents loudnorm as supporting single- and double-pass operation with dynamic and linear modes. Linear normalization requires measured input values and conditions on loudness range and true peak; when those conditions are not met, it can revert to dynamic mode. FFmpeg loudnorm reference.

The practical implication for our proposed workflow is to retain the reported processing mode rather than equating a requested option with an observed result.

First settle the editorial mix: which voice take, which music segments, and which intended effects. Then measure that defined mix, apply the chosen normalization process, and record the settings and resulting report. If the soundtrack changes, do not reuse measurements from an earlier mix.

This ordering keeps the review understandable. Otherwise a second agent may inherit a measurement file without knowing which source combination it described.

Reopen and remeasure the release artifact

Our recommendation is to inspect the encoded delivery file separately from the project configuration. Keep the pre-export report, but do not use it as the sole release evidence.

The proposed sequence is:

  1. Identify the exact output file and intended audio stream.
  2. Confirm the expected stream and channel arrangement.
  3. Measure that final stream against the agreed target and tolerance.
  4. Listen to representative sections and the transitions between them.
  5. Record unresolved numeric or perceptual issues before distribution.

Check a quiet explanation, a dense music section, an abrupt transition and the ending. Choose these sections from the actual editorial structure rather than sampling an arbitrary five seconds.

If only selected sections were auditioned, say so. Do not label a sample review as a complete listen. A short deliverable may justify listening to the whole file; a long one needs an explicit review scope.

A good number cannot explain every bad mix

Consider a fictional tutorial whose music obscures a critical instruction. Increasing the level of the combined soundtrack does not tell the agent whether the voice-to-music relationship matches the user's intention.

Our recommended response is to revisit the mix decision, not repeatedly chase a louder target. Ask which words must remain clear, where music is allowed to dominate, and whether the source recording itself needs attention. Make any requested repair a separate, reviewable change.

For a candidate skill, prepare a non-sensitive fixture containing narration, a music transition and a quiet closing line. Define the expected experience before generation: the intended voice remains understandable, the music transition is deliberate, and the ending is intact. Pair that listening checklist with the destination's numeric requirements.

This is a proposed acceptance exercise, not a claim that a particular skill passed. Keep listening observations and measurement results in separate fields so neither is mistaken for the other.

Limitations and a useful receipt

This workflow does not guarantee identical playback on every device or destination, repair every recording defect, or replace professional mixing judgment. A target selected for one use may be inappropriate for another.

A compact release receipt should identify the source mix, output artifact, selected stream, delivery specification, processing report, final measurements, listening scope and open issues. If a check was unavailable, report it as unavailable instead of converting an assumption into a pass.

Our video timing guide handles the separate question of cue timing and frame-rate changes. Browse the skills directory for production candidates, but ask them to deliver this audio evidence alongside the video. “Made louder” describes an action. It does not describe an accepted result.

Your Video Agent Made It Louder. Did It Check the Final Mix? | OpenAgentSkill