Dataset release

Audio/Video Engineering Agentic Tasks 1M

Media production agents / 1,029,459 operations / MIT

A focused 1.03M-operation dataset for training agents inside DAW and NLE environments, built around mid-session troubleshooting, dense conversational instructions, timeline reasoning, routing changes, sonic repair, color matching, and edit execution.

AI media engineer operating audio and video production systems
1,029,459audio/video task operations
156Mprompt-only task tokens
459.8Mgeneration tokens spent
25professional archetypes
2time-based software environments
127.75average words per instruction

Technical overview

What this release is built to train.

The dataset captures the messy, high-pressure language of professional audio engineers, composers, colorists, video editors, sound designers, and post-production operators mid-session.

It focuses on conversational actuation: translating context-heavy human problem reports into exact software actions, parameter changes, routing edits, timeline cuts, and node-level fixes.

DAW coverage includes mixing, mastering, recording, restoration, synthesis, Foley, game audio, broadcast audio, live sound, composition, vocals, and beat-making.

NLE and cinematic post-production coverage includes video editing, color grading, motion graphics, documentary workflows, podcast production, creator workflows, and music education media.

The mid-session crisis as a training objective

The 1,029,459 operations model the way media professionals actually speak while a session is underway: context-heavy, diagnostic, time-constrained, and often emotionally charged. Instead of requesting a generic effect, a prompt can describe a low-end routing conflict, drifting multicam sync, shot-matching failure, or unstable synthesis patch and require exact parameter, timeline, routing, node, or automation changes.

The target behavior is conversational actuation. A multimodal agent must separate symptoms from causes, preserve creative intent, infer dependencies across a timeline or signal chain, and translate the request into precise GUI operations without requiring the user to decompose every step.

Twenty-five roles across DAW and NLE environments

DAW coverage includes production, recording, mixing, mastering, sound design, live sound, restoration, game audio, Foley, broadcast, orchestral and screen composition, commercial composition, electronic production, songwriting, vocal performance, and beat making. These tasks require psychoacoustic, spectral, dynamic, routing, spatial, and arrangement reasoning.

NLE and post-production coverage includes editing, color grading, motion graphics, documentary work, podcast production, creator workflows, and music education. These tasks exercise temporal reasoning, media organization, sync, pacing, editorial rhythm, color-science changes, and delivery continuity across many dependent frames and layers.

Generation architecture, scale, and use

Curated role taxonomies replace generic software terms with node-level concepts such as dynamic spectrum analyzers, multicam sync bins, and binaural panners. More than 1,000 workers sampled deterministic tool intersections tied to batch IDs, while hash-based behavior injection varied the voice from clinical diagnosis to high-pressure troubleshooting without giving up reproducibility.

The prompts contain approximately 156 million task tokens, average 127.75 words, and required approximately 459.8 million generation tokens. Batch-ordered, Zstandard-compressed Parquet records preserve the professional and group labels. The MIT-licensed release targets DAW/NLE agents, timeline-aware multimodal models, troubleshooting copilots, environment simulations, and execution-policy research.