Music edits for tracks
A Claude agent that edits batches of short music edits to a track excerpt: footage, effects and captions
Task
The system makes batches of short music edits to a 12–20 s excerpt of a track, for social media. All of them stay within shared frames of what looks good, yet differ in footage, finish and caption font — and none repeats past edits on other tracks. A Claude agent edits to the track’s score from a footage catalogue, Remotion renders, a check enforces the rules. What qualifies is decided by blind arenas.
How it works
- Track → score: bars, hits, drop, mood
- Lyrics: vocal stem → Whisper → caption screens
- The system assigns each edit its footage, format, finish and caption style
- Claude agent writes the edit to the score
- Render — Remotion; check — rules and uniqueness across all past edits
- Blind arena: only approved styles reach a batch
What had to be solved
The edits came out as copies
Left to choose on their own, the agents converged on one look: in an arena of six edits, four took the same caption style. Now the batch generator assigns each edit its own set — caption style and finish direction — and the check holds it. Uniqueness is measured against every past edit on every track, not only within the batch.
Commits
- f6b8a3403.10.2026the batch generator assigns each edit its own set and diversity sheet: the copies in w2 came from leaving the choice to the agents
- ce9c8d503.10.2026only approved styles and uniqueness across the whole history: every edit must differ from all past ones, not just its neighbours in the batch
- 31950de02.10.2026ears: the agent is deaf — it gets the grid, hits, drop and mood as a score; the grid is checked against attacks instead of by ear
- 5f266d006.10.2026first batch run the way the interface launches it: $2.06 for two edits — editing and finish by one agent cost less than the old finish alone
- d463d5806.10.2026captions: word timing by the sound, screens on sixteenth notes and no shorter than 0.3 s