OpenMontage + Antigravity Changed My Editing Game (It's Free)
A start-from-scratch walkthrough of Open Montage, the 36,000-star repo that turns a coding agent into a video editor: how its 1,000+ files reduce to five parts, the exact Windows install path (Python, Node, FFmpeg, venv, Remotion render engine, Piper local voice) done inside Antigravity for about zero dollars, and a live side-by-side of a vague prompt producing AI slop versus a detailed prompt that returns a timestamped implementation plan.
AI with Surya12 minTranscript found
Quick learning frame
Read this before watching.
Creative automation uses agents to accelerate production while keeping human taste in story, pacing, selection, and critique.
New playlist item from AI with Surya; queued for transcript-backed review, topic mapping, and a practical learning artifact.
Skill you build: The ability to install and drive a bring-your-own-model agent toolkit like Open Montage, and to prompt it with enough specificity (catalog first, plan first, cadence named) that it produces controlled motion graphics instead of generic AI slop.
Watch for the shift from claim to mechanism. The learning value is the point where the transcript reveals a repeatable action, tool boundary, context move, review habit, or artifact.
Concept diagram
Where this video fits.
01Brief
02Source
03Generation
04Selection
05Edit
06Taste Review
Deep lesson
Turn this video into working knowledge.
2,306 cleaned transcript words reviewed across 648 timed caption segments.
Thesis
OpenMontage + Antigravity Changed My Editing Game (It's Free) teaches a practical creative automation move: A start-from-scratch walkthrough of Open Montage, the 36,000-star repo that turns a coding agent into a video editor: how its 1,000+ files reduce to five parts, the exact Windows install path (Python, Node, FFmpeg, venv, Remotion render engine, Piper local voice) done inside Antigravity for about zero dollars, and a live side-by-side of a vague prompt producing AI slop versus a detailed prompt that returns a timestamped implementation plan.
The goal is not to remember the video. The goal is to extract the operating principle, tie it to timestamped evidence, test how far the claim transfers, and make something reusable.
1:50
Five parts, no brain
“And five, and this is a very important part, there is no brain in the box. No large language model inside the repo. So, your coding agent reads the manuals, follows the recipes, uses a tool, but then...”
Open Montage's 1,000+ files are really five things: markdown instruction files that teach an agent every video-production job, 12 pipelines (explainer, talking head, documentary, step-by-step) as recipes, about 50 small Python scripts as the hands, render engines that write real code per scene (fake terminals, Apple-style glass widgets), and critically no LLM inside the repo at all: your coding agent reads the manuals and you supply the model. Open the cloned repo and map each of its top-level folders onto one of the five categories, then read one pipeline file end to end so you know which recipe your first edit should use.
6:12
Zero-dollar local stack
“locally on this machine, minus the cost to the large language model API calls, okay? So, the next piece would be for us to record the clip live, okay? I'm going to record a fresh clip right now.”
The whole setup is Python plus Node plus FFmpeg, then clone, create and activate a venv, install requirements, install the Remotion render folder, and install Piper as a free local voice engine (11 Labs optional as the paid swap); copying the sample config means literally zero API keys, so the entire pipeline runs locally at $0 and the only spend is the LLM doing the reasoning. Run the install yourself in a fresh venv and record which of the three prerequisites you were missing, then hand the agent a raw one-take clip so you have real footage to iterate on.
8:50
Plan-first beats slop
“now is we're going to give it a bit more detailed command, right? So, I'm going to start a new chat. I'm going to give it a prompt like this, which is my footage is this. Read the...”
"Add cool graphics to my video" is the classic definition of AI slop because the agent applies whatever it feels like; the fix is a prompt that names the footage, tells the agent to read the motion skills catalog first and state which registry blocks it will use, and specifies cadence (motion graphics roughly every five seconds), which yields an implementation plan with visual concepts and exact time ranges you can edit before rendering, and choosing a fast model like Gemini 3.6 Flash keeps that loop cheap where heavier thinking models do not. Write both prompts for the same clip, run them in separate chats, and diff the two outputs so you can see exactly which specifics (catalog read, registry blocks, cadence) bought the improvement.
01
Brief
Start with this video's job: A start-from-scratch walkthrough of Open Montage, the 36,000-star repo that turns a coding agent into a video editor: how its 1,000+ files reduce to five parts, the exact Windows install path (Python, Node, FFmpeg, venv, Remotion render engine, Piper local voice) done inside Antigravity for about zero dollars, and a live side-by-side of a vague prompt producing AI slop versus a detailed prompt that returns a timestamped implementation plan. Treat "Brief" as the outcome you are trying to make visible, not a topic label. Anchor it to 1:50, where the video says: “And five, and this is a very important part, there is no brain in the box. No large language model inside the repo. So, your coding agent reads the manuals, follows the recipes, uses a tool, but then...”
02
Source
Use "Source" to locate the part of the creative automation workflow the video is demonstrating. Ask what changes in your real setup if this claim is true. Anchor it to 6:12, where the video says: “locally on this machine, minus the cost to the large language model API calls, okay? So, the next piece would be for us to record the clip live, okay? I'm going to record a fresh clip right now.”
03
Generation
Turn "Generation" into the reusable artifact for this lesson: A creative workflow board with critique criteria and review checkpoints. This is where watching becomes something you can inspect and reuse.
04
Selection
Use "Selection" as the application surface. Decide whether the idea touches a browser flow, a local file, a model choice, a source document, a UI, or a review step.
05
Edit
Use "Edit" to prove the lesson. The evidence should connect back to the video title, transcript anchors, and a concrete output, not a generic best-practice claim.
06
Taste Review
Use "Taste Review" to carry the idea forward: save the prompt, checklist, diagram, or operating rule that would make the next agent run better.
Example
Source-backed work packet
Convert the video into a scoped task that includes the transcript claim, target workflow, acceptance criteria, and proof. The output should be a creative workflow board with critique criteria and review checkpoints..
Example
Claim vs. demo brief
Separate what the speaker claims, what the demo actually proves, and what still needs outside verification before you adopt the workflow.
Example
Teach-back module
Transform the lesson into a definition, a mechanism diagram, one misconception, one practice exercise, and a check-for-understanding question.
Do not learn it wrong
Treating the title as the lesson without checking what the transcript actually says.
Letting the prompt drift into generic advice that could apply to any video in the playlist.
Copying the tool setup without identifying the operating principle that transfers to your own stack.
Skipping the artifact, which means the learning never becomes operational or inspectable.
Do not count this as learned until these are true.
01
State the transcript-backed claim in your own words: A start-from-scratch walkthrough of Open Montage, the 36,000-star repo that turns a coding agent into a video editor: how its 1,000+ files reduce to five parts, the exact Windows install path (Python, Node, FFmpeg, venv, Remotion render engine, Piper local voice) done inside Antigravity for about zero dollars, and a live side-by-side of a vague prompt producing AI slop versus a detailed prompt that returns a timestamped implementation plan.
02
Explain the practical stakes without hype: New playlist item from AI with Surya; queued for transcript-backed review, topic mapping, and a practical learning artifact.
03
Map the idea onto the Brief -> Source -> Generation -> Selection -> Edit -> Taste Review sequence and name the weakest link.
04
Produce the artifact and include the evidence that proves it: A creative workflow board with critique criteria and review checkpoints.
Put it into practice
Give this grounded prompt to Codex or Claude after watching.
You are helping me turn one specific YouTube video into real, durable learning.
Source video:
- Title: OpenMontage + Antigravity Changed My Editing Game (It's Free)
- URL: https://www.youtube.com/watch?v=LMQwGVSSdDk
- Topic: Creative Automation
- My current learning frame: Install Open Montage in a fresh Python environment, record one unedited 30-second take, then edit it twice: once with a one-line vague prompt and once with a plan-first prompt that names the catalog, registry blocks, and a five-second graphics cadence, and iterate on the returned plan until the timings match your script.
- Why this matters: New playlist item from AI with Surya; queued for transcript-backed review, topic mapping, and a practical learning artifact.
Transcript anchors from this exact video:
- 0:19 / Evidence 1: "background music. Listen. This is generated, too. These captions you are reading, also Open Montage. Actually, the only thing here that is not generated is me. And it knows that. So, what is this thing? 36,000 stars. And..."
- 1:50 / Evidence 2: "And five, and this is a very important part, there is no brain in the box. No large language model inside the repo. So, your coding agent reads the manuals, follows the recipes, uses a tool, but then..."
- 3:44 / Evidence 3: "these are the three commands that you will need to run. The next step will be really to go ahead and clone the repo. So, I'm in this particular folder and I'm going to clone the repo directly..."
- 6:12 / Evidence 4: "locally on this machine, minus the cost to the large language model API calls, okay? So, the next piece would be for us to record the clip live, okay? I'm going to record a fresh clip right now."
- 8:50 / Evidence 5: "now is we're going to give it a bit more detailed command, right? So, I'm going to start a new chat. I'm going to give it a prompt like this, which is my footage is this. Read the..."
- 11:23 / Evidence 6: "is absolutely possible. And that's why I wanted to show you exactly live how I did it with all the commands and prompt, etc. So, So this helps and you're able to leverage this in order to optimize..."
Your task:
1. Use the transcript anchors above as the primary source packet. If you add outside context, label it clearly as outside context and keep it secondary.
2. Create a source-check table with columns: timestamp, claim, what the demo proves, confidence, and what still needs verification.
3. Extract the actual teachable claims from the video. Do not invent claims that are not supported by the title, lesson frame, or transcript anchors.
4. Build a reusable learning artifact: A creative workflow board with critique criteria and review checkpoints.
5. Include:
- a plain-English definition of the core idea
- a diagram or structured model using this sequence: Brief -> Source -> Generation -> Selection -> Edit -> Taste Review
- 3 concrete examples that apply the video idea to real agentic work
- 2 failure modes the video helps prevent
- a checklist I can use the next time I run Codex or Claude
- one practical exercise with a clear done signal
6. Add a "learning transfer" section: what changes in my workflow tomorrow if I actually learned this?
7. Add a "source check" section that cites which transcript anchor supports each major takeaway.
Quality bar:
- Make this specific to "OpenMontage + Antigravity Changed My Editing Game (It's Free)", not a generic Creative Automation essay.
- Prefer operational examples, failure modes, and reusable artifacts over broad definitions.
- Call out uncertainty instead of smoothing over weak evidence.
- If evidence is weak, say what transcript segment or timestamp needs review instead of guessing.
- Finish with a concise artifact I could paste into my learning app.
Misconceptions
What to stop believing.
Creative AI removes the need for taste.
It increases the need for taste because output volume explodes.
The best prompt is enough.
References, critique, iteration, and post-production matter just as much.
Practice studio
Learning only counts when you make something.
01
Transcript evidence map
Separate what the video actually says from what you already believe about the topic.
3 source-backed takeaways with timestamps, confidence, and a transfer note.02
One useful artifact
Apply the video to a real workflow and produce a creative workflow board with critique criteria and review checkpoints..
A reusable artifact with a done signal and one verification step.03
Teach-back card
Explain the lesson to someone who has not watched the video yet.
A 90-second explanation, one diagram, one example, and one misconception to avoid.
Recall check
Answer first, then reveal — without rewatching.
Open Montage has over 1,000 files: what are its five components, and what is deliberately absent?
What must be installed before Open Montage runs, and what does the pipeline actually cost?
How did the detailed prompt differ from "add cool graphics to my video," and what did it produce?
Source shelf
Use the video as a doorway, then verify with primary sources.