Stop Guessing Which Model to Use. Build THIS Instead.
Instead of trusting generic model-launch tutorials and public benchmarks, this video shows how to build a personal benchmark: have an AI mine your own past conversations for the tasks you actually do, then score new models against your own rubric with a /benchmark slash command.
Mark Kashef8 minTranscript found
Quick learning frame
Read this before watching.
Creative automation uses agents to accelerate production while keeping human taste in story, pacing, selection, and critique.
New playlist item from Mark Kashef; queued for transcript-backed review, topic mapping, and a practical learning artifact.
Skill you build: The ability to build and run your own task-based model evaluation from your real work history so you can decide, on evidence, whether a new model or effort level is worth switching to.
Watch for the shift from claim to mechanism. The learning value is the point where the transcript reveals a repeatable action, tool boundary, context move, review habit, or artifact.
Concept diagram
Where this video fits.
01Brief
02Source
03Generation
04Selection
05Edit
06Taste Review
Deep lesson
Turn this video into working knowledge.
1,786 cleaned transcript words reviewed across 488 timed caption segments.
Thesis
Stop Guessing Which Model to Use. Build THIS Instead. teaches a practical creative automation move: Instead of trusting generic model-launch tutorials and public benchmarks, this video shows how to build a personal benchmark: have an AI mine your own past conversations for the tasks you actually do, then score new models against your own rubric with a /benchmark slash command.
The goal is not to remember the video. The goal is to extract the operating principle, tie it to timestamped evidence, test how far the claim transfers, and make something reusable.
0:11
Benchmarks miss your work
“them structured in the exact same way. Insert name of model is absolutely insane and changes everything. You then are lured to click on set tutorials. And then in the tutorial, you watch someone build a website you're...”
Every model release triggers near-identical tutorials building apps you will never use, and public benchmarks like WebArena or Math Olympiad do not reflect your actual work, which is client emails, briefs, repo audits, and lead lists, so those numbers rarely tell you if the model helps you. Write down the handful of concrete tasks that make up most of your day and ask which public benchmark, if any, actually tests them.
3:18
Mine your own history
“whatever brand new model comes out, whether it's Claude Code, Codex, Kimmy, you can keep evolving the skill as you want. And like I'll show you in just a few moments with all of these open tabs, you'll...”
Point Claude or Codex at your past chat histories (the creator has 1000+ conversations, about 2GB) to isolate the recurring tasks you really execute, then re-run those exact tasks on the new model; to avoid the bias of a model grading itself, you write your own rubric based on the metrics you care about, including vibe, verbosity, and sassiness. Have an AI scan your conversation history, list your most common recurring tasks, and draft a personal rubric defining what a good result means to you.
5:35
The /benchmark command
“you would choose. So, this is an example of a completed report. And it tells you exactly what you're looking for, which is it tested the task, but it was dead even. Meaning, it's not worth it for...”
A reusable /benchmark slash command compares models and effort levels (for example Opus 4.8 vs Opus 5 vs Fable 5 on low/high) against your rubric, spawning a live artifact table that runs 30-40 minutes and scores quality, instruction fidelity, token efficiency, speed, and number of turns, then returns a verdict on whether switching is actually worth it. Run a /benchmark-style comparison of two models or two effort levels on one of your real tasks and read the verdict to decide if the costlier option earns its price.
01
Brief
Start with this video's job: Instead of trusting generic model-launch tutorials and public benchmarks, this video shows how to build a personal benchmark: have an AI mine your own past conversations for the tasks you actually do, then score new models against your own rubric with a /benchmark slash command. Treat "Brief" as the outcome you are trying to make visible, not a topic label. Anchor it to 0:11, where the video says: “them structured in the exact same way. Insert name of model is absolutely insane and changes everything. You then are lured to click on set tutorials. And then in the tutorial, you watch someone build a website you're...”
02
Source
Use "Source" to locate the part of the creative automation workflow the video is demonstrating. Ask what changes in your real setup if this claim is true. Anchor it to 3:18, where the video says: “whatever brand new model comes out, whether it's Claude Code, Codex, Kimmy, you can keep evolving the skill as you want. And like I'll show you in just a few moments with all of these open tabs, you'll...”
03
Generation
Turn "Generation" into the reusable artifact for this lesson: A creative workflow board with critique criteria and review checkpoints. This is where watching becomes something you can inspect and reuse.
04
Selection
Use "Selection" as the application surface. Decide whether the idea touches a browser flow, a local file, a model choice, a source document, a UI, or a review step.
05
Edit
Use "Edit" to prove the lesson. The evidence should connect back to the video title, transcript anchors, and a concrete output, not a generic best-practice claim.
06
Taste Review
Use "Taste Review" to carry the idea forward: save the prompt, checklist, diagram, or operating rule that would make the next agent run better.
Example
Source-backed work packet
Convert the video into a scoped task that includes the transcript claim, target workflow, acceptance criteria, and proof. The output should be a creative workflow board with critique criteria and review checkpoints..
Example
Claim vs. demo brief
Separate what the speaker claims, what the demo actually proves, and what still needs outside verification before you adopt the workflow.
Example
Teach-back module
Transform the lesson into a definition, a mechanism diagram, one misconception, one practice exercise, and a check-for-understanding question.
Do not learn it wrong
Treating the title as the lesson without checking what the transcript actually says.
Letting the prompt drift into generic advice that could apply to any video in the playlist.
Copying the tool setup without identifying the operating principle that transfers to your own stack.
Skipping the artifact, which means the learning never becomes operational or inspectable.
Do not count this as learned until these are true.
01
State the transcript-backed claim in your own words: Instead of trusting generic model-launch tutorials and public benchmarks, this video shows how to build a personal benchmark: have an AI mine your own past conversations for the tasks you actually do, then score new models against your own rubric with a /benchmark slash command.
02
Explain the practical stakes without hype: New playlist item from Mark Kashef; queued for transcript-backed review, topic mapping, and a practical learning artifact.
03
Map the idea onto the Brief -> Source -> Generation -> Selection -> Edit -> Taste Review sequence and name the weakest link.
04
Produce the artifact and include the evidence that proves it: A creative workflow board with critique criteria and review checkpoints.
Put it into practice
Give this grounded prompt to Codex or Claude after watching.
You are helping me turn one specific YouTube video into real, durable learning.
Source video:
- Title: Stop Guessing Which Model to Use. Build THIS Instead.
- URL: https://www.youtube.com/watch?v=3ICM9ZdflZA
- Topic: Creative Automation
- My current learning frame: Have an AI extract your three most common real tasks from your chat history, write a rubric scoring quality, instruction fidelity, token efficiency, speed, and turns, then benchmark two models or effort levels on those tasks and record which one actually wins for you.
- Why this matters: New playlist item from Mark Kashef; queued for transcript-backed review, topic mapping, and a practical learning artifact.
Transcript anchors from this exact video:
- 0:11 / Evidence 1: "them structured in the exact same way. Insert name of model is absolutely insane and changes everything. You then are lured to click on set tutorials. And then in the tutorial, you watch someone build a website you're..."
- 1:45 / Evidence 2: "how it was executed by prior models, and run those exact same tasks and projects with this brand new model. So, whether it's rewriting a cold email and testing that out on different effort levels, or drafting some..."
- 3:18 / Evidence 3: "whatever brand new model comes out, whether it's Claude Code, Codex, Kimmy, you can keep evolving the skill as you want. And like I'll show you in just a few moments with all of these open tabs, you'll..."
- 5:35 / Evidence 4: "you would choose. So, this is an example of a completed report. And it tells you exactly what you're looking for, which is it tested the task, but it was dead even. Meaning, it's not worth it for..."
- 7:56 / Evidence 5: "content and resources that I put out all the time including our Claude Code and Codex Living Course, then check out the first thing down below and join me in my early A adopters community. And for the..."
Your task:
1. Use the transcript anchors above as the primary source packet. If you add outside context, label it clearly as outside context and keep it secondary.
2. Create a source-check table with columns: timestamp, claim, what the demo proves, confidence, and what still needs verification.
3. Extract the actual teachable claims from the video. Do not invent claims that are not supported by the title, lesson frame, or transcript anchors.
4. Build a reusable learning artifact: A creative workflow board with critique criteria and review checkpoints.
5. Include:
- a plain-English definition of the core idea
- a diagram or structured model using this sequence: Brief -> Source -> Generation -> Selection -> Edit -> Taste Review
- 3 concrete examples that apply the video idea to real agentic work
- 2 failure modes the video helps prevent
- a checklist I can use the next time I run Codex or Claude
- one practical exercise with a clear done signal
6. Add a "learning transfer" section: what changes in my workflow tomorrow if I actually learned this?
7. Add a "source check" section that cites which transcript anchor supports each major takeaway.
Quality bar:
- Make this specific to "Stop Guessing Which Model to Use. Build THIS Instead.", not a generic Creative Automation essay.
- Prefer operational examples, failure modes, and reusable artifacts over broad definitions.
- Call out uncertainty instead of smoothing over weak evidence.
- If evidence is weak, say what transcript segment or timestamp needs review instead of guessing.
- Finish with a concise artifact I could paste into my learning app.
Misconceptions
What to stop believing.
Creative AI removes the need for taste.
It increases the need for taste because output volume explodes.
The best prompt is enough.
References, critique, iteration, and post-production matter just as much.
Practice studio
Learning only counts when you make something.
01
Transcript evidence map
Separate what the video actually says from what you already believe about the topic.
3 source-backed takeaways with timestamps, confidence, and a transfer note.02
One useful artifact
Apply the video to a real workflow and produce a creative workflow board with critique criteria and review checkpoints..
A reusable artifact with a done signal and one verification step.03
Teach-back card
Explain the lesson to someone who has not watched the video yet.
A 90-second explanation, one diagram, one example, and one misconception to avoid.
Recall check
Answer first, then reveal — without rewatching.
Why does the video argue that standard model benchmarks and launch tutorials are not useful for most people?
How do you build a personal benchmark from your own work, and how do you avoid biased grading?
What does the /benchmark slash command produce, and what does it measure?
Source shelf
Use the video as a doorway, then verify with primary sources.