Creative Automation / Foundation

Is Inkling AI the Ultimate Open Source Model? Full Test

A hands-on test of Inkling, a 975B-parameter (41B-active) mixture-of-experts model built for agentic coding and tool use, covering its self-fine-tuning demo, native multimodal reasoning, and three live tests: a Kanban app build, a calibrated Mars-2035 probability forecast, and a formatted long-form writing task.

Ray Codes11 minTranscript found

Quick learning frame

Read this before watching.

Creative automation uses agents to accelerate production while keeping human taste in story, pacing, selection, and critique.

New playlist item from Ray Codes; queued for transcript-backed review, topic mapping, and a practical learning artifact.

Skill you build: The ability to design test prompts that probe an AI model across coding execution, calibrated forecasting, and long-form formatting, rather than relying on a single benchmark number.

Watch for the shift from claim to mechanism. The learning value is the point where the transcript reveals a repeatable action, tool boundary, context move, review habit, or artifact.

Concept diagram

Where this video fits.

01Brief
02Source
03Generation
04Selection
05Edit
06Taste Review

Deep lesson

Turn this video into working knowledge.

2,275 cleaned transcript words reviewed across 670 timed caption segments.

Thesis

Is Inkling AI the Ultimate Open Source Model? Full Test teaches a practical creative automation move: A hands-on test of Inkling, a 975B-parameter (41B-active) mixture-of-experts model built for agentic coding and tool use, covering its self-fine-tuning demo, native multimodal reasoning, and three live tests: a Kanban app build, a calibrated Mars-2035 probability forecast, and a formatted long-form writing task.

The goal is not to remember the video. The goal is to extract the operating principle, tie it to timestamped evidence, test how far the claim transfers, and make something reusable.

0:00

Sparse but Massive

“Writing code with AI is great, but getting it to use external tools is much harder. Today we are looking at Inkling, a newly dropped model that excels at agentic workflows. Let us look at the sheer scale...”

Inkling is a 975B-parameter mixture-of-experts model that only activates 41B parameters per query (with a lighter 'Inkling small' variant using 12B active parameters), supports a 1M-token context, and reasons natively over text, image, and audio without converting audio to text first; it was even used to fine-tune itself entirely on its own, writing its own synthetic training data to learn a constrained 'no letter E' language task. Write down Inkling's active-vs-total parameter split (41B of 975B) next to a model you already use, to build intuition for how mixture-of-experts routing changes the hardware you actually need.

4:42

Agentic Coding Muscle

“functional web application in single short just from a simple text prompt. And then it used an embedded browser agent to actually click and interact with the final application. But it is not just good at single prompt...”

Inkling was purpose-built for agentic coding and tool use, with official demos showing it build a full web app from a single prompt and then use an embedded browser agent to click through and test it, plus a multiplayer snake game with real-time server bots; it matches Nemotron 3 Ultra's score on the Terminal Bench 2.1 coding benchmark while using only a third of the tokens, and it's deployable locally via SGLang, VLLM, or Unsloth, with a special NVFP4 checkpoint for Blackwell GPUs and a 98% strong-reject safety score. Pick one of the three deployment paths (SGLang, VLLM, Unsloth) and read its docs for loading a model with an adjustable thinking-effort setting, since that's the lever Inkling exposes for trading cost against quality.

7:40

Calibrated Forecasting

“handle all the complex logic and the strict constraints that I had given in the initial prompt. And the second task was to check whether the human will land on Mars and return to Earth by 2035. And...”

Asked to forecast the probability of a human Mars round trip by 2035, Inkling didn't just guess: it gave roughly a 1% probability, walked through launch-window arithmetic, SpaceX's 'no earlier than 2028' Starship timeline, and life-support and certification requirements, then explicitly stated what evidence (like reusable orbital refueling by 2028) would raise its estimate above 10%. Give a reasoning model your own uncertain prediction question and require it to state both a probability and the specific evidence that would change that probability, then check whether its stated conditions are falsifiable.

01

Brief

Start with this video's job: A hands-on test of Inkling, a 975B-parameter (41B-active) mixture-of-experts model built for agentic coding and tool use, covering its self-fine-tuning demo, native multimodal reasoning, and three live tests: a Kanban app build, a calibrated Mars-2035 probability forecast, and a formatted long-form writing task. Treat "Brief" as the outcome you are trying to make visible, not a topic label. Anchor it to 0:00, where the video says: “Writing code with AI is great, but getting it to use external tools is much harder. Today we are looking at Inkling, a newly dropped model that excels at agentic workflows. Let us look at the sheer scale...”

02

Source

Use "Source" to locate the part of the creative automation workflow the video is demonstrating. Ask what changes in your real setup if this claim is true. Anchor it to 4:42, where the video says: “functional web application in single short just from a simple text prompt. And then it used an embedded browser agent to actually click and interact with the final application. But it is not just good at single prompt...”

03

Generation

Turn "Generation" into the reusable artifact for this lesson: A creative workflow board with critique criteria and review checkpoints. This is where watching becomes something you can inspect and reuse.

04

Selection

Use "Selection" as the application surface. Decide whether the idea touches a browser flow, a local file, a model choice, a source document, a UI, or a review step.

05

Edit

Use "Edit" to prove the lesson. The evidence should connect back to the video title, transcript anchors, and a concrete output, not a generic best-practice claim.

06

Taste Review

Use "Taste Review" to carry the idea forward: save the prompt, checklist, diagram, or operating rule that would make the next agent run better.

Example

Source-backed work packet

Convert the video into a scoped task that includes the transcript claim, target workflow, acceptance criteria, and proof. The output should be a creative workflow board with critique criteria and review checkpoints..

Example

Claim vs. demo brief

Separate what the speaker claims, what the demo actually proves, and what still needs outside verification before you adopt the workflow.

Example

Teach-back module

Transform the lesson into a definition, a mechanism diagram, one misconception, one practice exercise, and a check-for-understanding question.

Do not learn it wrong
  • Treating the title as the lesson without checking what the transcript actually says.
  • Letting the prompt drift into generic advice that could apply to any video in the playlist.
  • Copying the tool setup without identifying the operating principle that transfers to your own stack.
  • Skipping the artifact, which means the learning never becomes operational or inspectable.

Transcript-derived moments

Use timestamps to study the actual video.

Quality check

Do not count this as learned until these are true.

01

State the transcript-backed claim in your own words: A hands-on test of Inkling, a 975B-parameter (41B-active) mixture-of-experts model built for agentic coding and tool use, covering its self-fine-tuning demo, native multimodal reasoning, and three live tests: a Kanban app build, a calibrated Mars-2035 probability forecast, and a formatted long-form writing task.

02

Explain the practical stakes without hype: New playlist item from Ray Codes; queued for transcript-backed review, topic mapping, and a practical learning artifact.

03

Map the idea onto the Brief -> Source -> Generation -> Selection -> Edit -> Taste Review sequence and name the weakest link.

04

Produce the artifact and include the evidence that proves it: A creative workflow board with critique criteria and review checkpoints.

Put it into practice

Give this grounded prompt to Codex or Claude after watching.

You are helping me turn one specific YouTube video into real, durable learning.

Source video:
- Title: Is Inkling AI the Ultimate Open Source Model? Full Test
- URL: https://www.youtube.com/watch?v=U7IX2607jCM
- Topic: Creative Automation
- My current learning frame: Run the same three-test battery (a self-contained coding build, a calibrated probability forecast, and a long-form formatted writing task) against a local open-source model you can access, and compare its outputs to Inkling's results described in the video.
- Why this matters: New playlist item from Ray Codes; queued for transcript-backed review, topic mapping, and a practical learning artifact.

Transcript anchors from this exact video:
- 0:00 / Evidence 1: "Writing code with AI is great, but getting it to use external tools is much harder. Today we are looking at Inkling, a newly dropped model that excels at agentic workflows. Let us look at the sheer scale..."
- 2:52 / Evidence 2: "ultimate coding test where I'm going to test the agentic coding and the tool use capabilities of the Inkling model to see if it is actually able to build functional applications. I'm asking it to build a full..."
- 4:42 / Evidence 3: "functional web application in single short just from a simple text prompt. And then it used an embedded browser agent to actually click and interact with the final application. But it is not just good at single prompt..."
- 7:40 / Evidence 4: "handle all the complex logic and the strict constraints that I had given in the initial prompt. And the second task was to check whether the human will land on Mars and return to Earth by 2035. And..."
- 10:33 / Evidence 5: "raw multi-model power with the local hardware limits. I'll be putting the link to the official blog, the steps to install this model in your local system, or to try it on online platforms, and the three test..."

Your task:
1. Use the transcript anchors above as the primary source packet. If you add outside context, label it clearly as outside context and keep it secondary.
2. Create a source-check table with columns: timestamp, claim, what the demo proves, confidence, and what still needs verification.
3. Extract the actual teachable claims from the video. Do not invent claims that are not supported by the title, lesson frame, or transcript anchors.
4. Build a reusable learning artifact: A creative workflow board with critique criteria and review checkpoints.
5. Include:
   - a plain-English definition of the core idea
   - a diagram or structured model using this sequence: Brief -> Source -> Generation -> Selection -> Edit -> Taste Review
   - 3 concrete examples that apply the video idea to real agentic work
   - 2 failure modes the video helps prevent
   - a checklist I can use the next time I run Codex or Claude
   - one practical exercise with a clear done signal
6. Add a "learning transfer" section: what changes in my workflow tomorrow if I actually learned this?
7. Add a "source check" section that cites which transcript anchor supports each major takeaway.

Quality bar:
- Make this specific to "Is Inkling AI the Ultimate Open Source Model? Full Test", not a generic Creative Automation essay.
- Prefer operational examples, failure modes, and reusable artifacts over broad definitions.
- Call out uncertainty instead of smoothing over weak evidence.
- If evidence is weak, say what transcript segment or timestamp needs review instead of guessing.
- Finish with a concise artifact I could paste into my learning app.

Misconceptions

What to stop believing.

Creative AI removes the need for taste.

It increases the need for taste because output volume explodes.

The best prompt is enough.

References, critique, iteration, and post-production matter just as much.

Practice studio

Learning only counts when you make something.

01

Transcript evidence map

Separate what the video actually says from what you already believe about the topic.

3 source-backed takeaways with timestamps, confidence, and a transfer note.
02

One useful artifact

Apply the video to a real workflow and produce a creative workflow board with critique criteria and review checkpoints..

A reusable artifact with a done signal and one verification step.
03

Teach-back card

Explain the lesson to someone who has not watched the video yet.

A 90-second explanation, one diagram, one example, and one misconception to avoid.

Recall check

Answer first, then reveal — without rewatching.

How many of Inkling's 975 billion total parameters are actually activated per query, and why does that matter for running it on consumer hardware?

What benchmark result does the video cite to show Inkling's coding efficiency compared to Nemotron 3 Ultra?

What probability did Inkling assign to a human landing on Mars and returning safely by 2035, and what would change that estimate?

Source shelf

Use the video as a doorway, then verify with primary sources.

ReadingComfyUIwww.comfy.org/ReadingAffinityaffinity.serif.com/