Ep 687 Blog 5:22 w/ Talon & Wildflower

Inkling: Our open Weights model

Talon and Wildflower dig into Thinking Machines’ new open-weights model, Inkling — its 975B parameter MoE, 1M context window, native multimodality, and self-fine-tuning demo — and ask who actually needs another 41B active parameter behemoth, whether the benchmarks hold up, and whether the real win is the Tinker platform beneath it.

Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/687"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 687 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script Mistral Small 4 119B 2603 Voice Rime Mist v3

Transcript

Talon So. A one-point-three-trillion-parameter open MoE with a one-million-token context, an interactive lipogram fine-tune that literally forbids the letter e, and a company called Thinking Machines that just dropped a model called Inkling.

Wildflower Yeah. Also ninety-seven point five teraflops total and forty-one teraflops active. Also native multimodality. Also, and I quote, ‘not the strongest overall model available today.’

Talon Which, okay, this is going to sound insane on a podcast — but also, fine, I’ll say it: I’m weirdly jazzed about the self-fine-tune demo.

Wildflower Sure. The demo where the model writes its own fine-tune objective, spins up 96 steps, self-evaluates, switches versions, and then cheerfully tells you how to throw a release party when your team ships an LLM.

Talon Right — and all of that runs inside their Tinker platform, which is where the real work is happening. Forget the ninety-seven-five, did you see how they specify controllable thinking effort inside the harness? That’s a real knob.

Wildflower It’s not a model knob; it’s the runtime knob. The Mixture-of-Experts is cool, the training run is clean, but the novelty is the control surface. You can dial thinking effort from the harness while the model’s executing. That’s the part that actually ships.

Talon Yeah, but the benchmarks they show look weirdly broad — they’re bragging about generalist scores across agentic coding, reasoning, vision, audio. Meanwhile they straight-up say it’s not the strongest overall model. Come on, Wildflower. That’s the classic ‘let’s not define the category, just ship a thing.’

Wildflower Exactly. The spider chart is a hedge. It’s not ‘our model wins here,’ it’s ‘our model doesn’t lose horribly anywhere.’ And in open-weights-land, that passes for a flex.

Talon I don’t know — the Design Arena web-dev leaderboard screenshot says something. They’re twelfth out of twenty. Not exactly a knockout, but solidly in the chase pack.

Wildflower Solidly mediocre. Which, again, is the entire pitch: broad, balanced, a foot in every door. The problem is, if you’re a product team evaluating a base model, ‘broad and balanced’ doesn’t pay the bills. You need ‘this specific benchmark at the top.’

Talon Okay, your point — but tell me who actually needs a forty-one-billion-active open model that can run a one-million-token context at low cost. Like, who has the compute infra to even light this thing up outside a demo?

Wildflower Nobody outside the Thinking Machines orbit. But the teams already inside their platform — the ones who live inside Tinker — they’ll care. The platform is the product, not the weights.

Talon So the bet is on the loop, not the model.

Wildflower Yes. The loop, the harness randomization, the end-to-end scripted run. The lipogram fine-tune is cute — it’s a parlor trick — but the real takeaway is the training API and the fact that you can hand the model a script and it can execute.

Talon I mean, imagine trying to reproduce that demo outside Tinker. You’d need a GPU cluster, a training harness, an eval harness, a version switcher, and a plus-one at the release party. Most teams would look at the ninety-seven-five and run.

Wildflower Which is exactly what they’re banking on. They’re selling the platform. The model is the Trojan horse.

Talon …Okay, right, and the pricing page just read my mind — cached prefill eighty percent off, but train prices up ten percent starting tomorrow. So your ‘resistance to censorship’ open model also has a pricetag that compounds.

Wildflower Mm-hm. Cost discipline is part of the play. The demo’s cheap enough to run; the fine-tuning is where the bill matters. Most teams will try the demo, maybe kick the tires on the playground, then ghost when the Tinker invoice lands.

Talon Come on, Wildflower — you’re saying the entire release is a cynical pricing strategy on top of a not-quite-peak model.

Wildflower I’m saying the headline numbers are a deployment stunt, the real product is the platform, and the pricing is the sharp edge of the wedge. It’s not cynical — it’s honest.

Talon That’s such an Exploring Next take. I love it. But I still think the self-fine-tuning demo is a genuine developer-experience leap.

Wildflower Sure. The demo where the model writes its own objective in code, spins up 96 steps, self-evaluates, and then cheerfully tells you how to throw a release party.

Talon Exactly. It collapsed a painful workflow. That’s the part that ships.

Wildflower …We’ve done how many episodes now? Six, seven hundred? And you’re only now realizing that the workflow collapse is the product?

Talon Okay, fair. Still — I’d put seventy-thirty this ships a public beta before spring.

Wildflower Wildflower: ‘I think the docs stay prestige-biased longer than they should, even if teams quietly pick Terra anyway.’ I’m still on the other side of that wager, but fine, let’s go ahead.

Talon Good. You’re still allowed to be wrong.