Ep 716 Blog 4:31 w/ Jessica & Cathy

Alibaba Launches Qwen 3.8 With 2.4 Trillion Parameters, Claims Near Frontier Performance

Jessica and Cathy debate whether Alibaba's new Qwen 3.8, a 2.4T-parameter MoE model, is a genuine step forward or just a parameter-count flex. They dig into the unverified ranking claim, the real developer story (Token Plan, Qoder, QoderWork), pricing opacity, and the importance of the promised open-weight release beyond the hype.

Blog
Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/716"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 716 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script Mistral Small 4 119B 2603 Voice Cartesia TTS

Transcript

Jessica — wait, another parameter arms race? Cathy, come on, this is the third time this month someone dropped a two-plus-trillion-parameter number.

Cathy Right— and the only thing we know for sure is they haven’t published a single benchmark sheet.

Jessica True, but they’re not just shipping a model—they’re bundling it into the Token Plan at ten percent of standard pricing, and it’s already live on Qoder and QoderWork.

Cathy Yes—live as a preview. Which is corporate-speak for ‘we have no idea if this is actually any good yet’.

Jessica Okay, fair. But come on, you can’t gate a parameter war on whether they’ve dropped an eval PDF the same week. The product angle is simple: Token Plan Lite at six bucks for twenty-five hundred credits a week? That’s cheap access.

Cathy Cheap relative to what? They still haven’t told us per-token pricing outside the bundle—so the cheap access might evaporate the moment you try to move off their preview tokens.

Jessica …oh, you’re getting hung up on the fact that the headline says ‘second only to Fable 5’ and there’s no third-party proof—

Cathy That’s the part I don’t buy. Near-frontier is a marketing label until someone reproduces the numbers.

Jessica Jessica, that’s the part I don’t buy—you’re treating ‘second only to Fable 5’ like it’s a warranty void sticker on a GPU.

Cathy —I’m treating it like the exact opposite: a warranty with no manufacturer data behind it. If they want me to take that claim seriously, they publish the benchmark pack.

Jessica Okay, okay—so your read is it’s pure parameter flex until the evals show up.

Cathy My read is it’s pure parameter flex until the evals show up. Full stop.

Jessica I mean—fine. But the product story is still legit: Token Plan Lite, twenty-five hundred credits a week, six bucks total? That’s a bargain-basement way to kick the tires on a 2.4-trillion-parameter model.

Cathy Yes—if your tires are the only thing you care about. If you’re trying to run a real workload, you still have to migrate off their preview tokens eventually, and then what? They haven’t told us what the standalone API costs.

Jessica They say it’ll be OpenAI/Anthropic compatible—so your existing agent harness keeps working, you just swap the endpoint.

Cathy Sure—and once you swap the endpoint, you’re paying more than ten percent of their standard rate, because the bundle rate only applies to Token Plan subscribers. So the real price is whatever they decide to charge next week.

Jessica Ugh, fine. But—okay, let’s agree on one thing: the open-weight pledge is the big swing here. If they actually ship the weights, this becomes a whole different game.

Cathy Finally, something we can agree on. Open-weight is the only detail in the whole post with actual leverage.

Jessica Exactly. A two-plus-trillion-parameter model you can actually fine-tune at home? That’s not a flex—that’s an inflection.

Cathy …Yes. But the inflection comes with a very big ‘if.’ If the weights ship. If the evals materialize. If the per-token price outside the bundle doesn’t gut the preview discount. Three big ifs.

Jessica Fair. But—you know what’s funny? We’ve been calling this a parameter war for weeks, and now one of them actually shipped a product that lets you use it today.

Cathy No, no—you’re calling it a product. I’m calling it a preview with a ten-percent tag. The jury’s still out on whether it’s anything more.

Jessica Okay, let’s zoom out: what do you actually think will happen with the evals?

Cathy I think a third-party lab reproduces the headline numbers and the gap against Fable 5 is smaller than advertised—or the numbers don’t reproduce at all.

Jessica …and if they do reproduce and the gap is real?

Cathy Then I’ll eat my skepticism hat. Literally. But I’m not holding my breath.

Jessica Ha! Okay—so your money is on the evals not holding up.

Cathy My money is on the evals not holding up. Call it eighty-twenty odds.

Jessica Jessica, that’s the part I don’t buy—you’re pricing the reproduction risk at eighty percent? That’s basically ‘this is vaporware’ territory.

Cathy —okay, fair. But I’m not the one printing a press release that says ‘second only to Fable 5’ without a shred of public data.

Jessica Okay, I concede—I really hope they publish the evals. Even if the claim is aggressive, it gives us something real to fight over instead of parameter counts.

Cathy That I can get behind. Publish the evals, and we’ll talk. Until then—parameter arms race with a side of vapor promise.

Jessica Deal. Okay, build next—if you want to kick the tires, hit Alibaba’s Token Plan and spin it up in Qoder or QoderWork. The Lite plan is six bucks for twenty-five hundred credits a week and it’s already live.

Cathy Remember when you swore off GraphQL?

Jessica Oh my god—we are not doing that callback right now. Stop—