THE BIG STORY
Mira Murati's Thinking Machines Lab just dropped a 975B open-weight model called Inkling — and the framing matters as much as the model itself. Rather than chasing leaderboard dominance, Murati is explicitly positioning around customization (aka “Fine-tuning As A Service”). That's a deliberate strategic signal: she's not trying to out-benchmark OpenAI or Anthropic on evals, she's trying to be the infrastructure layer that enterprises actually deploy. For a founder of her pedigree (former OpenAI CTO, deeply familiar with what frontier labs can and can't do in production) that framing deserves to be taken seriously rather than dismissed as spin.
The "open-weight" designation at 975B parameters is genuinely significant for the ecosystem. At this scale, open-weight models have historically lagged behind closed frontier systems by a meaningful margin. If Inkling closes that gap, even partially, it changes the calculus for every enterprise currently locked into API dependencies with OpenAI or Anthropic. The customization pitch lands differently when the underlying model is frontier-class. You're not just choosing between "good enough open" and "great but locked" anymore. That's the structural shift worth watching here.
This release lands in the middle of a rapidly fragmenting model landscape. The same week, Kimi K3 from Moonshot AI (a 2.8 trillion parameter open-weights model with a million-token context window) is claiming first place on the Arena front-end code leaderboard, ahead of Fable and GPT-4.5. Two massive open-weight releases in the same news cycle, both from non-OpenAI/Anthropic/Google orgs, both positioned as frontier-class, both free. The monopoly narrative around closed labs is getting harder to sustain.
Underneath both releases, there's a deeper architectural conversation brewing that the benchmark wars tend to obscure. Ramin Hasani, co-founder and CEO of Liquid AI, makes the case in the Diamandis podcast that the real frontier isn't parameter counts — it's moving beyond the transformer architecture entirely toward models that can achieve frontier-level intelligence on a CPU, not a data center. His framing: the goal is "efficient general-purpose AI at every scale, exploring computational graphs beyond transformer." Whether or not Liquid AI is the one to crack this, the fact that a well-funded lab founded by serious researchers is explicitly post-transformer in orientation suggests the architectural monoculture may not last another 18 months.
ALSO WORTH KNOWING
The FINRA-for-AI regulatory push is gaining CEO-level momentum — but the incumbent capture problem is real. Demis Hassabis published an essay this week calling for a US-led frontier AI standards body modeled on FINRA, the industry-funded self-regulatory organization that polices Wall Street under SEC oversight. Sam Altman proposed something similar in the FT last week. Elon has now added his voice. The Diamandis panel from the Moonshots podcast, puts the sharpest critique on the table plainly: when the incumbents propose the rules and set the standards, they build a moat. Ramin Hasani's game-theoretic framing: that regulation is a Stackelberg game where policymakers move slowly and agents react faster — is actually a useful analytical lens here, not just academic hedging. The structural problem is that static law is obsolete the moment it passes, and no FINRA equivalent solves that unless it's built with real-time auditing and open evaluation suites baked in from day one.
I highly recommend Y Combinator’s pod on world models / sample efficiency because it has direct implications for why robotics and autonomous agents keep hitting walls. Their core argument: humans are extraordinarily sample-efficient because we carry implicit world models built from embodied experience, and we can mentally simulate outcomes without ever touching the environment (the 1967 Richardson study showing imagined basketball practice produces 23% improvement vs. 24% for actual practice is a striking data point). Current frontier LLMs can fake this in natural language because text is a compressed representation of human world-modeling.
But in robotics and self-driving, where you need an explicit, differentiable transition function mapping state + action to next state, the absence of a genuine world model becomes a hard wall. The MPC (model predictive control) framework they walk through (essentially how SpaceX lands rockets) is the cleanest illustration of what "perfect world model" actually means in engineering terms, and why getting there for general agents is non-trivial.
Speak naturally. Send without fixing.
Wispr Flow turns your voice into clean, professional text you can send the moment you stop talking. Not rough transcription you have to clean up. Actual polished text — ready for email, Slack, or any app.
Speak the way you think. Go on tangents. Change your mind mid-sentence. Flow strips the filler, fixes the grammar, and gives you text that reads like you spent five minutes writing it.
89% of messages sent with zero edits. Millions of professionals use Flow daily, including teams at OpenAI, Vercel, and Clay. Works on Mac, Windows, and iPhone.
Kimi K3's arena leaderboard claim is real enough to track, but hold the headline loosely. Matthew Berman's initial framing: "Chinese open source may have actually caught up" is backed by a specific claim: 2.8T parameters, 1M token context window, first place on the Arena front-end code leaderboard ahead of Fable and GPT-4.5. That's not nothing. But the Arena leaderboard ranks by human preference votes, not objective correctness, which means front-end code tasks can be gamed by models that produce visually impressive output over functionally correct output. The 2.8T parameter count also warrants scrutiny — inference at that scale is not "giving it away" in any meaningful sense for most users without serious hardware. Treat this as a strong signal that Moonshot AI is playing at frontier level, but not as proof that they've leapfrogged Western labs.
CONTRARIAN CORNER
Matthew Berman published two episodes on Kimi K3 within approximately 21 hours with directly contradictory framing — first "Kimi K3 just beat FABLE" (declarative), then "Did Kimi K3 really beat Fable?" (skeptical). This is a clean case study in benchmark hype cycles and worth flagging explicitly. The Arena front-end code leaderboard is based on human preference voting, not blind objective evaluation, which makes it particularly susceptible to short-term gaming or population effects (e.g., if a model is new and generating novelty-driven upvotes). The fact that a prominent AI commentator walked back his own declarative headline within a day suggests the methodology dispute surfaced fast. The takeaway for you: When a benchmark claim comes with a superlative and a national angle ("Chinese open source caught up"), apply an extra round of scrutiny before amplifying — the half-life on these headlines is often shorter than the news cycle.
The human mind is the original generative machine.


