The model is broken. Not the code — the narrative.
A press release, a blog post, and a single metric: MCP score. That’s the entire evidence package for Thinking Machines Lab’s Inkling model, touted as the "best Western open-source model." My 2018 audit of Bancor’s integer overflow taught me one thing: trust is not a transitive property of reputation. It must be verified at the stack level. Here, the stack is empty. Math has no mercy, and this claim has no proof.
Context: The Silence, the Hype, and the Void
Mira Murati’s team emerged from two years of stealth with a single announcement: a model trained (or fine-tuned) for Agent tasks, scoring "impressively" on the Model Context Protocol (MCP) benchmark. It’s live on OpenRouter, open-source, and marketed as the best offering from the Western world. The implicit comparison is to DeepSeek, Qwen, and others—but the only explicit metric is MCP. This is a classic signal gap. In DeFi, we see this pattern with yield farms: a single APY figure with no auditor’s signature. Rug pulls are just bad code. Here, the code isn’t even visible. The protocol’s GitHub? Not specified. The license? Unknown. The training compute? Redacted.
Core: Dissecting the Stack of Hype
Let’s apply the forensic methodology I developed during the 2020 DeFi yield trap analysis. When Compound offered 20% APY, I didn’t celebrate—I built a token emission model. The same logic applies here.

Technical Teardown: The sole metric—MCP score—is a proxy for context-aware tool calling, not for general intelligence. No MMLU, no HumanEval, no GSM8K. For an “open-source” model to claim superiority without publishing standard benchmarks is like a lending protocol claiming solvency without publishing a balance sheet. The absence of data is the data. It suggests either the model performs poorly on broad tests, or the team is deliberately narrowing the evaluation scope. In either case, the “best” label is a marketing vector, not a technical truth.

In my 2022 Terra/Luna post-mortem, I flagged the UST death spiral because the model lacked external collateral. Here, Inkling’s “architecture” lacks any external validation. We don’t know if it’s a fine-tuned Llama 3.1, a Mixture-of-Experts variant, or a 7B parameter toy. The compute cost? The inference latency? The energy per token? Silence. High yield, high graveyard. A model with no publicly verifiable stack is a liability, not an asset.
Commercial Teardown: “Open-source” and “API revenue” are often a contradictory pair. If Inkling is truly open-weight (Apache 2.0), its commercial value is limited to enterprise support and custom deployments—a thin margin business. If it’s a restrictive license (like “OpenRAIL-M”), then the “open” label is diluted. The OpenRouter listing gives a clue: the platform aggregates models for developers to test. This is a customer acquisition play, not a revenue model. The unit economics are opaque. What’s the per-token cost? Is the company solvent? Without this, we’re speculating.
Systemic Risk: The model’s focus on MCP (tool use) introduces a new vector: agentic autonomy. A model that can call APIs and execute actions is a cascade of risk. A single toxicity in the training data, a single malicious prompt, and the agent can delete your database or transfer funds. The Terra death spiral was a system-of-systems failure. An “open-source agent” is a system-of-systems waiting to fail. Trust the stack, not the story.
Contrarian: What If the Bulls Are Right?
Counter-intuitively, the lack of data could be a deliberate strategy to avoid hype, not to inflate it. Mira Murati’s background at OpenAI includes safety and alignment—she knows the damage of premature claims. The MCP focus might be a genuine innovation: if Inkling sets a new standard for agent frameworks, it could become the computational primitive of the next wave of automation. In that case, the “best Western” claim would be self-fulfilling—adoption creates quality.
But here’s the math: a model’s value is not its peak score on a custom benchmark. It’s the integral of its utility across diverse, adversarial real-world tasks. Without those tests, we’re betting on a name, not a proof. The 2024 Bitcoin ETF custody scrutiny taught me that institutional trust requires independent audits. Inkling needs a trust minimisation protocol: open-weight release, reproducible benchmarks, and a public challenge to the community.
Takeaway: Verify or Ignore
Thinking Machines Lab has two paths: publish the full technical report and enable third-party audit, or remain a narrative-driven project. My framework says: Math has no mercy. No data, no investment of time. The highest yield is still in scepticism. I’ll wait for the GitHub repo, the Hugging Face page, and the standard evaluations. Until then, this model is a black box with a shiny sticker. Rug pulls are just bad code. This one hasn’t shown its code.