Gelalens

Market Prices

Coin Price 24h
BTC Bitcoin
$75,816.7 -2.84%
ETH Ethereum
$2,402.91 -4.46%
SOL Solana
$97.1 -5.49%
BNB BNB Chain
$715.1 -0.54%
XRP XRP Ledger
$1.29 -9.36%
DOGE Dogecoin
$0.0801 -4.38%
ADA Cardano
$0.1950 -6.47%
AVAX Avalanche
$7.26 -4.26%
DOT Polkadot
$0.9418 -6.15%
LINK Chainlink
$10.92 -5.58%

Fear & Greed

51

Neutral

Market Sentiment

Event Calendar

{{年份}}
15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

12
05
halving BCH Halving

Block reward halving event

28
03
unlock Arbitrum Token Unlock

92 million ARB released

18
03
unlock Sui Token Unlock

Team and early investor shares released

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

Altseason Index

42

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
1
Bitcoin
BTC
$75,816.7
1
Ethereum
ETH
$2,402.91
1
Solana
SOL
$97.1
1
BNB Chain
BNB
$715.1
1
XRP Ledger
XRP
$1.29
1
Dogecoin
DOGE
$0.0801
1
Cardano
ADA
$0.1950
1
Avalanche
AVAX
$7.26
1
Polkadot
DOT
$0.9418
1
Chainlink
LINK
$10.92

🐋 Whale Tracker

🔵
0xbc58...f672
3h ago
Stake
31,760 BNB
🔴
0x95c0...67a6
12m ago
Out
3,788.82 BTC
🔵
0xb6b9...e84c
5m ago
Stake
2,227,222 USDT

💡 Smart Money

0x88b5...03dc
Early Investor
+$4.8M
87%
0x23c1...7acd
Early Investor
+$0.8M
68%
0xc5f8...b90f
Market Maker
+$2.6M
90%

🧮 Tools

All →
Analysis

DeepSeek's 50-Point Leap: A Self-Reported Mirage or Genuine Breakthrough?

RayEagle

The numbers are almost too clean. DeepSWE jumps from 12.8 to 62.7—a 49.9-point gain in a single model revision. CyberGym climbs from 52.7 to 83.3. AutomationBench from 12.8 to 31.8. DeepSeek's V4-Pro-0813 self-test report, leaked via monitoring by Dongcha Beating, paints a picture of a model that has reshaped its own frontier. But in the world of immutable code and cryptographic proofs, I've learned one thing: self-reported metrics are the equivalent of a smart contract audit performed by the project's own developer. You trust it at your own risk.

Context: The DeepSeek V4-Pro Lineage

DeepSeek has been a quiet contender in the large language model race, known for its aggressive pricing and open-weight releases. The V4-Pro Preview launched earlier this year at a cost of 3 yuan per million input tokens and 6 yuan per million output tokens—roughly $0.42 and $0.84, respectively. That pricing undercut most Western competitors by a factor of 10, making it a darling for cost-sensitive developers building AI agents. The new 0813 revision, according to the leaked report, maintains the exact same API pricing. No price increase, despite the claimed performance surge.

The benchmarks in question are specifically designed for AI agent performance: DeepSWE measures software engineering task completion, CyberGym evaluates cybersecurity capabilities, AutomationBench tests autonomous workflow execution, and Terminal Bench 2.1 assesses command-line reasoning. These are not your standard MMLU or GSM8K scores. They are high-stakes, multi-step evaluation suites that reflect real-world utility for AI agents executing on-chain tasks, interacting with APIs, or managing smart contracts. For a crypto-native researcher like myself, the parallels are immediate. The same way a Layer 2 project might claim 100x throughput without revealing its testnet conditions, DeepSeek is claiming agentic superiority without third-party validation.

Core: The Numbers and the Pattern

Let's dissect the data. The most striking figure is DeepSWE: from 12.8 to 62.7. That is a 389% improvement. In the context of AI benchmarks, such leaps are rare. Models typically improve by 5-10 points per generation. A 50-point jump suggests either a fundamentally different architecture, extensive fine-tuning on the evaluation set, or a flaw in the measurement harness. The report itself notes that "Agent evaluations heavily rely on Harness"—a reference to the evaluation framework that can be gamed if the model is trained on similar tasks. In crypto, we call this a 'testnet exploit.' The model learns the test, not the domain.

Terminal Bench 2.1: 87.9 vs. Claude Opus 4.8's 85.0. CyberGym: 83.3 vs. 78.3. DeepSWE: 62.7 vs. 58.0. AutomationBench: 31.8 vs. Fable 5's 29.1. DeepSeek's model claims superiority across the board. But note the margin: the largest gap is 5 points on CyberGym. The DeepSWE gap is 4.7 points. AutomationBench is 2.7 points. These are not the kind of blowout victories that would justify a 50-point self-improvement from the Preview version. The Preview's DeepSWE was 12.8—so the 0813 version is 62.7, but Claude Opus 4.8 scores 58.0. That means the Preview was 12.8, which is abysmal, and the new version is now 62.7. The gap between Claude and DeepSeek's new version is only 4.7 points. But the improvement from Preview is 49.9 points. That asymmetry raises a red flag: either the Preview was deliberately underperforming to make the 0813 look better, or the benchmark conditions changed.

From my experience auditing smart contracts, I've seen similar patterns. A project will release a 'v0.1' with glaring vulnerabilities, then a 'v1.0' that 'fixes' them—only to have the same vulnerabilities reappear under different conditions. The code does not lie, but it can be misled. Here, the harness is the contract. If DeepSeek optimized specifically for the DeepSWE evaluation suite, the 50-point jump is not a measure of general intelligence but of overfitting. In crypto, we call that a 'rug pull' of metrics.

Another angle: the price. The API costs remain unchanged. In a bull market for AI (which is analogous to the crypto bull market), companies typically raise prices when they claim superior performance. The fact that DeepSeek has not suggests either a strategic price war or a recognition that the improvements are not yet production-ready. The crypto analogy is clear: a project that claims a breakthrough but keeps its token price low is either generous or hiding something. I lean toward the latter.

Contrarian: The Harness Is the Vulnerability

The contrarian take is not that DeepSeek is lying, but that the metric itself is fragile. The leaked report mentions that "Agent evaluations heavily rely on Harness." What is a harness? It is a set of scripts, environments, and evaluation protocols that execute the agent's actions and compare them to expected outputs. If the harness is not perfectly isolated, the model can learn its patterns. In the world of zero-knowledge proofs, we talk about 'prover overhead'—the cost of verifying a computation. Here, the harness is the verifier. If the verifier is compromised, the proof is worthless.

Consider the implications for AI agents interacting with blockchain. An agent that scores 87.9 on Terminal Bench 2.1 might be great at executing shell commands in a controlled environment, but what about in a live Ethereum mainnet with gas fluctuations, mempool frontrunning, and reorgs? The harness does not include those variables. In crypto, we have the concept of 'mev-resistance'—an agent that performs well in a sandbox but fails in the wild is a liability. The same applies to DeepSeek's claims. Trust is a legacy variable. I do not trust self-reported benchmarks. I trust independent verification.

There is also the question of the test set. DeepSWE is a software engineering benchmark that involves editing code repositories. If DeepSeek trained on the same repositories, the model is essentially looking at the answer key. The 50-point jump could be a result of data contamination. In crypto, we see this with 'audit reports' that only cover surface-level issues. The code may be standard, but the economic attack surface is not tested. Here, the benchmark is the code, and the model is the auditor. If the model has seen the code before, the audit is meaningless.

Takeaway: Wait for the Third-Party Verdict

Until a reputable third party—like MLCommons, Stanford CRFM, or an independent lab—reproduces these results, treat them as aspirational marketing. The crypto-native mindset is to verify, not trust. DeepSeek's V4-Pro-0813 may indeed be a leap forward, but the 50-point surge in DeepSWE is too clean, too convenient, and too coincidental with the lack of price increase. I've seen this pattern before: in 2020, a DeFi project claimed a 10x gas reduction on its new smart contract. The code was released, and the community found that the optimization only worked for a single transaction type. The gas reduction was real, but only in a narrow, gamed scenario.

Code does not lie, but it can be misled. Benchmarks can be manipulated. The only way to know is to run the model yourself. For now, I remain skeptical. The DeepSeek team has a track record of solid engineering, but the leap from 12.8 to 62.7 defies the normal scaling laws of AI. Either they have discovered a new architecture, or they have discovered a new way to game the harness. My money is on the latter. In the world of crypto and AI, the only thing that matters is what happens when the training wheels come off. Until then, these numbers are just noise.