Claude Opus 4.8 Costs 57.1× More and Loses All Five Benchmarks. What Beat It Was Not a Model, but the Harness

Cost: 57.1× cheaper; Results: wins on all five third-party benchmarks

Same model, only the Harness changes

HLR: gain from switching the Harness ÷ gain from switching the model

Floatboat says the same DeepSeek-V4-Flash model outperformed Claude Opus 4.8 across five benchmarks when paired with Floatboat Harness.

LOS ANGELES, CA, UNITED STATES, August 28, 2026 /EINPresswire.com/ -- Floatboat says the same $0.14 DeepSeek-V4-Flash model, when paired with Floatboat Harness, outperformed Claude Opus 4.8 across all five benchmarks. The model did not change, and the cost did not change; the only variable was the Harness.

In January 2026, the Floatboat desktop app turned the file manager, editor, and browser into the agent runtime environment. In April, FloatIM connected people and agents — and agents and agents — into a shared office network. In May, FloatSchedule let agents begin work on their own, on schedule, without waiting for instructions. None of those launches was simply an interface innovation. To make Floatboat a real agent runtime environment, the Runtime had to support read/write privileges and sandbox boundaries. To keep context intact across handoffs, Infra had to support persistence and concurrency. To let agents begin unsupervised and still finish work, the Agent Loop had to converge on long-horizon tasks.

On August 7, Floatboat put that foundation on the leaderboard with complete evaluation data across five third-party benchmarks. The base model was DeepSeek-V4-Flash, one of the cheapest models on the market, and the results surpassed a roster of top-tier overseas models. This matters because the industry has long accepted the equation: Model + Harness = Agent. The model side is measured constantly. The other half — what the model can touch, how loops converge, what tools return, and where state lives — has not had a clear ruler. Floatboat says this data shows what that side is worth: more, faster, better, cheaper.

DeepSeek-V4-Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens. Under a typical 3:1 input-output ratio, that comes to a blended $0.175/M. Claude Opus 4.8 sits at a blended $10/M, or 57.1× more expensive. Floatboat says all five benchmarks were surpassed.

The same DeepSeek-V4-Flash, running on DeepSeek’s own official Harness, scored 54.4 on DeepSWE, below Opus 4.8’s 58.0. Put it on Floatboat, and it reached 67.25. On Terminal Bench 2.1, DeepSeek’s official Harness tied Opus 4.8 at 82.7 : 82.7 — and lost on the other four. In other words, the same model inside the model provider’s own system did not beat Opus 4.8 on a single benchmark. Connected to Floatboat, it won all five. The model weights did not change. The unit cost did not change. There was only one variable.

Both sides used the DeepSeek-V4-Flash 0731 model base. DeepSeek’s official public benchmarks used the minimalist mode of DeepSeek Harness, configured at top_p=0.95 and temperature=1.0. Floatboat used the Floatboat Evaluation Harness, with model inference for most benchmarks provided by DeepSeek’s official API. All tasks ran independently in isolated sandbox environments with identical input parameters. Floatboat says the cheapest model was chosen deliberately so the result could be attributed only to the execution system.

The same-base Harness deltas rise in order: 1.9% → 9.6% → 12.6% → 19.9% → 23.6%. Ordered by task horizon from short to long, the tiers are non-decreasing and the gains rise monotonically. The longer the task horizon, the larger the Harness gain. Real work, Floatboat argues, is long-horizon work. On BrowseComp — a leaderboard created by OpenAI itself — Floatboat’s 87.80 surpassed GPT-5.6 Terra (87.5), Claude Sonnet 5 (84.7), Opus 4.8 (84.3), and GPT-5.6 Luna (83.3).

Floatboat defines the Harness boundary simply: all work involved other than the model itself falls within the scope of the Harness. The companion metric is HLR, or Harness Leverage Ratio. HLR = Gain from changing the Harness / Gain from changing the model. On DeepSWE, DeepSeek’s official system scored 54.4, Claude Opus 4.8 scored 58.0, and the same DeepSeek model on Floatboat reached 67.25. That gives HLR = 12.85 / 3.6 = 3.57, or about 3.6×.

Floatboat says long-horizon tasks usually fail not because “some step went wrong,” but because the loop does not converge. Real work is not a textbook problem with a standard answer: human intent often becomes clear only as intermediate artifacts appear, and many critical criteria do not exist in the first sentence. In long-horizon tasks, model drift gets amplified turn by turn. A loop that only runs forward gets further wrong the longer it runs; a loop that looks back and checks can stay aligned.

That judgment maps directly to the stack: Runtime determines what the agent can touch; Agent Loop determines whether a long-horizon task can converge; Tools determine whether the model can correctly consume results; Infra determines whether failure can even be found. If any one of those layers belongs to someone else, the ceiling is still someone else’s ceiling.

The client environment that users actually run is stronger than the evaluation environment. Floatboat Desktop integrates the file manager, the file editor, and the browser natively into the product, and connects to 3,000+ online services and a range of multimodal models. Floatboat says the same benchmark set run in the full client environment produces higher scores than the figures disclosed in this article.

“Cheaper” is an open ledger: a blended $0.175/M against $10/M, with the beaten rival costing 57.1× more. “Faster” refers to how that price changes usage behavior. When long-horizon work costs almost nothing, rerunning a failed run also costs almost nothing. More, faster, better, cheaper — Floatboat says all four land together here.

Every number above can be recomputed. The entry point is floatboat.ai. AOE Tech Labs has also launched a DeepSeek-based edition — ¥5 / $1 for the first month — built on the same base model, DeepSeek-V4-Flash. The full report is available at https://floatboat.ai/news/harness-benchmark.

About Floatboat: Floatboat, developed by AOE Tech Labs, is an AI office environment and agent workstation platform focused on runtime, context persistence, and automated multi-step work.

Floatboat
AOE
+1 302-203-8567
contact@floatboat.ai

Legal Disclaimer:

EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

The Asia Reporter

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.