Harness engineering, not model intelligence, decided the winner in Jefferies' benchmark of eight US and Chinese AI agents.
Harness engineering, not model intelligence, decided the winner in Jefferies' benchmark of eight US and Chinese AI agents.

Harness engineering, not model intelligence, decided the winner in Jefferies' benchmark of eight US and Chinese AI agents.
Alibaba's Qianwen Office scored 95 points to top the field, beating Anthropic's Claude Cowork at 94 and OpenAI's Codex at 92 despite running a model with a lower intelligence score than its rivals.
"The harness advantage is sufficient to offset model intelligence gaps," Jefferies analysts wrote in the report, which tested agents across five enterprise tasks including multi-file retrieval, autonomous web research, browser control, PPT creation and marketing poster generation.
The test revealed a striking pattern: the same Claude Opus 4.6 model scored between 58.0% and 76.4% on Terminal-Bench 2.0 depending on which harness wrapped it — an 18.4-point spread. Gemini 3 Pro varied by 13.4 points across harnesses. Jefferies decomposed agent capability into 60% model and 40% harness weight, then reverse-engineered each product's implied harness score. Qianwen Office's harness ranked highest, surpassing both Claude Cowork and Codex.
The finding reshapes how enterprises evaluate AI vendors. Model intelligence scores — Qwen 3.8 Max at 56, Claude Opus 5 at 61, GPT-5.6 Sol at 59 — no longer predict which agent product will actually get work done. Qianwen Office's API costs about $1.1 per million tokens versus $3.9 for Opus 5 and $4.4 for GPT-5.6 Sol, making the harness advantage even more compelling on a cost-adjusted basis.
Jefferies broke harness engineering into six layers: instructions, context, tools, boundaries, feedback and governance. Each layer handles what the model cannot manage on its own — defining the task and its constraints, supplying background information, providing the tools to execute, setting the scope of action, enabling self-correction through feedback loops, and allowing organizations to manage fleets of agents.
The framework maps directly to how companies manage employees: brief the task, give information, provide tools, set permissions, offer feedback and supervise. A capable model dropped into a poorly managed environment underperforms, just as a smart hire fails without proper management.
Tencent's Workbuddy presents the most striking anomaly. Its 21 million monthly visits make it the most-used Chinese AI agent, yet its implied harness score ranked last among Chinese products in the Jefferies test. The explanation lies in distribution: Workbuddy is deeply embedded in Tencent's product stack — WeChat Docs, IMA and WeCom — and follows a model-agnostic approach, letting users switch between Kimi K3, DeepSeek V4 and GLM 5.2. Heavy marketing spend compounds the reach.
Jefferies' assessment: short-term, Workbuddy wins on distribution; long-term, its 21 million users will generate the real-world trace data needed to improve its harness and potentially train proprietary models.
The report also found task-specific strengths and weaknesses. US agents stumbled on marketing poster generation — Claude Cowork and Gemini Spark both failed that task — while Chinese agents struggled with browser control, where Workbuddy and MiniMax Code fell short. Speed varied too: Gemini Spark was fastest among US agents, Workbuddy fastest among Chinese.
The report identified four structural advantages for Chinese harness builders: super-app integration (agents embedded directly in WeCom, Feishu and DingTalk), model-agnostic harnesses, token prices averaging 70-80% lower than US equivalents, and faster iteration cycles driven by larger user bases and more failure cases.
Four weaknesses offset these: the model gap persists (weaker models depend more heavily on harness quality), enterprise software monetization remains difficult in China, international expansion faces a software stack built for local platforms, and compute constraints limit the heavy inference workloads agent workflows demand.
For investors, the benchmark suggests the AI agent market's value accrues to companies that master deployment infrastructure, not just model weights. Alibaba's Qianwen Office demonstrates that a mid-tier model wrapped in a superior harness can outperform frontier models on real enterprise tasks at a fraction of the cost. Tencent's Workbuddy shows distribution can temporarily mask harness weakness, but the data flywheel from 21 million users gives it a path to catch up. The companies best positioned are those combining strong harness engineering with the user base to generate improvement data — a dynamic that favors Alibaba and Tencent in China, and Anthropic and OpenAI in the US.
This article is for informational purposes only and does not constitute investment advice.