Huawei's New Benchmark Gives AI Agents Months of Your Life—Then Watches Them Fail

DECRYPT ·

Researchers from Huawei and three partner institutions released Claw-Anything, a benchmark that evaluates AI agents on personal-assistant tasks. GPT-5.5, OpenAI's flagship model, scored only 34.5% on the pass@1 metric—far below its scores on existing benchmarks, suggesting current tests are measuring the wrong things. The team also released an automated data pipeline that produced 2,000 training environments; fine-tuning an open-weight model on that data improved task success by 23.7%. The pitch for AI personal assistants has always been the same: Give the agent access to your digital life and it handles the rest. Your emails, your calendar, your notes, your devices—all of it. Your AI knows. Your AI acts. You sleep. Researchers from Huawei Technologies, Beijing Institute of Technology, Peking University, and the Chinese Academy of Sciences just built a benchmark to see if that's actually true. Spoiler: It's not. Claw-Anything evaluates AI agents across three dimensions at once: long-horizon event streams covering more than three months of simulated user activity, interdependent backend services averaging 10.1 per task, and multi-device interaction across both CLI Linux environments and GUI Android environments.

AI 시장 분석

Huawei has introduced a new benchmark designed to evaluate the performance of AI agents. This benchmark reveals that current AI agents still face significant limitations in handling complex, real-world scenarios, suggesting they are not yet capable of fully mimicking or solving intricate human-like situations.

상승 영향

하락 영향

AI가 생성한 분석으로 투자 자문이 아닙니다.

DYAX Investor Sentiment

Bullish (Long) 60% · Bearish (Short) 40%

336 participants

Related News

원문 보기 — DECRYPT