A new benchmark designed to measure how well artificial intelligence handles realistic knowledge work has delivered a humbling result: even the best-performing model fully solved just 3 percent of the tasks it was given, according to The Decoder.
The finding cuts against the impression many people have formed from chatbots that can draft emails, summarize documents, and answer questions with apparent ease. When the bar is raised to the kind of multi-step, real-world work that office professionals actually do, The Decoder reports, today's leading models struggle badly.
The 3 percent figure refers to tasks the top model solved completely. That distinction matters: partially completing a task is not the same as delivering finished, reliable work that a person could hand off without checking. Benchmarks like this one are built to test that harder standard rather than rewarding answers that merely look plausible.
The Decoder frames the result as evidence of a gap between the hype around AI handling knowledge work and what the technology can currently accomplish on its own.
Why it matters: As companies weigh whether to lean on AI for substantive professional tasks, this benchmark is a reminder that current systems remain far from reliably doing real knowledge work without human oversight.