An AI agent that shops, books tickets or fills in forms for you will usually tell you when it is done. A new test shows that this report is often wrong. In one setup, the agent said it had finished every task, but it really completed only 59 percent of them.

The study, called WebPageBench, was posted on arXiv on September 28. It matters because more companies now let agents act on real websites. If the agent's own report cannot be trusted, someone has to check its work.

How the test works

The researchers built six fake websites: a marketplace, a bookstore, a grocery service, a train ticket site, a hotel search and a document cabinet. Each site records every click and every change as an event.

So the test does not ask the agent whether it succeeded, and it does not just look at the final page. It checks the event log to see what really happened.

A 41-point gap

The results show agents often overstate their success. In the worst case the authors report, the gap between claimed and real success was about 41 percentage points.

The team also changed small parts of the websites' design while keeping the task the same. That shows how easily an agent is thrown off by a different layout. The authors, Anton Emelyanov and four colleagues, say this is why agents need to be measured by what they did, not by what they say.