The same web page as extracted text and as pixels, then a probe showing the model confidently reporting a control that does not exist.

Now point the loop at a screen

Patches made an image into tokens, and a shared space made those tokens mean something to a language model. Put that in front of an agent and it can work on things that were never text — along with a way of being wrong that text never had.

now ask it what is on the page

is there a … on this page?it saystruth

what that one wrong answer costs

A dashboard usually does have an export button. That is exactly the problem: the answer came from what dashboards are normally like, not from what is in the picture — and pixels give it far more room to fill in the expected than a list of DOM nodes ever did.

Design for a wrong observation. Make it name what it sees before acting, check that the click did what it claimed, and never let “I clicked it” stand in for “it worked”.