The same web page as extracted text and as pixels, then a probe showing the model confidently reporting a control that does not exist.
Now point the loop at a screen
Patches made an image into tokens, and a shared space made those tokens mean something to a language model. Put that in front of an agent and it can work on things that were never text — along with a way of being wrong that text never had.
now ask it what is on the page
is there a … on this page?it saystruth
what that one wrong answer costs
A dashboard usually does have an export button. That is exactly the problem: the answer came from what dashboards are normally like, not from what is in the picture — and pixels give it far more room to fill in the expected than a list of DOM nodes ever did.
Design for a wrong observation. Make it name what it sees before acting, check that the click did what it claimed, and never let “I clicked it” stand in for “it worked”.