Accessibility snapshots, not screenshots
The README describes the server as one that lets language models interact with web pages through structured accessibility snapshots, bypassing the need for screenshots or visually-tuned models. That is the design choice that separates it from screenshot-driven browser agents: the model receives a labelled element tree, so it can name a target instead of guessing at pixel coordinates.