Why verify transitions, not screens
A login screen that appears after a tap on “Sign in” is expected. The same screen, pixel for pixel, after a tap on “View order” is an attack. GUI agents act through interfaces built by third parties, and attacks on them redirect what happens after an action while every screen stays plausible. A monitor that inspects only screens can be defeated by reusing a legitimate one.
Judging a transition requires an expectation of what should have followed the action, which is what a world model provides. Existing GUI world models render that expectation as text, code or images, so checking it against the observed screen takes a second model to judge the two.
LGWM predicts in the space where observations are encoded. Verification becomes a vector comparison, and the same residual shows whether a mismatch is harmful. World models have served as simulators and as planners; LGWM gives them a third role, verification.

