LGWM · Latent GUI World Model

Right Screen, Wrong TransitionWorld Models as Verifiers for GUI Agents

Jiaming Zhang1, Xuan Wang1, Fuyao Zhang1, Yang Cao2, Lingjuan Lyu3, Wei Yang Bryan Lim1

1Nanyang Technological University · 2Institute of Science Tokyo · 3Sony AI

LGWM is an action-conditioned world model that predicts the next screen in representation space and checks it against the screen that actually appears. One cosine distance verifies every step of a GUI agent, with no judge model and no attack labels.

Left: a generative world model renders its prediction as text, image or code and needs a VLM judge, taking 93 to 124 seconds, while LGWM compares predicted and observed screens in one representation space in 17 ms. Right: the same login screen is a hijack after a tap on View order and safe after a tap on Sign in.
Safety is a property of the transition, not the screen. Left: generative world models render the future and need a judge; LGWM compares predicted and observed screens in one representation space. Right: the same login screen is legitimate after “Sign in” and an attack after “View order”.

Why verify transitions, not screens

A login screen that appears after a tap on “Sign in” is expected. The same screen, pixel for pixel, after a tap on “View order” is an attack. GUI agents act through interfaces built by third parties, and attacks on them redirect what happens after an action while every screen stays plausible. A monitor that inspects only screens can be defeated by reusing a legitimate one.

Judging a transition requires an expectation of what should have followed the action, which is what a world model provides. Existing GUI world models render that expectation as text, code or images, so checking it against the observed screen takes a second model to judge the two.

LGWM predicts in the space where observations are encoded. Verification becomes a vector comparison, and the same residual shows whether a mismatch is harmful. World models have served as simulators and as planners; LGWM gives them a third role, verification.

How LGWM works

A ViT-B encoder initialized from DINOv2 maps a screenshot to 512 patch tokens. An action encoder embeds the action type, its text and its touch coordinates, which share one positional basis with the visual patches. A 12-layer Transformer predictor combines both and outputs the tokens of the next screen. There is no decoder and no autoregressive generation.

Training: a screen encoder and an action encoder feed a latent predictor whose output is regressed onto the next-screen representation from an EMA encoder. Runtime: the predicted and observed representations are compared by cosine similarity for the integrity score, and a linear head on their residual gives the harm score.
One predicted future, two readouts. Training: an EMA encoder supplies the next-screen target, and patches changed by the action receive a higher loss weight. Runtime: the gap between predicted and observed screens gives the integrity score; a linear head on the residual gives the harm score.
Training-free

Integrity score

1 − cos(predicted, observed)

Flags a transition the model would not have produced. It uses no labeled examples, no attack-specific tuning and no parameters beyond the pretrained world model.

One linear layer

Harm score

σ(w · residual + b)

A surprise is not always harmful: a push notification, a layout change. A linear head on the residual, observed − predicted, separates harmful violations from benign ones.

Training is self-supervised on 1.85M real transitions from Android in the Wild, AndroidControl, GUIOdyssey, AMEX and MiniWoB++: 120,000 updates at batch size 512 on four A100 GPUs. The training graph has 261M parameters, of which 172.86M run at inference.

Results on RSWT-Bench

RSWT-Bench (Right Screen, Wrong Transition) is a donor-paired diagnostic. Each of 222 credential screens appears once as a legitimate outcome and once, pixel-identical, as a hijacked one, so any detector that sees only the resulting screen has an expected AUC of exactly 0.5. The test set has 966 transitions, and every method receives the same before screen, action and after screen.

Main results on RSWT-Bench. Detection is pairwise AUC, higher is better. Cost is per decision at batch size one on one A100-80GB.

MethodParamsRSWTSwapCredLatency (ms)TFLOPsVRAM (GB)
Closed-source VLMs
Gemini 3.7 Flash–0.9950.9980.997–––
GPT-5.6 Luna–0.9880.9710.979–––
Qwen3.7 Plus–0.9730.9540.950–––
Qwen3.7 Flash–0.9680.9530.935–––
Qwen3-VL-Plus–0.9180.8810.867–––
Claude Haiku 4.5–0.8800.8420.844–––
Qwen3-VL-Flash–0.7460.8180.669–––
Open VLMs
GLM-4.6V-Flash9B0.9090.8680.8303,711.615.31620.72
Qwen3-VL-8B8B0.9070.9050.8591,097.810.92417.72
Qwen3-VL-4B4B0.9050.8630.8461,075.15.5309.05
Qwen3-VL-2B2B0.6880.7120.605829.82.6514.39
Qwen2.5-VL-7B7B0.5740.5800.543819.411.42716.73
Generative world models
gWorld-8B8B0.8850.8970.837113,200.074.10018.18
SAWM8B0.8820.8360.8041,127.713.44917.77
Code2World7B0.8700.9010.85992,900.0110.71118.56
MobileWorld-HTML8B0.5770.6910.568124,300.0119.32118.64
Latent world model
LGWM (ours)261M0.9870.9700.98517.10.4081.09

RSWT compares hijacked and legitimate occurrences of identical screens. Swap and Cred compare benign mismatches with paired-safe transitions and legitimate credentials. Generative world-model cost excludes the VLM judge. Blind human annotators reach 0.958 RSWT on a 20-pair subset.

Within one AUC point of frontier VLMs, without a judge. The training-free integrity score reaches 0.987 RSWT AUC and orders 219 of the 222 donor pairs correctly. Gemini 3.7 Flash reaches 0.995 and GPT-5.6 Luna 0.988; every open VLM and generative world model scores 0.909 or lower.

Cheap enough to check every step. One decision costs 17.1 ms, 0.408 TFLOPs and 1.09 GB of memory: 64× lower latency and 16× less memory than Qwen3-VL-8B under the same protocol. World models that render the future (gWorld, Code2World, MobileWorld-HTML) need 93 to 124 s before their judge runs.

Surprise detects, direction interprets. A violated expectation is not always harmful. On harmful versus benign violations, the linear head on the residual direction reaches 0.954 AUC, where prompted VLMs score between 0.51 and 0.61.

The verifier is a world model

Four checks rule out degenerate explanations for the score.

Quick start

Score one transition from a before screenshot, an action and an after screenshot.

git clone https://github.com/jiamingzhang94/lgwm && cd lgwm
pip install -c requirements-tested.txt -e '.[text]'
hf download jiamingzz/lgwm --local-dir weights

echo '{"action_type": "tap", "x": 0.5, "y": 0.5}' > action.json
python tools/score_transition.py --before before.png --after after.png \
  --action action.json --harm-head weights/lgwm-harm-head

The command prints integrity and harm_score. The repository also covers evaluation on RSWT-Bench and training from scratch.

Frequently asked questions

What is LGWM?

LGWM (Latent GUI World Model) is an action-conditioned world model for GUI agents. Given the current screenshot and an action, it predicts the representation of the next screen and compares it with the screen that actually appears. One cosine distance tells whether the transition is the one the action should have produced.

What is interaction hijacking, and why do screen-only monitors miss it?

The interface answers an agent's action with a screen that is legitimate in another context, such as a real login page shown after a tap on “View order”. Nothing on the screen is malicious, so a monitor that inspects only the result cannot tell the hijacked case from the legitimate one. The evidence is in the transition.

What is RSWT-Bench?

RSWT-Bench (Right Screen, Wrong Transition) is a donor-paired benchmark with 966 test transitions. Each of 222 credential screens appears under both a legitimate and a hijacked transition, so any detector that sees only the outcome screen scores 0.5 AUC by construction.

How does LGWM differ from generative GUI world models?

Generative world models render the predicted future as text, code or pixels, then need a second model to judge it against the observed screen. LGWM keeps the prediction in representation space, so verification is a single vector comparison that takes 17 ms. Rendering alone takes 93 to 124 seconds in gWorld, Code2World and MobileWorld-HTML.

Does LGWM need attack labels or a judge model?

No. The world model is trained self-supervised on 1.85M real GUI transitions, and the integrity score uses no labels, no attack-specific tuning and no judge. Only the optional harm head, a single linear layer on the prediction residual, is supervised.

Is LGWM open source?

Yes. The code and model weights are released under the Apache-2.0 license on GitHub and Hugging Face, together with the processed training data and the RSWT-Bench manifests.

Citation

@misc{zhang2026lgwm,
  title  = {Right Screen, Wrong Transition:
            World Models as Verifiers for GUI Agents},
  author = {Zhang, Jiaming and Wang, Xuan and Zhang, Fuyao and
            Cao, Yang and Lyu, Lingjuan and Lim, Wei Yang Bryan},
  year   = {2026},
  howpublished = {\url{https://github.com/jiamingzhang94/lgwm}}
}