Table 1: Comparison of test-time interventions on Qwen3-VL-32B and Gemini-3-Flash across GUI benchmarks in diverse environments. The best and the second best overall success rate results are in bold and underline , respectively. HiViG delivers gains in all benchmarks. WALv2: WebArenaLitev2 for web, ALab: AndroidaLab for mobile, WAA: WindowsAgentArena for desktop, and Avg.: average overall success rate. Detailed performances for each benchmark are in Table 6 , Table 7 , and Table 8 .
Method
Qwen3-VL-32B-Thinking
Gemini-3-Flash
WALv2
ALab
WAA
Avg.
WALv2
ALab
WAA
Avg.
Base agent
13.0
44.2
35.7
31.0
30.5
58.0
35.8
41.4
OpenCUA
14.9
47.1
35.2
32.4
29.9
57.2
35.8
41.0
SE-WSM
11.7
45.7
31.6
29.7
30.5
57.2
35.1
40.9
CGI (32B)
16.9
46.9
33.7
32.5
29.9
55.8
37.9
41.2
CGI (8B)
12.3
44.2
31.0
29.2
26.6
55.1
37.9
39.9
GUI-Critic-R1
13.6
49.3
34.4
32.4
22.1
53.6
32.5
36.1
HiViG (Ours)
25.3
51.5
38.0
38.3
45.5
61.6
44.2
50.4
Table 2: Impact of verbal feedback components. Experiments with verbal feedback components of HiViG on WebArenaLitev2 evaluated with Qwen3-VL-32B and Gemini-3-Flash as CUA. VisAnalysis: Visually-grounded error analysis, HisAnalysis: History state tracking. Combining two components leads to better performance.
VisAnalysis
HisTrack
Qwen3-VL
G-3-Flash
✗
✗
13.5
30.0
✓
✗
21.4
35.1
✗
✓
23.4
42.9
✓
✓
25.3
45.5
Table 3: Ablation of visual grounding strategies. We evaluate the impact of masking the policy’s verbal intent (30% of SFT samples) and injecting a visual marker (i.e., ’X’ marker). Combining both strategies forces the critic to break its text reliance and evaluate the raw spatial coordinates, yielding visually accurate verbal feedback. WALv2: WebArenaLitev2, ALab: AndroidLab.
Intent Masking
Visual Marker
WALv2
ALab
✓
✗
20.8
46.4
✗
✓
25.3
47.1
✓
✓
25.3
51.5
Table 4: Ablation on critic training. We compare the overall success rate of a unified 8B critic trained on a mixed dataset (Mixed training HiViG ) against separately trained 8B critics (Separate HiViG ) and zero-shot baselines that use Qwen3-VL models for historical progress grounding across WebArenaLitev2. VisAnalysis: Visually grounded error analysis, HisTrack: Historical state tracking.
Both Qwen3-VL-8B and Qwen3-VL-32B are thinking models.
VisAnalysis
HisTrack
Overall
Separate HiViG
Separate HiViG
23.4
Separate HiViG
Qwen3-VL-8B
24.7
Separate HiViG
Qwen3-VL-32B
25.3
Mixed training HiViG (Ours)
25.3
Table 5: Ablation on intent masking. We evaluate the impact of the intent masking ratio for visual grounding. While a 50% mask yields slight improvements on WebArenaLitev2, the 30% ratio provides the most robust and balanced performance across platforms. WALv2: WebArenaLitev2, ALab: AndroidLab. Avg.: average overall success rate.
Intent Masking Ratio
WALv2
ALab
Avg.
0%
25.3
47.1
36.2
30%
25.3
51.5
38.4
50%
26.0
49.3
37.7
Table 6: Comparison of test-time intervention methods on Qwen3-VL-32B-Thinking and Gemini-3-Flash in WebArenalitev2 benchmark.
Method
Admin
GitLab
Shopping
Map
Reddit
Overall
Computer Use Agent: Qwen3-VL-32B-Thinking
Base agent
11.4
16.7
15.9
3.9
15.8
13.0
OpenCUA
14.3
16.7
20.5
3.9
15.8
14.9
SE-WSM
11.4
16.7
15.9
3.9
5.3
11.7
CGI (32B)
17.1
16.7
25.0
3.9
15.8
16.9
CGI (8B)
14.3
13.3
13.6
3.9
15.8
12.3
GUI-Critic-R1
14.3
13.3
15.9
3.9
21.1
13.6
HiViG (Ours)
28.6
13.3
31.8
23.1
26.3
25.3
Computer Use Agent: Gemini-3-Flash
Base agent
25.7
56.7
25.0
19.2
26.3
30.5
OpenCUA
22.9
53.3
22.7
23.1
31.6
29.9
SE-WSM
25.7
53.3
31.8
23.1
10.5
30.5
CGI (32B)
37.1
33.3
31.8
19.2
21.1
29.9
CGI (8B)
17.1
50.0
25.0
11.5
31.6
26.6
GUI-Critic-R1
8.6
36.7
27.3
15.4
21.1
22.1
HiViG (Ours)
42.9
70.0
38.6
34.6
42.1
45.5
Table 7: Comparison of test-time intervention methods on Qwen3-VL-32B-Thinking and Gemini-3-Flash in AndroidLab benchmark.
Method
Bluecoins
Calendar
Cantook
Clock
Contacts
Map
Pimusic
Setting
Zoom
Overall
Computer Use Agent: Qwen3-VL-32B-Thinking
Base agent
20.0
35.7
33.3
74.1
53.3
6.7
16.7
69.6
40.0
44.2
OpenCUA
33.3
28.6
41.7
70.4
53.3
6.7
25.0
73.9
60.0
47.1
SE-WSM
20.0
42.9
41.7
70.4
53.3
13.3
8.3
69.6
60.0
45.7
CGI (32B)
26.7
42.9
33.3
74.1
53.3
6.7
33.3
76.5
0.0
46.9
CGI (8B)
26.7
42.9
25.0
77.8
53.3
0.0
16.7
69.6
20.0
44.2
GUI-Critic-R1
40.0
42.9
16.7
77.9
53.3
13.3
25.0
73.9
60.0
49.3
HiViG (Ours)
40.0
42.9
50.0
70.4
60.0
13.3
16.7
78.3
60.0
51.5
Computer Use Agent: Gemini-3-Flash
Base agent
66.7
42.9
75.0
74.1
53.3
13.3
25.0
78.3
80.0
58.0
OpenCUA
60.0
42.9
50.0
81.5
26.7
40.0
33.3
78.3
80.0
57.2
SE-WSM
53.3
42.9
58.3
77.9
40.0
40.0
33.3
73.9
80.0
57.2
CGI (32B)
60.0
42.9
50.0
74.1
33.3
40.0
33.3
78.3
60.0
55.8
CGI (8B)
60.0
35.7
58.3
88.9
26.7
26.7
33.3
73.9
40.0
55.1
GUI-Critic-R1
60.0
42.9
50.0
66.7
46.7
26.7
25.0
73.9
80.0
53.6
HiViG (Ours)
73.3
42.9
50.0
88.9
40.0
40.0
33.3
78.3
80.0
61.6
Table 8: Comparison of test-time intervention methods on Qwen3-VL-32B-Thinking and Gemini-3-Flash in WindowsAgentArena benchmark.