← 返回阅读笔记 / Back to reading note

A History-Aware Visually Grounded Critic for Computer Use Agents (HiViG) 原表

Y26-133 原文表格

抽取日期:2026-09-07。来自原文。保留原表行列,未统一实验配置;数值解释见中文分析。

Table 1: Comparison of test-time interventions on Qwen3-VL-32B and Gemini-3-Flash across GUI benchmarks in diverse environments. The best and the second best overall success rate results are in bold and underline , respectively. HiViG delivers gains in all benchmarks. WALv2: WebArenaLitev2 for web, ALab: AndroidaLab for mobile, WAA: WindowsAgentArena for desktop, and Avg.: average overall success rate. Detailed performances for each benchmark are in Table 6 , Table 7 , and Table 8 .
Method Qwen3-VL-32B-Thinking   Gemini-3-Flash
WALv2 ALab WAA Avg. WALv2 ALab WAA Avg.
Base agent 13.0 44.2 35.7 31.0 30.5 58.0 35.8 41.4
OpenCUA 14.9 47.1 35.2 32.4 29.9 57.2 35.8 41.0
SE-WSM 11.7 45.7 31.6 29.7 30.5 57.2 35.1 40.9
CGI (32B) 16.9 46.9 33.7 32.5 29.9 55.8 37.9 41.2
CGI (8B) 12.3 44.2 31.0 29.2 26.6 55.1 37.9 39.9
GUI-Critic-R1 13.6 49.3 34.4 32.4 22.1 53.6 32.5 36.1
HiViG (Ours) 25.3 51.5 38.0 38.3 45.5 61.6 44.2 50.4
Table 2: Impact of verbal feedback components. Experiments with verbal feedback components of HiViG on WebArenaLitev2 evaluated with Qwen3-VL-32B and Gemini-3-Flash as CUA. VisAnalysis: Visually-grounded error analysis, HisAnalysis: History state tracking. Combining two components leads to better performance.
VisAnalysis HisTrack Qwen3-VL G-3-Flash
13.5 30.0
21.4 35.1
23.4 42.9
25.3 45.5
Table 3: Ablation of visual grounding strategies. We evaluate the impact of masking the policy’s verbal intent (30% of SFT samples) and injecting a visual marker (i.e., ’X’ marker). Combining both strategies forces the critic to break its text reliance and evaluate the raw spatial coordinates, yielding visually accurate verbal feedback. WALv2: WebArenaLitev2, ALab: AndroidLab.
Intent Masking Visual Marker WALv2 ALab
20.8 46.4
25.3 47.1
25.3 51.5
Table 4: Ablation on critic training. We compare the overall success rate of a unified 8B critic trained on a mixed dataset (Mixed training HiViG ) against separately trained 8B critics (Separate HiViG ) and zero-shot baselines that use Qwen3-VL models for historical progress grounding across WebArenaLitev2. VisAnalysis: Visually grounded error analysis, HisTrack: Historical state tracking. Both Qwen3-VL-8B and Qwen3-VL-32B are thinking models.
VisAnalysis HisTrack Overall
Separate HiViG Separate HiViG 23.4
Separate HiViG Qwen3-VL-8B 24.7
Separate HiViG Qwen3-VL-32B 25.3
Mixed training HiViG (Ours)   25.3
Table 5: Ablation on intent masking. We evaluate the impact of the intent masking ratio for visual grounding. While a 50% mask yields slight improvements on WebArenaLitev2, the 30% ratio provides the most robust and balanced performance across platforms. WALv2: WebArenaLitev2, ALab: AndroidLab. Avg.: average overall success rate.
Intent Masking Ratio WALv2 ALab Avg.
0% 25.3 47.1 36.2
30% 25.3 51.5 38.4
50% 26.0 49.3 37.7
Table 6: Comparison of test-time intervention methods on Qwen3-VL-32B-Thinking and Gemini-3-Flash in WebArenalitev2 benchmark.
Method Admin GitLab Shopping Map Reddit Overall
Computer Use Agent: Qwen3-VL-32B-Thinking
Base agent 11.4 16.7 15.9 3.9 15.8 13.0
OpenCUA 14.3 16.7 20.5 3.9 15.8 14.9
SE-WSM 11.4 16.7 15.9 3.9 5.3 11.7
CGI (32B) 17.1 16.7 25.0 3.9 15.8 16.9
CGI (8B) 14.3 13.3 13.6 3.9 15.8 12.3
GUI-Critic-R1 14.3 13.3 15.9 3.9 21.1 13.6
HiViG (Ours) 28.6 13.3 31.8 23.1 26.3 25.3
Computer Use Agent: Gemini-3-Flash
Base agent 25.7 56.7 25.0 19.2 26.3 30.5
OpenCUA 22.9 53.3 22.7 23.1 31.6 29.9
SE-WSM 25.7 53.3 31.8 23.1 10.5 30.5
CGI (32B) 37.1 33.3 31.8 19.2 21.1 29.9
CGI (8B) 17.1 50.0 25.0 11.5 31.6 26.6
GUI-Critic-R1 8.6 36.7 27.3 15.4 21.1 22.1
HiViG (Ours) 42.9 70.0 38.6 34.6 42.1 45.5
Table 7: Comparison of test-time intervention methods on Qwen3-VL-32B-Thinking and Gemini-3-Flash in AndroidLab benchmark.
Method Bluecoins Calendar Cantook Clock Contacts Map Pimusic Setting Zoom Overall
Computer Use Agent: Qwen3-VL-32B-Thinking
Base agent 20.0 35.7 33.3 74.1 53.3 6.7 16.7 69.6 40.0 44.2
OpenCUA 33.3 28.6 41.7 70.4 53.3 6.7 25.0 73.9 60.0 47.1
SE-WSM 20.0 42.9 41.7 70.4 53.3 13.3 8.3 69.6 60.0 45.7
CGI (32B) 26.7 42.9 33.3 74.1 53.3 6.7 33.3 76.5 0.0 46.9
CGI (8B) 26.7 42.9 25.0 77.8 53.3 0.0 16.7 69.6 20.0 44.2
GUI-Critic-R1 40.0 42.9 16.7 77.9 53.3 13.3 25.0 73.9 60.0 49.3
HiViG (Ours) 40.0 42.9 50.0 70.4 60.0 13.3 16.7 78.3 60.0 51.5
Computer Use Agent: Gemini-3-Flash
Base agent 66.7 42.9 75.0 74.1 53.3 13.3 25.0 78.3 80.0 58.0
OpenCUA 60.0 42.9 50.0 81.5 26.7 40.0 33.3 78.3 80.0 57.2
SE-WSM 53.3 42.9 58.3 77.9 40.0 40.0 33.3 73.9 80.0 57.2
CGI (32B) 60.0 42.9 50.0 74.1 33.3 40.0 33.3 78.3 60.0 55.8
CGI (8B) 60.0 35.7 58.3 88.9 26.7 26.7 33.3 73.9 40.0 55.1
GUI-Critic-R1 60.0 42.9 50.0 66.7 46.7 26.7 25.0 73.9 80.0 53.6
HiViG (Ours) 73.3 42.9 50.0 88.9 40.0 40.0 33.3 78.3 80.0 61.6
Table 8: Comparison of test-time intervention methods on Qwen3-VL-32B-Thinking and Gemini-3-Flash in WindowsAgentArena benchmark.
Method Office Web Browsing Windows System Code Media Windows Utilities Overall
Computer Use Agent: Qwen3-VL-32B-Thnking
Base agent 4.7 57.1 50.0 58.3 27.7 50.0 35.7
OpenCUA 2.3 61.9 70.8 50.0 24.5 25.0 35.2
SE-WSM 2.3 52.4 58.3 50.0 22.9 25.0 31.6
CGI (32B) 4.7 47.6 58.3 50.0 37.2 25.0 33.7
CGI (8B) 4.7 47.6 58.3 41.7 28.1 25.0 31.0
GUI-Critic-R1 4.7 61.9 45.8 54.2 28.1 41.7 34.4
HiViG (Ours) 7.0 57.1 58.3 58.3 28.9 50.0 38.0
Computer Use Agent: Gemini-3-Flash
Base agent 4.7 52.4 58.3 45.8 32.9 58.3 35.8
OpenCUA 4.7 52.4 58.3 45.8 32.9 58.3 35.8
SE-WSM 4.7 42.9 62.5 45.8 28.1 66.7 35.1
CGI (32B) 4.7 47.6 62.5 58.3 32.9 50.0 37.9
CGI (8B) 4.7 57.1 62.5 54.2 37.7 41.7 37.9
GUI-Critic-R1 2.3 57.1 58.3 45.8 19.3 41.7 32.5
HiViG (Ours) 23.3 57.1 62.5 54.2 28.9 66.7 44.2