← 返回阅读笔记 / Back to reading note

Benchmarking and Improving GUI Agents in High-Dynamic Environments (DynamicGUIBench / DynamicUI) 原表

Y26-097 原文表格

抽取日期:2026-09-07。来自原文。保留原表行列,未统一实验配置;数值解释见中文分析。

Table 1. Comparison of existing GUI benchmarks and our benchmark in terms of application diversity, task scale, evaluation setting (Mode), anomaly presence (An.), and dynamic tasks (Dyn.). ✓ and ✗ denote the presence and absence of anomaly or dynamic tasks, respectively.
Benchmark Venue Apps/Web Tasks Mode An. Dyn.
Mind2Web (Deng et al., 2023) NeurIPS’23 137 2,350 Offline
OSWorld (Xie et al., 2024) NeurIPS’24 10 369 Online
GUI Odyssey (Lu et al., 2025) ICCV’25 201 7,735 Offline
WorldGUI (Hengyuan Zhao et al., 2025) arXiv’25 10 315 Offline
Android World (Rawles et al., ) ICLR’25 20 116 Offline
OSWorld-G (Xie et al., 2025) NeurIPS’25 10 564 Online
GUI-Robust (Yang et al., 2025b) NeurIPS’25 392 5,318 Offline
D-GARA (Chen et al., 2025a) AAAI’26 7 152 Online
Ours 2026 10 149 Online
Table 2. Key statistics in DynamicGUIBench.
Statistic Number Statistic Number
Chrome 38 (25.5%) Multi apps 33 (22.4%)
Gimp 10 (6.8%) Os 7 (4.7%)
Libreofficecalc{}_{\text{calc}} 14 (9.5%) Thunderbrid 9 (6.1%)
Libreofficeimpress{}_{\text{impress}} 7 (4.7%) VLC 5 (3.4%)
Libreofficewriter{}_{\text{writer}} 14 (9.5%) VS Code 12 (8.1%)
Feasible 146 (98%) Infeasible 3 (2.0%)
Table 3. Comparison with state-of-the-art methods on the DynamicGUIBench. Abbreviations: Th. (Thunderbird), Multi. (Multi-Apps), Vs. (Visual Studio Code), Imp. (Impress), Wri. (Writer). The best scores are in bold.
Agent Model Step Acc (%)
Chrome Gimp Calc Imp. Wri. Multi. Os Th. Vlc Vs. Avg.
o3 (OpenAI, 2025) 15 7.9 10.0 7.1 0.0 7.1 6.1 0.0 0.0 20.0 8.3 6.7
Qwen3-vl-4B (Bai et al., 2025) 15 18.4 10.0 7.1 0.0 7.1 9.1 14.3 0.0 40.0 0.0 10.7
Qwen3-vl-8B (Bai et al., 2025) 15 21.1 20.0 14.3 0.0 14.3 9.1 14.3 0.0 40.0 16.7 14.8
o3 (OpenAI, 2025) 50 10.5 0.0 14.3 0.0 7.1 6.1 42.9 0.0 20.0 0.0 8.7
UITARS-1.5-7B (Qin et al., 2025b) 50 13.2 10.0 7.1 0.0 7.1 9.1 0.0 0.0 20.0 0.0 8.1
Qwen3-vl-4B (Bai et al., 2025) 50 13.2 10.0 7.1 0.0 7.1 9.1 28.6 0.0 60.0 0.0 10.7
Qwen3-vl-8B (Bai et al., 2025) 50 26.3 22.0 14.3 0.0 7.1 7.5 14.3 0.0 60.0 8.3 15.1
doubao-1-5-0717(Guo et al., 2025) 50 5.3 0.0 7.1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 2.8
Seed1.8-VL (Seed, ) 50 10.5 10.0 0.0 14.3 7.1 3.0 0.0 0.0 0.0 16.7 6.7
EvoCUA-8B (Xue et al., 2026) 50 18.4 33.0 0.0 0.0 7.1 4.5 28.6 0.0 20.0 16.7 11.7
Ours w/ Qwen3-vl-8B 15 29.0 10.0 14.3 0.0 7.1 12.2 14.3 11.1 0.0 16.7 15.5
Ours w/ Qwen3-vl-8B 50 36.8 30.0 14.3 14.3 7.1 9.1 28.6 44.4 40.0 8.3 22.1
Table 4. Ablation on the DynamicGUIBench.
Dynamic Perceiver Reflection Module Refinement Strategy Acc(%)
- - - 15.1
- 17.4
- 17.4
- 20.8
22.1
Table 5. Comparison with state-of-the-art methods on the OSWorld ( Xie et al., 2024 ) benchmark. We use ‘*’ to denote the results evaluated by us (which will be updated if improved evaluation scripts become available).
Agent Model Step Acc (%)
Chrome Gimp Calc Imp. Wri. Multi. Os Th. Vlc Vs. Avg.
o3 (OpenAI, 2025) 50 21.7 38.5 8.5 4.3 21.7 11.8 37.5 20.0 17.6 21.7 17.2
Opencua-a3b (Wang et al., 2025a) 50 21.6 53.9 2.1 23.3 34.8 6.5 37.5 6.7 11.8 43.5 19.9
UITARS-72b-dpo (Qin et al., 2025b) 50 33.2 61.5 12.8 25.5 43.5 6.7 33.3 33.3 23.5 47.8 25.8
UITARS-1.5-7B (Qin et al., 2025b) 50 32.8 53.9 8.5 39.3 41.3 8.6 25.6 46.7 24.4 52.2 27.2
Qwen3-vl-4B* (Bai et al., 2025) 50 30.4 21.7 12.8 21.8 39.1 5.2 33.3 46.7 23.0 13.0 19.8
Opencua-7B (Wang et al., 2025a) 50 38.6 43.6 13.2 32.6 33.3 12.1 43.5 42.2 28.3 47.1 28.2
Qwen3-vl-8B* (Bai et al., 2025) 50 23.9 34.8 17.0 31.8 56.5 9.2 33.3 66.7 40.4 21.7 25.8
Ours w/ UITARS-1.5-7B 50 32.8 57.7 8.5 39.3 41.3 8.6 29.8 50.0 27.5 54.4 28.2
Ours w/ Qwen3-vl-8B 50 34.8 69.6 17.0 19.6 26.1 13.1 62.5 40.0 17.7 43.5 28.4
Table 6. Comparison of uniform frame sampling baselines and our Dynamic Perceiver (DP) across four POMDP categories on DynamicGUIBench.
Method Frames Interruptive UI Ephemeral Ref. DynList ContentTrig Avg.
Qwen3-vl-8B (Bai et al., 2025) - 25.0 12.3 13.6 19.2 15.1
Uniform 1 25.0 15.3 31.8 3.9 16.8
3 25.0 17.0 13.6 11.5 16.4
Ours DP 43.8 20.0 22.7 15.4 22.1
Table 7. Comparison of different general-purpose models for the reflection module on DynamicGUIBench.
Method Interruptive UI Ephemeral Ref. DynList ContentTrig Avg.
Qwen3-vl-8B (Bai et al., 2025) 25.0 9.3 9.1 11.5 11.3
GPT-5.4-mini (Singh et al., 2025) 43.8 20.0 22.7 15.4 22.1