← 返回阅读笔记 / Back to reading note

Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents 原表

Y26-136 原文表格

抽取日期:2026-09-07。来自原文。保留原表行列,未统一实验配置;数值解释见中文分析。

Table 1 : Per-domain success rate (%) and total token usage per task (K) on WebArena. Domain values are means over three independent runs (mean ± \pm std ); per-run details are in Appendix G . Avg. is the task-weighted mean over tasks of all four domains.
Shopping (187) Reddit (106) Admin (182) GitLab (180) Avg.
Method SR \uparrow Tok \downarrow SR \uparrow Tok \downarrow SR \uparrow Tok \downarrow SR \uparrow Tok \downarrow SR \uparrow Tok \downarrow
Gemini 3 Flash
Vanilla-IB 47.77±\pm1.9 45.7±\pm0.8 47.48±\pm0.5 71.7±\pm2.2 55.68±\pm1.6 99.0±\pm4.0 29.07±\pm0.6 78.0±\pm1.8 44.78 73.6
AWM 41.18±\pm2.1 70.8±\pm3.0 47.48±\pm1.4 87.2±\pm2.0 47.43±\pm0.6 142.7±\pm5.4 24.44±\pm1.0 92.3±\pm1.1 39.34 99.3
ASI 44.74±\pm1.2 82.8±\pm0.3 44.97±\pm2.0 94.6±\pm4.8 52.75±\pm1.1 139.4±\pm11.2 22.96±\pm2.0 107.7±\pm5.8 41.02 107.3
RBank 45.45±\pm1.4 54.6±\pm1.3 39.93±\pm1.4 76.8±\pm2.8 48.90±\pm1.9 124.6±\pm3.3 22.96±\pm0.9 72.7±\pm1.0 39.33 82.6
GPT-5.4-mini
Vanilla-IB 38.50±\pm0.9 54.3±\pm0.8 34.59±\pm4.8 96.2±\pm6.6 35.89±\pm1.3 124.1±\pm8.0 22.22±\pm2.2 89.8±\pm4.4 32.67 90.2
AWM 29.95±\pm1.9 62.0±\pm2.0 29.56±\pm1.1 80.4±\pm12.6 32.24±\pm3.2 113.4±\pm16.8 17.22±\pm1.1 95.7±\pm2.7 27.02 88.5
ASI 33.15±\pm0.9 73.4±\pm9.0 28.93±\pm2.2 87.2±\pm4.9 32.96±\pm3.8 138.5±\pm10.9 20.74±\pm0.6 95.7±\pm1.9 29.00 99.8
RBank 26.38±\pm3.9 57.9±\pm1.6 22.01±\pm2.9 69.2±\pm2.9 34.25±\pm1.8 123.1±\pm2.3 14.45±\pm2.9 69.3±\pm2.2 24.58 81.0
Qwen 3.6-27B
Vanilla-IB 45.99±\pm1.4 58.8±\pm1.5 44.34±\pm2.5 89.7±\pm4.3 50.73±\pm1.4 130.3±\pm9.5 28.15±\pm1.2 101.8±\pm6.5 42.14 95.5
AWM 41.35±\pm1.9 71.8±\pm4.5 43.08±\pm3.0 103.6±\pm13.8 46.15±\pm3.1 178.5±\pm12.5 23.33±\pm1.1 102.2±\pm4.8 38.01 115.0
ASI 43.67±\pm3.5 78.8±\pm5.5 43.08±\pm1.4 120.5±\pm4.7 49.08±\pm2.3 182.2±\pm4.2 25.74±\pm1.4 120.6±\pm3.6 40.15 125.8
RBank 39.75±\pm0.8 65.4±\pm0.2 41.19±\pm1.4 99.1±\pm5.4 47.62±\pm1.7 156.9±\pm8.6 20.93±\pm1.7 86.0±\pm5.7 37.00 101.9
Table 3 : Token breakdown by component under Gemini 3 Flash on WebArena , per task in thousands (K). Actor: tokens from LLM calls made during interaction steps. Modules: tokens from auxiliary components (workflow induction, skill synthesis, retrieval, verification). Both are further split into prompt ( P. ) and completion ( C. ) tokens. Values are means over three runs. Vanilla-IB has no module calls by construction.
Shopping Reddit Admin GitLab
Actor Modules Actor Modules Actor Modules Actor Modules
Method P. C. P. C. P. C. P. C. P. C. P. C. P. C. P. C.
Vanilla-IB 45.2 0.5 0 0 71.3 0.4 0 0 98.4 0.6 0 0 77.3 0.7 0 0
AWM 64.1 0.4 6.1 0.2 82.9 0.4 3.7 0.2 135.0 0.5 6.9 0.3 86.1 0.6 5.4 0.2
ASI 61.8 0.5 20.0 0.5 84.6 0.4 9.3 0.3 118.1 0.5 20.2 0.6 85.2 0.5 21.4 0.5
RBank 36.6 1.1 16.5 0.4 53.3 1.0 22.1 0.4 96.9 1.2 26.1 0.4 51.3 1.8 19.2 0.4
Table 4 : Any-of-3 and All-of-3 success rates (%) on the Shopping domain of WebArena across three models. Any-of-3 is the fraction of tasks succeeding in at least one run; All-of-3 is the fraction succeeding in all three runs. Δ ↑ \Delta_{\uparrow} and Δ ↓ \Delta_{\downarrow} are the gaps from the three-run mean (Table 1 ).
Gemini 3 Flash GPT-5.4-mini Qwen 3.6-27B
Method Any 𝚫\bm{\Delta_{\uparrow}} All 𝚫\bm{\Delta_{\downarrow}} Any 𝚫\bm{\Delta_{\uparrow}} All 𝚫\bm{\Delta_{\downarrow}} Any 𝚫\bm{\Delta_{\uparrow}} All 𝚫\bm{\Delta_{\downarrow}}
Vanilla-IB 54.01 +6.2 42.25 -5.5 46.52 +8.0 27.81 -10.7 54.01 +8.0 38.50 -7.5
AWM 47.06 +5.9 35.83 -5.4 37.43 +7.5 21.39 -8.6 49.73 +8.4 33.16 -8.2
ASI 51.34 +6.6 37.43 -7.3 40.64 +7.5 26.74 -6.4 52.41 +8.7 36.36 -7.3
RBank 53.48 +8.0 38.50 -7.0 35.29 +8.9 17.64 -8.7 47.06 +7.3 34.22 -5.5
Table 6: Augmented baselines with and without Vanilla-IB ’s pruning procedure on the Shopping domain of WebArena, under Gemini 3 Flash and Qwen 3.6-27B. Each baseline is shown in its released configuration and with the same rule-based accessibility-tree pruning used by Vanilla-IB ( + pruning ). Values are means over three runs (mean ± \pm std ). Success rate (%) and total token usage per task (K).
Gemini 3 Flash Qwen 3.6-27B
Method SR (%) \uparrow Tok (K) \downarrow SR (%) \uparrow Tok (K) \downarrow
AWM 41.18±\pm2.1 70.8±\pm3.0 41.35±\pm1.9 71.8±\pm4.5
   + pruning 43.32±\pm0.5 66.9±\pm4.3 39.04±\pm2.3 68.9±\pm2.7
ASI 44.74±\pm1.2 82.8±\pm0.3 43.67±\pm3.5 78.8±\pm5.5
   + pruning 45.63±\pm0.8 71.6±\pm5.6 45.81±\pm2.4 76.7±\pm6.3
RBank 45.45±\pm1.4 54.6±\pm1.3 39.75±\pm0.8 65.4±\pm0.2
   + pruning 42.25±\pm4.2 51.1±\pm0.2 39.21±\pm2.2 59.3±\pm1.0
Vanilla-IB 47.77±\pm1.9 45.7±\pm0.8 45.99±\pm1.4 58.8±\pm1.5
Table 7: Success rate (%) and total token usage per task (K) at a higher step budget , under Qwen 3.6-27B on WebArena. The augmented methods use a 15-step actor horizon and Vanilla-IB a 20-step horizon. Values are means over three runs (mean ± \pm std ).
Shopping Reddit Admin
Method SR (%) \uparrow Tok (K) \downarrow SR (%) \uparrow Tok (K) \downarrow SR (%) \uparrow Tok (K) \downarrow
Vanilla-IB - 20 steps 46.52±\pm0.9 60.8±\pm4.1 47.17±\pm3.8 91.6±\pm3.3 51.65±\pm2.9 161.4±\pm22.4
AWM - 15 steps 41.89±\pm3.4 79.5±\pm3.8 44.03±\pm0.5 105.0±\pm1.4 47.43±\pm4.2 212.6±\pm6.9
ASI - 15 steps 45.63±\pm2.9 96.4±\pm1.4 43.08±\pm1.1 127.4±\pm13.7 51.83±\pm3.0 206.8±\pm11.3
RBank - 15 steps 42.24±\pm2.8 79.1±\pm1.4 41.83±\pm2.7 111.6±\pm5.4 50.18±\pm1.1 198.0±\pm24.2
Table 8 : ASI verification failures at the first high-level function call on WebArena. Values are averages over three runs. Attempts : total number of function induction attempts by ASI across all tasks. First-step fail : number of attempts where the induced function failed before the second agent action. Recov. rate : fraction of first-step failures where the verification episode was nonetheless judged correct, causing a potentially broken function to be stored in the shared library.
Domain Attempts First-step fail Recov. rate
Gemini 3 Flash
Shopping 58.0 19.3 69.0%
Reddit 30.0 9.0 59.3%
Admin 64.3 6.3 63.2%
GPT-5.4-mini
Shopping 44.3 32.0 54.2%
Reddit 25.0 12.3 45.9%
Admin 51.7 22.7 47.1%
Table 9 : AWM workflow-induction statistics on WebArena, averaged per run. Induced : total induction events triggered by the judge for a run. From failed (%) : fraction originating from tasks the ground-truth evaluator classified as failed. Final WFs : distinct workflows in the final library after duplicate suppression.
Domain Induced From failed (%) Final WFs
Gemini 3 Flash
Shopping 105.7 49.5% 37.3
Reddit 45.7 10.9% 16.7
Admin 99.7 42.1% 48.3
GPT-5.4-mini
Shopping 78.3 52.3% 33.7
Reddit 31.3 19.1% 13.7
Admin 50.0 42.0% 24.3
Table 10 : ReasoningBank memory-corpus statistics on WebArena, averaged over three runs. Tasks processed per run: 187 (Shopping), 106 (Reddit), 182 (Admin). Succ.-labeled : tasks labeled successful by the judge. FP in succ. : fraction of success-labeled tasks from trajectories that failed according to the ground-truth evaluator. FN in fail : fraction of fail-labeled tasks from trajectories that succeeded according to the ground-truth evaluator.
Domain Total mem. Succ.- labeled FP in succ. (%) FN in fail (%)
Gemini 3 Flash
Shopping 518 96 52.9% 40.7%
Reddit 294 32 30.2% 25.3%
Admin 507 86 42.6% 40.4%
GPT-5.4-mini
Shopping 544 67 59.5% 18.4%
Reddit 307 20 60.0% 17.4%
Admin 532 62 48.1% 25.2%
Table 11: Example failures of augmented agents where Vanilla-IB succeeds. Tasks are selected from WebArena (single runs shown; all from Gemini 3 Flash).
Method Intent Augmented agent behavior Vanilla-IB behavior
AWM “Buy the highest-rated product in Ceiling Light within a budget above $1000.” Retrieved workflow for “Buy the highest-rated product” prescribes sorting then clicking the top result. Agent buys the wrong product because the workflow does not account for rating ties or the budget threshold. Applies sorting, checks rating and price against the budget, and selects the correct product.
ASI “Change the delivery address for my most recent order to 3 Oxford St, Cambridge, MA.” Calls navigate_to_contact_page and fill_contact_form using the element ID from the docstring example (’1381’). Receives ValueError and TypeError; form submission fails. Navigates to Address Book, adds the new address, recognizes that the order address cannot be changed through the UI, and reports accordingly.
RBank “What are the top-5 best-selling products in 2023?” Retrieved memory on handling ranking ties directs attention to tie-breaking. Agent misidentifies the 4th and 5th items (reports Sparta Gym Tank and Angel Light Running Short instead of Sprite Stasis Ball and Hawkeye Yoga Short). Reads the same report without distraction and returns the correct ranked list.
Table 12: Per-run success rate (%) and token usage per task (K) across three independent runs under Gemini 3 Flash on WebArena.
Success Rate (%) \uparrow Tok/task (K) \downarrow
Method Run a Run b Run c Mean±\pmStd Run a Run b Run c Mean±\pmStd
Shopping
Vanilla-IB 47.59 49.73 45.99 47.77±\pm1.9 46.6 45.5 45.1 45.7±\pm0.8
AWM 43.32 39.04 41.18 41.18±\pm2.1 68.0 70.6 73.9 70.8±\pm3.0
ASI 43.32 45.45 45.45 44.74±\pm1.2 82.4 83.0 83.0 82.8±\pm0.3
RBank 44.39 44.91 47.05 45.45±\pm1.4 56.0 54.1 53.6 54.6±\pm1.3
Reddit
Vanilla-IB 48.11 47.17 47.17 47.48±\pm0.5 71.7 69.5 73.8 71.7±\pm2.2
AWM 49.06 46.23 47.17 47.48±\pm1.4 87.6 88.9 85.0 87.2±\pm2.0
ASI 47.17 43.40 44.34 44.97±\pm2.0 98.5 89.3 96.0 94.6±\pm4.8
RBank 39.62 38.68 41.50 39.93±\pm1.4 78.9 73.6 77.8 76.8±\pm2.8
Admin
Vanilla-IB 53.85 56.59 56.59 55.68±\pm1.6 95.2 98.6 103.2 99.0±\pm4.0
AWM 47.80 47.80 46.70 47.43±\pm0.6 136.6 147.0 144.4 142.7±\pm5.4
ASI 53.85 51.65 52.75 52.75±\pm1.1 132.0 152.3 133.9 139.4±\pm11.2
RBank 47.80 51.10 47.80 48.90±\pm1.9 120.8 126.4 126.5 124.6±\pm3.3
GitLab
Vanilla-IB 28.33 29.44 29.44 29.07±\pm0.6 79.8 76.2 78.0 78.0±\pm1.8
AWM 25.00 23.33 25.00 24.44±\pm1.0 93.1 91.0 92.6 92.3±\pm1.1
ASI 22.78 21.11 25.00 22.96±\pm2.0 102.3 107.0 113.8 107.7±\pm5.8
RBank 22.78 23.89 22.22 22.96±\pm0.9 71.9 72.4 73.8 72.7±\pm1.0
Table 13: Per-run success rate (%) and token usage per task (K) across three independent runs under GPT-5.4-mini on WebArena.
Success Rate (%) \uparrow Tok/task (K) \downarrow
Method Run a Run b Run c Mean±\pmStd Run a Run b Run c Mean±\pmStd
Shopping
Vanilla-IB 37.43 39.04 39.04 38.50±\pm0.9 55.1 53.5 54.2 54.3±\pm0.8
AWM 27.81 31.02 31.02 29.95±\pm1.9 64.2 61.6 60.2 62.0±\pm2.0
ASI 34.22 32.62 32.62 33.15±\pm0.9 70.6 83.5 66.2 73.4±\pm9.0
RBank 21.93 27.81 29.41 26.38±\pm3.9 59.3 58.2 56.2 57.9±\pm1.6
Reddit
Vanilla-IB 38.68 35.85 29.25 34.59±\pm4.8 89.4 102.5 96.6 96.2±\pm6.6
AWM 30.19 30.19 28.30 29.56±\pm1.1 77.1 94.4 69.8 80.4±\pm12.6
ASI 26.42 30.19 30.19 28.93±\pm2.2 83.8 85.0 92.8 87.2±\pm4.9
RBank 18.87 24.53 22.64 22.01±\pm2.9 67.8 72.5 67.3 69.2±\pm2.9
Admin
Vanilla-IB 35.16 35.16 37.36 35.89±\pm1.3 117.4 133.0 121.9 124.1±\pm8.0
AWM 28.57 33.52 34.62 32.24±\pm3.2 129.3 95.9 114.9 113.4±\pm16.8
ASI 28.57 35.16 35.16 32.96±\pm3.8 126.1 146.5 142.9 138.5±\pm10.9
RBank 36.26 32.97 33.52 34.25±\pm1.8 121.0 125.5 122.8 123.1±\pm2.3
GitLab
Vanilla-IB 20.00 24.44 22.22 22.22±\pm2.2 94.4 85.5 89.6 89.8±\pm4.4
AWM 16.11 18.33 17.22 17.22±\pm1.1 92.5 97.0 97.5 95.7±\pm2.7
ASI 21.11 21.11 20.00 20.74±\pm0.6 94.8 97.9 94.4 95.7±\pm1.9
RBank 11.11 15.56 16.67 14.45±\pm2.9 70.9 70.3 66.8 69.3±\pm2.2
Table 14: Per-run success rate (%) and token usage per task (K) across three independent runs under Qwen 3.6-27B on WebArena.
Success Rate (%) \uparrow Tok/task (K) \downarrow
Method Run a Run b Run c Mean±\pmStd Run a Run b Run c Mean±\pmStd
Shopping
Vanilla-IB 47.06 44.39 46.52 45.99±\pm1.4 57.9 60.5 57.9 58.8±\pm1.5
AWM 39.57 43.30 41.18 41.35±\pm1.9 73.6 66.7 75.1 71.8±\pm4.5
ASI 47.06 43.85 40.10 43.67±\pm3.5 72.6 81.1 82.8 78.8±\pm5.5
RBank 39.57 40.64 39.04 39.75±\pm0.8 65.6 65.2 65.3 65.4±\pm0.2
Reddit
Vanilla-IB 47.17 43.40 42.45 44.34±\pm2.5 94.5 86.5 88.0 89.7±\pm4.3
AWM 44.34 45.28 39.62 43.08±\pm3.0 110.3 87.8 112.8 103.6±\pm13.8
ASI 43.40 41.51 44.34 43.08±\pm1.4 118.1 125.9 117.6 120.5±\pm4.7
RBank 39.62 41.51 42.45 41.19±\pm1.4 96.1 95.9 105.4 99.1±\pm5.4
Admin
Vanilla-IB 49.45 52.20 50.54 50.73±\pm1.4 141.2 125.3 124.4 130.3±\pm9.5
AWM 45.60 43.40 49.45 46.15±\pm3.1 192.8 172.3 170.3 178.5±\pm12.5
ASI 51.65 48.35 47.25 49.08±\pm2.3 177.7 183.1 185.9 182.2±\pm4.2
RBank 46.15 47.25 49.45 47.62±\pm1.7 160.8 162.9 147.1 156.9±\pm8.6
GitLab
Vanilla-IB 27.22 29.44 27.78 28.15±\pm1.2 109.0 96.2 100.2 101.8±\pm6.5
AWM 23.33 22.22 24.44 23.33±\pm1.1 102.9 106.6 97.0 102.2±\pm4.8
ASI 27.22 25.55 24.44 25.74±\pm1.4 120.4 117.0 124.3 120.6±\pm3.6
RBank 20.56 22.78 19.44 20.93±\pm1.7 81.2 84.6 92.3 86.0±\pm5.7