Table 1 : Per-domain success rate (%) and total token usage per task (K) on WebArena. Domain values are means over three independent runs (mean ± \pm std ); per-run details are in Appendix G . Avg. is the task-weighted mean over tasks of all four domains.
Shopping (187)
Reddit (106)
Admin (182)
GitLab (180)
Avg.
Method
SR
Tok
SR
Tok
SR
Tok
SR
Tok
SR
Tok
Gemini 3 Flash
Vanilla-IB
47.771.9
45.70.8
47.480.5
71.72.2
55.681.6
99.04.0
29.070.6
78.01.8
44.78
73.6
AWM
41.182.1
70.83.0
47.481.4
87.22.0
47.430.6
142.75.4
24.441.0
92.31.1
39.34
99.3
ASI
44.741.2
82.80.3
44.972.0
94.64.8
52.751.1
139.411.2
22.962.0
107.75.8
41.02
107.3
RBank
45.451.4
54.61.3
39.931.4
76.82.8
48.901.9
124.63.3
22.960.9
72.71.0
39.33
82.6
GPT-5.4-mini
Vanilla-IB
38.500.9
54.30.8
34.594.8
96.26.6
35.891.3
124.18.0
22.222.2
89.84.4
32.67
90.2
AWM
29.951.9
62.02.0
29.561.1
80.412.6
32.243.2
113.416.8
17.221.1
95.72.7
27.02
88.5
ASI
33.150.9
73.49.0
28.932.2
87.24.9
32.963.8
138.510.9
20.740.6
95.71.9
29.00
99.8
RBank
26.383.9
57.91.6
22.012.9
69.22.9
34.251.8
123.12.3
14.452.9
69.32.2
24.58
81.0
Qwen 3.6-27B
Vanilla-IB
45.991.4
58.81.5
44.342.5
89.74.3
50.731.4
130.39.5
28.151.2
101.86.5
42.14
95.5
AWM
41.351.9
71.84.5
43.083.0
103.613.8
46.153.1
178.512.5
23.331.1
102.24.8
38.01
115.0
ASI
43.673.5
78.85.5
43.081.4
120.54.7
49.082.3
182.24.2
25.741.4
120.63.6
40.15
125.8
RBank
39.750.8
65.40.2
41.191.4
99.15.4
47.621.7
156.98.6
20.931.7
86.05.7
37.00
101.9
Table 3 : Token breakdown by component under Gemini 3 Flash on WebArena , per task in thousands (K). Actor: tokens from LLM calls made during interaction steps. Modules: tokens from auxiliary components (workflow induction, skill synthesis, retrieval, verification). Both are further split into prompt ( P. ) and completion ( C. ) tokens. Values are means over three runs. Vanilla-IB has no module calls by construction.
Shopping
Reddit
Admin
GitLab
Actor
Modules
Actor
Modules
Actor
Modules
Actor
Modules
Method
P.
C.
P.
C.
P.
C.
P.
C.
P.
C.
P.
C.
P.
C.
P.
C.
Vanilla-IB
45.2
0.5
0
0
71.3
0.4
0
0
98.4
0.6
0
0
77.3
0.7
0
0
AWM
64.1
0.4
6.1
0.2
82.9
0.4
3.7
0.2
135.0
0.5
6.9
0.3
86.1
0.6
5.4
0.2
ASI
61.8
0.5
20.0
0.5
84.6
0.4
9.3
0.3
118.1
0.5
20.2
0.6
85.2
0.5
21.4
0.5
RBank
36.6
1.1
16.5
0.4
53.3
1.0
22.1
0.4
96.9
1.2
26.1
0.4
51.3
1.8
19.2
0.4
Table 4 : Any-of-3 and All-of-3 success rates (%) on the Shopping domain of WebArena across three models. Any-of-3 is the fraction of tasks succeeding in at least one run; All-of-3 is the fraction succeeding in all three runs. Δ ↑ \Delta_{\uparrow} and Δ ↓ \Delta_{\downarrow} are the gaps from the three-run mean (Table 1 ).
Gemini 3 Flash
GPT-5.4-mini
Qwen 3.6-27B
Method
Any
All
Any
All
Any
All
Vanilla-IB
54.01
+6.2
42.25
5.5
46.52
+8.0
27.81
10.7
54.01
+8.0
38.50
7.5
AWM
47.06
+5.9
35.83
5.4
37.43
+7.5
21.39
8.6
49.73
+8.4
33.16
8.2
ASI
51.34
+6.6
37.43
7.3
40.64
+7.5
26.74
6.4
52.41
+8.7
36.36
7.3
RBank
53.48
+8.0
38.50
7.0
35.29
+8.9
17.64
8.7
47.06
+7.3
34.22
5.5
Table 6: Augmented baselines with and without Vanilla-IB ’s pruning procedure on the Shopping domain of WebArena, under Gemini 3 Flash and Qwen 3.6-27B. Each baseline is shown in its released configuration and with the same rule-based accessibility-tree pruning used by Vanilla-IB ( + pruning ). Values are means over three runs (mean ± \pm std ). Success rate (%) and total token usage per task (K).
Gemini 3 Flash
Qwen 3.6-27B
Method
SR (%)
Tok (K)
SR (%)
Tok (K)
AWM
41.182.1
70.83.0
41.351.9
71.84.5
+ pruning
43.320.5
66.94.3
39.042.3
68.92.7
ASI
44.741.2
82.80.3
43.673.5
78.85.5
+ pruning
45.630.8
71.65.6
45.812.4
76.76.3
RBank
45.451.4
54.61.3
39.750.8
65.40.2
+ pruning
42.254.2
51.10.2
39.212.2
59.31.0
Vanilla-IB
47.771.9
45.70.8
45.991.4
58.81.5
Table 7: Success rate (%) and total token usage per task (K) at a higher step budget , under Qwen 3.6-27B on WebArena. The augmented methods use a 15-step actor horizon and Vanilla-IB a 20-step horizon. Values are means over three runs (mean ± \pm std ).
Shopping
Reddit
Admin
Method
SR (%)
Tok (K)
SR (%)
Tok (K)
SR (%)
Tok (K)
Vanilla-IB - 20 steps
46.520.9
60.84.1
47.173.8
91.63.3
51.652.9
161.422.4
AWM - 15 steps
41.893.4
79.53.8
44.030.5
105.01.4
47.434.2
212.66.9
ASI - 15 steps
45.632.9
96.41.4
43.081.1
127.413.7
51.833.0
206.811.3
RBank - 15 steps
42.242.8
79.11.4
41.832.7
111.65.4
50.181.1
198.024.2
Table 8 : ASI verification failures at the first high-level function call on WebArena. Values are averages over three runs. Attempts : total number of function induction attempts by ASI across all tasks. First-step fail : number of attempts where the induced function failed before the second agent action. Recov. rate : fraction of first-step failures where the verification episode was nonetheless judged correct, causing a potentially broken function to be stored in the shared library.
Domain
Attempts
First-step fail
Recov. rate
Gemini 3 Flash
Shopping
58.0
19.3
69.0%
Reddit
30.0
9.0
59.3%
Admin
64.3
6.3
63.2%
GPT-5.4-mini
Shopping
44.3
32.0
54.2%
Reddit
25.0
12.3
45.9%
Admin
51.7
22.7
47.1%
Table 9 : AWM workflow-induction statistics on WebArena, averaged per run. Induced : total induction events triggered by the judge for a run. From failed (%) : fraction originating from tasks the ground-truth evaluator classified as failed. Final WFs : distinct workflows in the final library after duplicate suppression.
Domain
Induced
From failed (%)
Final WFs
Gemini 3 Flash
Shopping
105.7
49.5%
37.3
Reddit
45.7
10.9%
16.7
Admin
99.7
42.1%
48.3
GPT-5.4-mini
Shopping
78.3
52.3%
33.7
Reddit
31.3
19.1%
13.7
Admin
50.0
42.0%
24.3
Table 10 : ReasoningBank memory-corpus statistics on WebArena, averaged over three runs. Tasks processed per run: 187 (Shopping), 106 (Reddit), 182 (Admin). Succ.-labeled : tasks labeled successful by the judge. FP in succ. : fraction of success-labeled tasks from trajectories that failed according to the ground-truth evaluator. FN in fail : fraction of fail-labeled tasks from trajectories that succeeded according to the ground-truth evaluator.
Domain
Totalmem.
Succ.-labeled
FP insucc. (%)
FN infail (%)
Gemini 3 Flash
Shopping
518
96
52.9%
40.7%
Reddit
294
32
30.2%
25.3%
Admin
507
86
42.6%
40.4%
GPT-5.4-mini
Shopping
544
67
59.5%
18.4%
Reddit
307
20
60.0%
17.4%
Admin
532
62
48.1%
25.2%
Table 11: Example failures of augmented agents where Vanilla-IB succeeds. Tasks are selected from WebArena (single runs shown; all from Gemini 3 Flash).
Method
Intent
Augmented agent behavior
Vanilla-IB behavior
AWM
“Buy the highest-rated product in Ceiling Light within a budget above $1000.”
Retrieved workflow for “Buy the highest-rated product” prescribes sorting then clicking the top result. Agent buys the wrong product because the workflow does not account for rating ties or the budget threshold.
Applies sorting, checks rating and price against the budget, and selects the correct product.
ASI
“Change the delivery address for my most recent order to 3 Oxford St, Cambridge, MA.”
Calls navigate_to_contact_page and fill_contact_form using the element ID from the docstring example (’1381’). Receives ValueError and TypeError; form submission fails.
Navigates to Address Book, adds the new address, recognizes that the order address cannot be changed through the UI, and reports accordingly.
RBank
“What are the top-5 best-selling products in 2023?”
Retrieved memory on handling ranking ties directs attention to tie-breaking. Agent misidentifies the 4th and 5th items (reports Sparta Gym Tank and Angel Light Running Short instead of Sprite Stasis Ball and Hawkeye Yoga Short).
Reads the same report without distraction and returns the correct ranked list.
Table 12: Per-run success rate (%) and token usage per task (K) across three independent runs under Gemini 3 Flash on WebArena.
Success Rate (%)
Tok/task (K)
Method
Run a
Run b
Run c
MeanStd
Run a
Run b
Run c
MeanStd
Shopping
Vanilla-IB
47.59
49.73
45.99
47.771.9
46.6
45.5
45.1
45.70.8
AWM
43.32
39.04
41.18
41.182.1
68.0
70.6
73.9
70.83.0
ASI
43.32
45.45
45.45
44.741.2
82.4
83.0
83.0
82.80.3
RBank
44.39
44.91
47.05
45.451.4
56.0
54.1
53.6
54.61.3
Reddit
Vanilla-IB
48.11
47.17
47.17
47.480.5
71.7
69.5
73.8
71.72.2
AWM
49.06
46.23
47.17
47.481.4
87.6
88.9
85.0
87.22.0
ASI
47.17
43.40
44.34
44.972.0
98.5
89.3
96.0
94.64.8
RBank
39.62
38.68
41.50
39.931.4
78.9
73.6
77.8
76.82.8
Admin
Vanilla-IB
53.85
56.59
56.59
55.681.6
95.2
98.6
103.2
99.04.0
AWM
47.80
47.80
46.70
47.430.6
136.6
147.0
144.4
142.75.4
ASI
53.85
51.65
52.75
52.751.1
132.0
152.3
133.9
139.411.2
RBank
47.80
51.10
47.80
48.901.9
120.8
126.4
126.5
124.63.3
GitLab
Vanilla-IB
28.33
29.44
29.44
29.070.6
79.8
76.2
78.0
78.01.8
AWM
25.00
23.33
25.00
24.441.0
93.1
91.0
92.6
92.31.1
ASI
22.78
21.11
25.00
22.962.0
102.3
107.0
113.8
107.75.8
RBank
22.78
23.89
22.22
22.960.9
71.9
72.4
73.8
72.71.0
Table 13: Per-run success rate (%) and token usage per task (K) across three independent runs under GPT-5.4-mini on WebArena.
Success Rate (%)
Tok/task (K)
Method
Run a
Run b
Run c
MeanStd
Run a
Run b
Run c
MeanStd
Shopping
Vanilla-IB
37.43
39.04
39.04
38.500.9
55.1
53.5
54.2
54.30.8
AWM
27.81
31.02
31.02
29.951.9
64.2
61.6
60.2
62.02.0
ASI
34.22
32.62
32.62
33.150.9
70.6
83.5
66.2
73.49.0
RBank
21.93
27.81
29.41
26.383.9
59.3
58.2
56.2
57.91.6
Reddit
Vanilla-IB
38.68
35.85
29.25
34.594.8
89.4
102.5
96.6
96.26.6
AWM
30.19
30.19
28.30
29.561.1
77.1
94.4
69.8
80.412.6
ASI
26.42
30.19
30.19
28.932.2
83.8
85.0
92.8
87.24.9
RBank
18.87
24.53
22.64
22.012.9
67.8
72.5
67.3
69.22.9
Admin
Vanilla-IB
35.16
35.16
37.36
35.891.3
117.4
133.0
121.9
124.18.0
AWM
28.57
33.52
34.62
32.243.2
129.3
95.9
114.9
113.416.8
ASI
28.57
35.16
35.16
32.963.8
126.1
146.5
142.9
138.510.9
RBank
36.26
32.97
33.52
34.251.8
121.0
125.5
122.8
123.12.3
GitLab
Vanilla-IB
20.00
24.44
22.22
22.222.2
94.4
85.5
89.6
89.84.4
AWM
16.11
18.33
17.22
17.221.1
92.5
97.0
97.5
95.72.7
ASI
21.11
21.11
20.00
20.740.6
94.8
97.9
94.4
95.71.9
RBank
11.11
15.56
16.67
14.452.9
70.9
70.3
66.8
69.32.2
Table 14: Per-run success rate (%) and token usage per task (K) across three independent runs under Qwen 3.6-27B on WebArena.