← 返回阅读笔记 / Back to reading note

M2A: Multimodal Memory Agent with Dual-Layer Hybrid Memory for Long-Term Personalized Interactions

SUP-A27 原表

来源:作者原文。2026-09-07抓取,原始行列,不是统一复跑成绩。

Table 1 : Experimental results on LoCoMo dataset. The best results are highlighted in bold , and the second-best results are underlined . Models are evaluated on LLM-as-a-Judge with evaluation model Qwen3-VL-32B(Q), GPT-4o(G) and their average(Avg).
Model Method Multi Hop Temporal Open Domain Single Hop Visual Total
Q G Avg Q G Avg Q G Avg Q G Avg Q G Avg Q G Avg
GPT 4o-mini LoCoMo 27.10 24.82 25.96 17.45 16.82 17.14 36.14 29.17 32.66 43.93 41.97 42.95 31.10 30.27 30.68 34.19 32.35 33.27
Mem0 25.89 26.60 26.25 15.62 17.13 16.38 27.08 31.25 29.17 45.07 43.87 44.47 36.39 36.94 36.67 34.66 34.80 34.73
A-MEM 28.36 25.18 26.77 24.61 22.12 23.37 30.21 34.38 32.30 45.18 44.23 44.71 36.39 36.70 36.54 36.79 35.73 36.26
M2A (ours) 27.30 28.01 27.66 31.15 32.40 31.78 40.63 36.46 38.55 56.24 56.71 56.48 44.03 42.51 43.27 44.61 44.67 44.64
Qwen3VL 8b LoCoMo 24.82 23.76 24.29 28.97 29.28 29.13 31.25 30.21 30.73 43.88 48.04 45.96 37.31 35.17 36.24 36.64 37.98 37.31
Mem0 34.75 35.82 35.29 33.64 34.58 34.11 34.38 33.33 33.86 54.34 51.96 53.15 41.59 39.14 40.37 44.56 43.33 43.95
A-MEM 40.07 40.78 40.43 31.78 30.84 31.31 32.29 30.21 31.25 44.63 45.07 44.85 39.14 38.53 38.84 40.14 40.07 40.10
M2A (ours) 43.26 43.62 43.44 39.25 42.67 40.96 51.04 48.96 50.00 66.71 63.14 64.93 54.13 51.68 52.90 55.44 53.94 54.69
GLM-4.6V Flash LoCoMo 25.18 24.47 24.83 24.61 22.12 23.37 32.29 27.08 29.69 45.18 43.76 44.47 35.17 32.11 33.64 36.21 34.23 35.22
Mem0 43.97 45.04 44.51 34.27 33.02 33.65 39.58 35.42 37.50 57.91 57.67 57.79 40.37 39.45 39.91 47.73 47.19 47.46
A-MEM 41.84 39.36 40.60 29.90 31.15 30.53 35.42 36.46 35.94 49.94 45.90 47.92 44.65 41.90 43.28 43.60 41.19 42.39
M2A (ours) 48.58 47.51 48.05 39.87 42.05 40.96 57.30 52.08 54.69 66.11 66.47 66.29 55.66 52.91 54.29 56.67 56.29 56.48
Table 2 : Comparison of M 2 A and its variation on Qwen3-VL-8B.
Variant Qwen3-VL GPT-4o Avg.
M2A (Full) 55.44 53.94 54.69
w/o Dual-layer (semantic only) 41.56 41.19 41.38
w/o Iterative (single retrieval) 38.08 39.26 38.67
w/o Tri-path (text dense only) 51.15 50.03 50.59