Overview
MemAudit evaluates long-term agent memory as an artifact that can be inspected after an interaction. Agents assist simulated users, then a separate recovery step reconstructs hidden user attributes from the stored memory. The benchmark includes 50 users, 31 hidden dimensions per user, and five memory systems, with both full-store and retrieval-limited access.
The results distinguish successful task completion from faithful user-state retention: an agent can finish its tasks while retaining only a partial understanding of the user. This provides a direct evaluation target for more reliable personalized agents.