New Benchmark Tests LLM Agent Memory in BIM Information Retrieval
A new benchmark called IFCMemoryBench has been developed by researchers to assess the long-term memory functions of LLM-based agents specifically within Building Information Modelling (BIM). This benchmark fills a void in current assessments, which typically focus on conversational recall in either open-domain or persona-based contexts. IFCMemoryBench evaluates an agent's ability to utilize information from previous interactions while engaging in a structured, domain-specific environment. It features 143 multi-session tasks across 19 projects and includes 4,016 prior sessions, sourced from incomplete-information inquiries in IFC-Bench v2. Each task introduces missing project context from earlier discussions, challenging agents to recall and utilize this information in later sessions. The findings are detailed in a paper available on arXiv (2607.26072).
Key facts
- IFCMemoryBench evaluates long-term memory of LLM-based agents in BIM.
- The benchmark contains 143 multi-session tasks across 19 projects.
- It includes 4,016 prior sessions derived from IFC-Bench v2.
- Tasks require agents to reuse information from earlier conversations.
- Existing evaluations focus on open-domain or persona-grounded recall.
- The benchmark tests agents in a live, structured, domain-specific environment.
- The paper is available on arXiv with ID 2607.26072.
- The work was announced as a cross-type submission.
Entities
Institutions
- arXiv