ITcon Vol. 31, pg. 990-1014, http://www.itcon.org/2026/42

Towards trustworthy LLMs for project management: Benchmarking a vector hybrid and graph RAG architectures

DOI:10.36680/j.itcon.2026.042
submitted:February 2026
published:August 2026
editor(s):Kumar B
authors:Eduardo Navarro-Bringas, Assistant Professor in Construction Management
Institute of Sustainable Built Environment (ISBE), Heriot-Watt University
https://orcid.org/0000-0003-0308-3860
E.Navarro_Bringas@hw.ac.uk
summary:Retrieval Augmented Generation (RAG) is a practical response, yet there is limited comparative evidence on how different RAG architectures trade off efficiency and answer quality for major project portfolio data. This article benchmarks two RAG pipelines using UK Government Major Projects Portfolio (GMPP) datasets (from 2020-2021 to 2023-2024) and a fixed query set, holding the foundation LLM model (i.e. GPT 4.0.), prompt template, and generation settings constant to isolate the effects of the two retrieval architectures. A vector-based hybrid pipeline (metadata filtering, dense retrieval, BM25 re-ranker) is compared with a labelled property graph-based pipeline using graph schema-aware retrieval. Results, evaluated through RAGAS metrics and a supplementary blind expert evaluation, indicate task-dependent trade-offs between the two architectures. While the vector-based RAG pipeline achieves lower query time latency and lower token consumption per query, the graph RAG pipeline improves answer quality based on RAGAS dimensions linked to trustworthiness. Particularly, graph RAG outperforms in terms of faithfulness (0.81 vs 0.78), context recall (0.87 vs 0.83) and answer relevancy (0.74 vs 0.69), suggesting better grounding in retrieved evidence and more complete coverage of the information required to provide relevant answers. Context precision is marginally higher for the vector pipeline (0.86 vs 0.85), which might be linked with the limited scoped document retrieval settings, while answer correctness performs comparably across both architectures. Expert evaluation also rated graph-based responses higher for accuracy, completeness and trustworthiness across the query subset, although preferences varied by query type. Overall, the findings highlight that architecture choice should be driven by expected query type, information risk, and the relative importance of efficiency, provenance and decision-support reliability in project and portfolio management.
keywords:LLMs, Retrieval Augmented Generation (RAG), vector index, knowledge graphs, major projects
full text: (PDF file, 1.055 MB)
citation:Navarro-Bringas, E. (2026). Towards trustworthy LLMs for project management: Benchmarking a vector hybrid and graph RAG architectures. Journal of Information Technology in Construction (ITcon), 31, 990-1014. https://doi.org/10.36680/j.itcon.2026.042
statistics: