Evaluating LLM-based metadata extraction for CRIS with a multi-criteria framework
Author(s)
Iwaniszewski, Maciej
PCG Academia
Issue Date
May 20, 2026
Publisher
euroCRIS
Type
Conference Paper
Abstract
This work builds on research and hands-on implementation in DSpace-based CRIS environments, that included designing, building, and maintaining metadata workflows. Metadata extraction from research documents remains a persistent bottleneck. Librarians and repository administrators are required to manually create or verify records for publications, projects, and persons, and these records feed directly into institutional systems that depend on high data quality. Metadata underpins discovery and retrieval, as well as repository interoperability, including exposure via OAI-PMH and aggregation by OpenAIRE. FAIR-aligned practice presupposes consistent, well-formed metadata. Citation linking through Crossref and DataCite can break when metadata is incomplete or inconsistent. Likewise, digital preservation frameworks such as OAIS and PREMIS rely on metadata to maintain essential context over time. When metadata is incomplete or incorrect, records may be dropped from aggregators, publication to researcher links can break, and institutional reporting becomes unreliable.
Large language models can automate selected components of this work. Contemporary models can process a research article and generate structured metadata, including title, authors, affiliations, funding information, and keywords, within a single prompting step. When integrated with validation, normalization, and external registry lookups, they can form the basis of a viable extraction pipeline. Model performance, however, is not adequately represented by a single aggregate score. It depends on the model family and configuration, prompting strategy, document language, academic domain, media type, and the availability of tool-supported registry queries. In CRIS settings, performance must also be judged in terms of operational reliability. Attributes such as author affiliations, funding identifiers, and ORCID links have institutional implications, and errors can propagate into reporting, assessment, and compliance processes. For this reason, an evaluation framework should record provenance, represent uncertainty explicitly, maintain auditable decision logs, and flag higher-risk fields for targeted human review.
I propose two contributions. First, an LLM-based metadata extraction workflow for CRIS, implemented and tested in a working DSpace environment. The model performs the primary extraction by reading documents, identifying fields, and producing structured output, while validation, normalization, and linking steps are used to detect and correct errors. Second, a CRIS-oriented evaluation framework that compares multiple models across key experimental factors, measuring how closely extracted values align with authoritative reference sources using fuzzy similarity methods such as normalized word overlap, substring containment, and namenormalized Jaccard similarity for author fields. The evaluation is organized around three questions: (1) which metadata fields are reliably extractable from PDF text alone, and which require external sources regardless of the model? (2) how do model family, language, and domain interact to affect extraction quality? (3) where does repair prompting improve schema compliance, and where does it introduce new errors?
Description
17 slides.-- Presentation delivered within CRIS2026 session "AI for Metadata, Workflows, and Operational Research Information Management".-- Includes extended abstract
URI
https://dspacecris.eurocris.org/handle/11366/10294
File(s)![Thumbnail Image]()
![Thumbnail Image]()
Name
CRIS2026_paper-50_Iwaniszewski_Evaluating-LLM-based-metadata-extraction_extended-abstract.pdf
Description
Extended abstract
Size
584.46 KB
Format
Adobe PDF
Checksum (MD5)
64965d66d82ed4a03a461ae9b0d0d949
Name
CRIS2026_Iwaniszewski_slides_Evaluating-LLM-based-metadata-extraction.pdf
Description
Presentation
Size
2.19 MB
Format
Adobe PDF
Checksum (MD5)
86a3b30225a520a32cba707eb8ea6182
Conference(s)
