Repository logo
Digital Collections
Research Outputs
People
Organizations
DRIS
Statistics
New user? Click here to register.Have you forgotten your password?
New user? Click here to register.Have you forgotten your password?
  1. Home
  2. EuroCRIS Research Output
  3. Conference
  4. Evaluating LLM-based metadata extraction for CRIS with a multi-criteria framework

Evaluating LLM-based metadata extraction for CRIS with a multi-criteria framework

Author(s)
Iwaniszewski, Maciej
PCG Academia
Keywords

research information ...

automated metadata ex...

artificial intelligen...

large language models...

automated metadata pr...

DSpace

Issue Date
May 20, 2026
Publisher
euroCRIS
Type
Conference Paper
Abstract
This work builds on research and hands-on implementation in DSpace-based CRIS environments, that included designing, building, and maintaining metadata workflows. Metadata extraction from research documents remains a persistent bottleneck. Librarians and repository administrators are required to manually create or verify records for publications, projects, and persons, and these records feed directly into institutional systems that depend on high data quality. Metadata underpins discovery and retrieval, as well as repository interoperability, including exposure via OAI-PMH and aggregation by OpenAIRE. FAIR-aligned practice presupposes consistent, well-formed metadata. Citation linking through Crossref and DataCite can break when metadata is incomplete or inconsistent. Likewise, digital preservation frameworks such as OAIS and PREMIS rely on metadata to maintain essential context over time. When metadata is incomplete or incorrect, records may be dropped from aggregators, publication to researcher links can break, and institutional reporting becomes unreliable.
Large language models can automate selected components of this work. Contemporary models can process a research article and generate structured metadata, including title, authors, affiliations, funding information, and keywords, within a single prompting step. When integrated with validation, normalization, and external registry lookups, they can form the basis of a viable extraction pipeline. Model performance, however, is not adequately represented by a single aggregate score. It depends on the model family and configuration, prompting strategy, document language, academic domain, media type, and the availability of tool-supported registry queries. In CRIS settings, performance must also be judged in terms of operational reliability. Attributes such as author affiliations, funding identifiers, and ORCID links have institutional implications, and errors can propagate into reporting, assessment, and compliance processes. For this reason, an evaluation framework should record provenance, represent uncertainty explicitly, maintain auditable decision logs, and flag higher-risk fields for targeted human review.
I propose two contributions. First, an LLM-based metadata extraction workflow for CRIS, implemented and tested in a working DSpace environment. The model performs the primary extraction by reading documents, identifying fields, and producing structured output, while validation, normalization, and linking steps are used to detect and correct errors. Second, a CRIS-oriented evaluation framework that compares multiple models across key experimental factors, measuring how closely extracted values align with authoritative reference sources using fuzzy similarity methods such as normalized word overlap, substring containment, and namenormalized Jaccard similarity for author fields. The evaluation is organized around three questions: (1) which metadata fields are reliably extractable from PDF text alone, and which require external sources regardless of the model? (2) how do model family, language, and domain interact to affect extraction quality? (3) where does repair prompting improve schema compliance, and where does it introduce new errors?
Description
17 slides.-- Presentation delivered within CRIS2026 session "AI for Metadata, Workflows, and Operational Research Information Management".-- Includes extended abstract
URI
https://dspacecris.eurocris.org/handle/11366/10294
File(s)
Thumbnail Image
Name

CRIS2026_paper-50_Iwaniszewski_Evaluating-LLM-based-metadata-extraction_extended-abstract.pdf

Description
Extended abstract
Size

584.46 KB

Format

Adobe PDF

Checksum (MD5)

64965d66d82ed4a03a461ae9b0d0d949

Thumbnail Image
Name

CRIS2026_Iwaniszewski_slides_Evaluating-LLM-based-metadata-extraction.pdf

Description
Presentation
Size

2.19 MB

Format

Adobe PDF

Checksum (MD5)

86a3b30225a520a32cba707eb8ea6182

Related items
Metrics
Conference(s)
euroCRIS

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify