Document Type

Article

Abstract

Knowledge graph (KG) schema engineering is labor-intensive and resists automation at scale. We investigate whether LLMs can generate domain-specific KG schemas of sufficient quality for downstream symbolic reasoning. We propose a tiered contextual framework that varies domain context richness across four levels: zero context, domain scope, task requirements, and data distribution. Generated schemas are evaluated intrinsically on BioRED (600 PubMed abstracts, multi-type entities and relations), where automated tiered schemas match an established KG construction baseline at 79.9% EC, with edge conformance rising from 47.5% at L1 to a stable 78–80% from L2 onward. Extrinsic evaluation on a 50-record MedHop controlled ablation shows that task requirements are the context level that improves utility, yielding 14% QA accuracy, a 6-point gain over domain scope alone. Data-distribution context achieves complete entity type coverage but does not further improve QA accuracy. An 86-point gap between the best schema-guided condition and the unstructured retrieval ceiling is closed by a schema-free KG under the same 1-hop retrieval, tracing the bottleneck to relations discarded during schema-guided extraction. LLM-generated schemas are viable, low-cost seeds for KG engineering pipelines; task requirements represent the minimum viable context threshold for neurosymbolic applications.

Rights

Copyright © 2027, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved

Share

COinS