On the information content of semi-structured databases

被引：0

作者：

Levene, Mark ^{[1
]}

机构：

[1] Department of Computer Science, University College London, Gower Street, London WC1E 6BT, United Kingdom

来源：

Acta Cybernetica | 1998年 / 13卷 / 03期

关键词：

D O I：

暂无

中图分类号：

学科分类号：

摘要：

In a semi-structured database there is no clear separation between the data and the schema, and the degree to which it is structured depends on the application. Semi-structured data is naturally modelled in terms of graphs which contain labels which give semantics to its underlying structure. Such databases subsume the modelling power of recent extensions of flat relational databases, to nested databases which allow the nesting (or encapsulation) of entities, and to object databases which, in addition, allow cyclic references between objects. Due to the flexibility of data modelling in a semi-structured environment, in any given application there may be different ways in which to enter the data, but it is not always clear when the semantics are the same. In order to compare different approaches to modelling the data we investigate a measure of the information content of typical semi-structured databases in order to test whether such databases are information-wise equivalent. For the purpose of our investigation we use a graph-based data model, called the hypernode model, as our model for semi-structured data and formalise flat, nested and object databases as subclasses of hypernode databases. We use formal language theory to define the context-free grammar induced by a hypernode database, and then formalise the information content of such a database as the language generated by this context-free grammar. Intuitively, the information content of a database provides us with a measure of how flexible the database is in modelling the information from different points of view. This enables us to prove the following results regarding the expressive power of databases: (1) in general, hypernode databases and thus semi-structured databases express the general class of context-free languages, (2) the class of flat databases expresses the class of finite languages whose words are of restricted length between one and four, (3) the class of nested databases expresses the class of finite languages, and (4) the class of object databases expresses the general class of regular languages. We then define two hypernode databases to be information-wise equivalent if they generate the same context-free language. This allows us to prove the following results regarding the computational complexity of determining whether two databases are information-wise equivalent or inequivalent: (1) the problem of determining information-wise equivalence of hypernode databases and thus semi-structured databases is, in general, undecidable, (2) the problem of determining information-wise equivalence of flat databases can be solved in time polynomial in the size of the two databases, (3) the problem of determining information-wise inequivalence of nested databases is NP-complete, and (4) the problem of determining information-wise inequivalence of object databases is PSPACE-complete.

引用

页码：257 / 275

共 50 条

[41] Keyword Search on Structured and Semi-Structured Data
Chen, Yi
Wang, Wei
Liu, Ziyang
Lin, Xuemin
ACM SIGMOD/PODS 2009 CONFERENCE, 2009, : 1005 - 1009
[42] Rationale in Semi-structured Processes
Kannengiesser, Udo
Zhu, Liming
BUSINESS PROCESS MANAGEMENT WORKSHOPS, 2011, 66 : 634 - +
[43] Autonomous vehicles in structured and semi-structured environments
Ozguner, U
Redmill, K
Ogras, U
Dagci, O
Launsbach, M
PROCEEDINGS OF THE 41ST IEEE CONFERENCE ON DECISION AND CONTROL, VOLS 1-4, 2002, : 124 - 129
[44] A Framework for Extracting Information from Semi-Structured Web Data Sources
Shaker, Malunoud
Ibrahim, Hamidah
Mustapha, Aida
Abdullah, Lili Nurliyana
THIRD 2008 INTERNATIONAL CONFERENCE ON CONVERGENCE AND HYBRID INFORMATION TECHNOLOGY, VOL 1, PROCEEDINGS, 2008, : 27 - 31
[45] Gathering services of IHWA from semi-structured web information sources
Jeong, JS
Oh, DI
IC'2001: PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON INTERNET COMPUTING, VOLS I AND II, 2001, : 375 - 378
[46] Information discovery from semi-structured sources - Application to astronomical literature
Dkaki, T
Dousset, B
Egret, D
Mothe, J
COMPUTER PHYSICS COMMUNICATIONS, 2000, 127 (2-3) : 198 - 206
[47] Low-Dimensionality Information Extraction Model for Semi-structured Documents
Belhadj, Djedjiga
Belaid, Abdel
Belaid, Yolande
COMPUTER ANALYSIS OF IMAGES AND PATTERNS, CAIP 2023, PT I, 2023, 14184 : 76 - 85
[48] A Knowledge Management System for Disseminating Semi-Structured Information in a Worldwide University
Maher, Peter E.
Kourik, Janet L.
2008 PORTLAND INTERNATIONAL CONFERENCE ON MANAGEMENT OF ENGINEERING & TECHNOLOGY, VOLS 1-5, 2008, : 1936 - 1942
[49] Low-Dimensionality Information Extraction Model for Semi-structured Documents
Belhadj, Djedjiga
Belaïd, Abdel
Belaïd, Yolande
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2023, 14184 LNCS : 76 - 85
[50] Reverse method for labeling the information from semi-structured web pages
Akbar, Z.
Handoko, L. T.
PROCEEDINGS OF THE 2009 INTERNATIONAL CONFERENCE ON SIGNAL PROCESSING SYSTEMS, 2009, : 551 - 555

← 1 2 3 4 5 →