On the information content of semi-structured databases

被引：0

作者：

Levene, Mark ^{[1
]}

机构：

[1] Department of Computer Science, University College London, Gower Street, London WC1E 6BT, United Kingdom

来源：

Acta Cybernetica | 1998年 / 13卷 / 03期

关键词：

D O I：

暂无

中图分类号：

学科分类号：

摘要：

In a semi-structured database there is no clear separation between the data and the schema, and the degree to which it is structured depends on the application. Semi-structured data is naturally modelled in terms of graphs which contain labels which give semantics to its underlying structure. Such databases subsume the modelling power of recent extensions of flat relational databases, to nested databases which allow the nesting (or encapsulation) of entities, and to object databases which, in addition, allow cyclic references between objects. Due to the flexibility of data modelling in a semi-structured environment, in any given application there may be different ways in which to enter the data, but it is not always clear when the semantics are the same. In order to compare different approaches to modelling the data we investigate a measure of the information content of typical semi-structured databases in order to test whether such databases are information-wise equivalent. For the purpose of our investigation we use a graph-based data model, called the hypernode model, as our model for semi-structured data and formalise flat, nested and object databases as subclasses of hypernode databases. We use formal language theory to define the context-free grammar induced by a hypernode database, and then formalise the information content of such a database as the language generated by this context-free grammar. Intuitively, the information content of a database provides us with a measure of how flexible the database is in modelling the information from different points of view. This enables us to prove the following results regarding the expressive power of databases: (1) in general, hypernode databases and thus semi-structured databases express the general class of context-free languages, (2) the class of flat databases expresses the class of finite languages whose words are of restricted length between one and four, (3) the class of nested databases expresses the class of finite languages, and (4) the class of object databases expresses the general class of regular languages. We then define two hypernode databases to be information-wise equivalent if they generate the same context-free language. This allows us to prove the following results regarding the computational complexity of determining whether two databases are information-wise equivalent or inequivalent: (1) the problem of determining information-wise equivalence of hypernode databases and thus semi-structured databases is, in general, undecidable, (2) the problem of determining information-wise equivalence of flat databases can be solved in time polynomial in the size of the two databases, (3) the problem of determining information-wise inequivalence of nested databases is NP-complete, and (4) the problem of determining information-wise inequivalence of object databases is PSPACE-complete.

引用

页码：257 / 275

共 50 条

[1] Designing good semi-structured databases
Lee, SY
Lee, ML
Ling, TW
Kalinichenko, LA
CONCEPTUAL MODELING - ER'99, 1999, 1728 : 131 - 145
[2] Geodesic search and retrieval of semi-structured databases
Rubin, Stuart H.
Chen, Shu-Ching
2007 IEEE INTERNATIONAL CONFERENCE ON SYSTEMS, MAN AND CYBERNETICS, VOLS 1-8, 2007, : 440 - +
[3] Conceptual graphs as schemas for semi-structured databases
Su, YF
Wong, KF
SEVENTH INTERNATIONAL CONFERENCE ON DATABASE SYSTEMS FOR ADVANCED APPLICATIONS, PROCEEDINGS, 2001, : 150 - 151
[4] Query translation for distributed heterogeneous structured and semi-structured databases
Al-Wasil, Fahad M.
Fiddian, N. J.
Gray, W. A.
FLEXIBLE AND EFFICIENT INFORMATION HANDLING, 2006, 4042 : 73 - 85
[5] Methods to Access Structured and Semi-Structured Data in Bioinformatics Databases: A Perspective
Moftah, Raja A.
Maatuk, Abdelsalam M.
White, Richard
2016 INTERNATIONAL CONFERENCE ON ENGINEERING & MIS (ICEMIS), 2016,
[6] Toward structured retrieval in semi-structured information spaces
Huffman, SB
Baudin, C
IJCAI-97 - PROCEEDINGS OF THE FIFTEENTH INTERNATIONAL JOINT CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOLS 1 AND 2, 1997, : 751 - 756
[7] Automatic Content Extraction on Semi-Structured Documents
dos Santos, Jose Eduardo Bastos
11TH INTERNATIONAL CONFERENCE ON DOCUMENT ANALYSIS AND RECOGNITION (ICDAR 2011), 2011, : 1235 - 1239
[8] Managing unstructured and semi-structured information in organisations
Aitken, Ashley M.
6th IEEE/ACIS International Conference on Computer and Information Science, Proceedings, 2007, : 712 - 717
[9] Hyperset Approach to Semi-structured Databases Computation of Bisimulation in the Distributed Case
Molyneux, Richard
Sazonov, Vladimir
DATASPACE: THE FINAL FRONTIER, PROCEEDINGS, 2009, 5588 : 78 - 90
[10] Extracting information from semi-structured Internet sources
Jeong, JS
Oh, DI
ISIE 2001: IEEE INTERNATIONAL SYMPOSIUM ON INDUSTRIAL ELECTRONICS PROCEEDINGS, VOLS I-III, 2001, : 1378 - 1381

← 1 2 3 4 5 →