Performance issues in distributed shared-nothing information-retrieval systems

被引:5
|
作者
Tomasic, A [1 ]
GarciaMolina, H [1 ]
机构
[1] STANFORD UNIV,DEPT COMP SCI,STANFORD,CA 94305
关键词
D O I
10.1016/S0306-4573(96)00019-2
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Many information-retrieval systems provides access to abstracts. For example, Stanford University, through its FOLIO system, provides access to the INSPEC database of abstracts of the literature on physics, computer science, electrical engineering, etc. In this article, this database is studied by using a trace-driven simulation. It focuses on a physical-index design that accommodates truncations, inverted-index caching, and database scaling in a distributed shared-nothing system. All three issues are shown to have a strong effect on response time and throughput. Database scaling is explored in two ways. One way assumes an ''optimal'' configuration for a single host and then linearly scales the database by duplicating the host architecture as needed. The second way determines the optimal number of hosts given a fixed database size. Copyright (C) 1996 Elsevier Science Ltd
引用
收藏
页码:647 / 665
页数:19
相关论文
共 50 条