On the Selection of an Optimal Set of Indexes

A problem of considerable interest in the design of a database is the selection of indexes. In this paper, we present a probabilistic model of transactions (queries, updates, insertions, and deletions) to a file. An evaluation function, which is based on the cost saving (in terms of the number of pa...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on software engineering 1983-03, Vol.SE-9 (2), p.135-143
Hauptverfasser:	Ip, M.Y.L., Saxton, L.V., Raghavan, V.V.
Format:	Artikel
Sprache:	eng
Schlagworte:	Algorithms Approximation algorithm attribute selection complexity Computer science Context modeling Cost function Councils Data base management systems Database systems Degradation Design Evaluation index selection Indexes Knapsack problem Polynomials Queries secondary index Selection Software Software algorithms Software engineering Transaction databases
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	A problem of considerable interest in the design of a database is the selection of indexes. In this paper, we present a probabilistic model of transactions (queries, updates, insertions, and deletions) to a file. An evaluation function, which is based on the cost saving (in terms of the number of page accesses) attributable to the use of an index set, is then developed. The maximization of this function would yield an optimal set of indexes. Unfortunately, algorithms known to solve this maximization problem require an order of time exponential in the total number of attributes in the file. Consequently, we develop the theoretical basis which leads to an algorithm that obtains a near optimal solution to the index selection problem in polynomial time. The theoretical result consists of showing that the index selection problem can be solved by solving a properly chosen instance of the knapsack problem. A theoretical bound for the amount by which the solution obtained by this algorithm deviates from the true optimum is provided. This result is then interpreted in the light of evidence gathered through experiments.
ISSN:	0098-5589 1939-3520
DOI:	10.1109/TSE.1983.236458