Abstract
Resource allocation problems for distributed systems have been extensively studied for years. However, not many studies consider failure/repair behaviors of the systems and fault-tolerant overhead. As a result, solutions may not be applicable to fault-tolerant systems. In this paper, we study resource allocation for a distributed system employing the primary site approach for fault tolerance. Two kinds of systems are considered in this paper. The first consists of fault-tolerant nodes where each node has many duplicated servers. One server is the primary, which serves user requests, and the rest are backup. The second does not have fault-tolerant nodes. To tolerate node failures, each node uses other nodes as backups. When a node fails, all requests initially allocated to the node are served by one of its backups. To study the resource allocation for such systems, we develop an approximate model for each system. With these models, efficient allocation algorithms are presented, which takes into account the failure/repair rates of the system and the fault-tolerant overheads. From many experiments, it is shown that the algorithms give the optimal or suboptimal allocations. The algorithms, which incur little overhead, can improve the system performance significantly over an intuitive allocation algorithm.
| Original language | English |
|---|---|
| Pages (from-to) | 108-119 |
| Number of pages | 12 |
| Journal | IEEE Transactions on Software Engineering |
| Volume | 19 |
| Issue number | 2 |
| DOIs | |
| State | Published - Feb 1993 |
Keywords
- Checkpoint
- fault tolerance
- recovery
- replication
- resource allocation
Fingerprint
Dive into the research topics of 'Resource Allocation for Primary-Site Fault-Tolerant Systems'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver