Skip to main navigation Skip to search Skip to main content

Resource Allocation for Primary-Site Fault-Tolerant Systems

  • Yennun Huang
  • , Satish K. Tripathi
  • Nokia

Research output: Contribution to journalArticlepeer-review

4 Scopus citations

Abstract

Resource allocation problems for distributed systems have been extensively studied for years. However, not many studies consider failure/repair behaviors of the systems and fault-tolerant overhead. As a result, solutions may not be applicable to fault-tolerant systems. In this paper, we study resource allocation for a distributed system employing the primary site approach for fault tolerance. Two kinds of systems are considered in this paper. The first consists of fault-tolerant nodes where each node has many duplicated servers. One server is the primary, which serves user requests, and the rest are backup. The second does not have fault-tolerant nodes. To tolerate node failures, each node uses other nodes as backups. When a node fails, all requests initially allocated to the node are served by one of its backups. To study the resource allocation for such systems, we develop an approximate model for each system. With these models, efficient allocation algorithms are presented, which takes into account the failure/repair rates of the system and the fault-tolerant overheads. From many experiments, it is shown that the algorithms give the optimal or suboptimal allocations. The algorithms, which incur little overhead, can improve the system performance significantly over an intuitive allocation algorithm.

Original languageEnglish
Pages (from-to)108-119
Number of pages12
JournalIEEE Transactions on Software Engineering
Volume19
Issue number2
DOIs
StatePublished - Feb 1993

Keywords

  • Checkpoint
  • fault tolerance
  • recovery
  • replication
  • resource allocation

Fingerprint

Dive into the research topics of 'Resource Allocation for Primary-Site Fault-Tolerant Systems'. Together they form a unique fingerprint.

Cite this