Abstract
A common assumption in the existing rollback techniques is that transients, the cause of most failures, subside very quickly, implying that a single retry of the program from the previous rollback point is sufficient. We discuss a general rollback strategy with n(n ≥ 2) retries which takes into consideration multiple transient failures as well as transients of long duration. Ways of deriving practical values of n for a given program are also discussed. Furthermore, we propose the use of a watchdog processor as an error detection tool to initiate recovery action through rollback, since the watchdog processor offers low error latency. We also discuss the merging of the watchdog processor with rollback recovery technique for enhancing the overall system reliability.
| Original language | English |
|---|---|
| Pages (from-to) | 87-95 |
| Number of pages | 9 |
| Journal | IEEE Transactions on Software Engineering |
| Volume | SE-12 |
| Issue number | 1 |
| DOIs | |
| State | Published - Jan 1986 |
Keywords
- Error detection
- error latency
- program retry
- recovery time
- roll back recovery
- transient errors
Fingerprint
Dive into the research topics of 'A Watchdog Processor Based General Rollback Technique with Multiple Retries'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver