Papers tagged fault tolerance
- A Byzantine Fault Tolerant Distributed Commit Protocol
- A History of Erlang
- A History of the Virtual Synchrony Replication Model
- A Hundred Impossibility Proofs for Distributed Computing
- A Persistent System in Real Use - Experiences of the First 13 Years -
- A Response to Cheriton and Skeen’s Criticism of Causal and Totally Ordered Communication
- A simple totally ordered broadcast protocol
- A Solution to the Network Challenges of Data Recovery in Erasure-coded Distributed Storage Systems: A Study on the Facebook Warehouse Cluster
- Bigtable: A Distributed Storage System for Structured Data
- Byzantine Chain Replication
- Cassandra - A Decentralized Structured Storage System
- Ceph: A Scalable, High-Performance Distributed File System
- Chain Replication for Supporting High Throughput and Availability
- Chord: A Scalable Peer-to-peer Lookup Service for Internet Applications
- Commodifying Replicated State Machines with OpenReplica
- Comparing the Robustness of POSIX Operating Systems
- Consensus in the Presence of Partial Synchrony
- Consistent Hashing and Random Trees: Distributed Caching Protocols for Relieving Hot Spots on the World Wide Web
- Copysets: Reducing the Frequency of Data Loss in Cloud Storage
- Crash-Only Software
- Distributed Snapshots: Determining Global States of Distributed Systems
- Distributed Systems Reading Group
- Dynamo: Amazon’s Highly Available Key-value Store
- Epidemic Broadcast Trees
- f4: Facebook’s Warm BLOB Storage System
- f4: Facebook’s Warm BLOB Storage System
- Freenet: A Distributed Anonymous Information Storage and Retrieval System
- GN&C Fault Protection Fundamentals
- Harvest, Yield, and Scalable Tolerant Systems
- Hints for Computer System Design
- HyParView: a membership protocol for reliable gossip-based broadcast
- HyperDex: A Distributed, Searchable Key-Value Store
- Impossibility of Distributed Consensus with One Faulty Process
- In Search of an Understandable Consensus Algorithm
- Kelips*: Building an Efficient and Stable P2P DHT Through Increased Memory and Background Overhead
- Large-scale cluster management at Google with Borg
- Making reliable distributed systems in the presence of software errors
- MapReduce: Simplified Data Processing on Large Clusters
- Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center
- Microreboot – A Technique for Cheap Recovery
- MillWheel: Fault-Tolerant Stream Processing at Internet Scale
- Modularity and Scalability in Calvin
- Orleans: Distributed Virtual Actors for Programmability and Scalability
- Pastry: Scalable, decentralized object location and routing for large-scale peer-to-peer systems
- Paxos Made Live - An Engineering Perspective
- Paxos Made Moderately Complex
- Paxos Made Simple
- Readings in distributed systems
- SIFT: Design and Analysis of a Fault-Tolerant Computer for Aircraft Control
- Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-intensive Systems
- Sinfonia: A New Paradigm for Building Scalable Distributed Systems
- Sparrow: Distributed, Low Latency Scheduling
- SWIM: Scalable Weakly-consistent Infection-style Process Group Membership Protocol
- TAO: Facebook’s Distributed Data Store for the Social Graph
- Tardigrade: Leveraging Lightweight Virtual Machines to Easily and Efficiently Construct Fault-Tolerant Services
- The Chubby lock service for loosely-coupled distributed systems
- The Google File System
- The Google File System
- The ϕ Accrual Failure Detector
- There Is More Consensus in Egalitarian Parliaments
- Tiered Replication: A Cost-effective Alternative to Full Cluster Geo-replication
- TOWARDS A CLOUD COMPUTING RESEARCH AGENDA
- Warp: Multi-Key Transactions for Key-Value Stores
- Zab: High-performance broadcast for primary-backup systems
- ZooKeeper: Wait-free coordination for Internet-scale systems