Papers in distributed systems
- A Byzantine Fault Tolerant Distributed Commit Protocol
- A History of the Virtual Synchrony Replication Model
- A Hundred Impossibility Proofs for Distributed Computing
- A Note on Distributed Computing
- A Response to Cheriton and Skeen’s Criticism of Causal and Totally Ordered Communication
- A simple totally ordered broadcast protocol
- A Universal Modular ACTOR Formalism for Artificial Intelligence
- A Versatile Scheme for Routing Highly Variable Traffic in Service Overlays and IP Backbones
- Beehive: O(1) Lookup Performance for Power-Law Query Distributions in Peer-to-Peer Overlays
- Brewer’s Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services
- Byzantine Chain Replication
- Chain Replication for Supporting High Throughput and Availability
- Chord: A Scalable Peer-to-peer Lookup Service for Internet Applications
- Commodifying Replicated State Machines with OpenReplica
- Consensus in the Presence of Partial Synchrony
- Consistent Global States of Distributed Systems: Fundamental Concepts and Mechanisms
- Consistent Hashing and Random Trees: Distributed Caching Protocols for Relieving Hot Spots on the World Wide Web
- Copysets: Reducing the Frequency of Data Loss in Cloud Storage
- Dapper, a Large-Scale Distributed Systems Tracing Infrastructure
- Distributed Snapshots: Determining Global States of Distributed Systems
- Eluding Carnivores: File Sharing with Strong Anonymity
- END-TO-END ARGUMENTS IN SYSTEM DESIGN
- Epidemic Algorithms for Replicated Database Maintenance
- f4: Facebook’s Warm BLOB Storage System
- Harvest, Yield, and Scalable Tolerant Systems
- Herbivore: A Scalable and Efficient Protocol for Anonymous Communication
- High-Level Specifications: Lessons from Industry
- Hoard: A Scalable Memory Allocator for Multithreaded Applications
- How the Hidden Hand Shapes the Market for Software Reliability
- Implementing Fault-Tolerant Services Using the State Machine Approach: A Tutorial
- Implementing the Omega failure detector in the crash-recovery failure model
- Impossibility of Distributed Consensus with One Faulty Process
- In Search of an Understandable Consensus Algorithm
- IronFleet: Proving Practical Distributed Systems Correct
- Kafka: a Distributed Messaging System for Log Processing
- Kelips*: Building an Efficient and Stable P2P DHT Through Increased Memory and Background Overhead
- Large-scale cluster management at Google with Borg
- Large-scale Incremental Processing Using Distributed Transactions and Notifications
- Life beyond Distributed Transactions: an Apostate’s Opinion
- Linearizability: A Correctness Condition for Concurrent Objects
- MapReduce: Simplified Data Processing on Large Clusters
- Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center
- MillWheel: Fault-Tolerant Stream Processing at Internet Scale
- Oblivious Routing of Highly Variable Traffic in Service Overlays and IP Backbones
- Omega: flexible, scalable schedulers for large compute clusters
- ON PROOF AND PROGRESS IN MATHEMATICS
- Orleans: Distributed Virtual Actors for Programmability and Scalability
- P⁵: A Protocol for Scalable Anonymous Communication
- Pastry: Scalable, decentralized object location and routing for large-scale peer-to-peer systems
- Paxos Made Live - An Engineering Perspective
- Paxos Made Moderately Complex
- Paxos Made Simple
- Practical Byzantine Fault Tolerance and Proactive Recovery
- Pregel: A System for Large-Scale Graph Processing
- Replication, History, and Grafting in the Ori File System
- Self-stabilizing Systems in Spite of Distributed Control
- SIFT: Design and Analysis of a Fault-Tolerant Computer for Aircraft Control
- Signal/Collect: Graph Algorithms for the (Semantic) Web
- Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-intensive Systems
- Sinfonia: A New Paradigm for Building Scalable Distributed Systems
- Solution of a Problem in Concurrent Programming Control
- Sparrow: Distributed, Low Latency Scheduling
- Sparse Partitions (Extended Abstract)
- Stronger Semantics for Low-Latency Geo-Replicated Storage
- The Akamai Network: A Platform for High-Performance Internet Applications
- The Chubby lock service for loosely-coupled distributed systems
- The Dining Cryptographers Problem: Unconditional Sender and Recipient Untraceability
- The Join Calculus: a Language for Distributed Mobile Programming
- The Part-Time Parliament
- THE SWIRLDS HASHGRAPH CONSENSUS ALGORITHM: FAIR, FAST, BYZANTINE FAULT TOLERANCE
- There Is More Consensus in Egalitarian Parliaments
- Tiered Replication: A Cost-effective Alternative to Full Cluster Geo-replication
- Tor: The Second-Generation Onion Router
- TOWARDS A CLOUD COMPUTING RESEARCH AGENDA
- Transactional Client-Server Cache Consistency: Alternatives and Performance
- Understanding the Limitations of Causally and Totally Ordered Communication
- Unicorn: A System for Searching the Social Graph
- Unikernels: Library Operating Systems for the Cloud
- VIEWING CONTROL STRUCTURES as PATTERNS of PASSING MESSAGES
- VL2: A Scalable and Flexible Data Center Network
- Zab: High-performance broadcast for primary-backup systems
- ZooKeeper: Wait-free coordination for Internet-scale systems