Aleksey Charapko

Do not Blame (only) Network for Your Paxos Scalability Issues. (PPM Part 1)

Playing Around

Aleksey Charapko

·

Feb 17, 2018

In the past few months our lab has been doing a lot of work with different flavors of paxos consensus algorithm. Paxos and its numerous flavors are widely used in today’s cloud infrastructure. Distributed systems rely on it for many different tasks to ensure safe operation. For instance, coordination services use some consensus protocol flavor…
Read More
Trace Synchronization with HLC

Playing Around

Aleksey Charapko

·

Oct 18, 2017

Event logging or tracing is one of the most common techniques for collecting data about the software execution. For simple application running on the same machine, a trace of events timestamped with the machine’s hardware clock is typically sufficient. When the system grows and becomes distributed over multiple nodes, each node is going to produce…
Read More
One Page Summary: Flease – Lease Coordination without a Lock Server

One Page Summary

Aleksey Charapko

·

Sep 4, 2017

This paper talks about a decentralized lease management solution. In the past, many lock/lease services have been centralized, placing a single authority to manage all locks in the system. Google’s Chubby, Apache ZooKeeper, etcd, and others rely on a centralized approach and backed by some flavor of a consensus algorithm for fault-tolerance. According to Flease authors,…
Read More
One Page Summary: “milliScope: a Fine-Grained Monitoring Framework for Performance Debugging of n-Tier Web Services”

One Page Summary

Aleksey Charapko

·

Jun 19, 2017

Authors of the ICDCS2017 milliScope paper attack an interesting monitoring problem for distributed systems: detecting and determining a cause of short-lived events in the system. In particular, they address the issue of identifying very short bottlenecks (VSBs) in distributed web services. VSBs manifest themselves as performance degradation of a small number of requests, however they…
Read More
Sonification of Distributed Systems with RQL

Playing Around

Aleksey Charapko

·

Jun 12, 2017

In the past, I have discussed sonification as a mean of representing monitoring data. Aside from some silly and toy examples, sonifications can be used for serious applications. In many monitoring cases, the presence of some phenomena is more important than the details about it. In such situations, simple sonification is a perfect way to…
Read More
Retroscoping Zookeeper Staleness

Other Thoughts

Aleksey Charapko

·

Apr 24, 2017

ZooKeeper is a popular coordination service used as part of many large scale distributed systems. ZooKeeper provides a file-system inspired abstraction to the users on top of its replicated key-value store. Like other Paxos-inspired protocols, ZooKeeper is typically deployed on at least 3 nodes, and can tolerate F node failure for a cluster of size…
Read More
Why Government IT is Expensive and Archaic

Other Thoughts

Aleksey Charapko

·

Apr 6, 2017

Disclaimer: I do not work for the government, and my rant below is based on my very limited exposure to how IT works at the US government setting. Why Government IT is Expensive and Archaic? I think, this can be a very long discussion, but I do have a quick answer: standards imposed by government…
Read More
The First Datastore-driven Vehicle

Playing Around

Aleksey Charapko

·

Mar 25, 2017

It is not a secret that procrastination is the favorite activity of most PhD students. I have been procrastinating today, even though my advisor probably wants me to keep writing. In the midst of my procrastination, I thought: “Why are there self-driving vehicles, but no database-driven vehicles?” As absurd as it sounds, I gave it…
Read More
One Page Summary: “Musketeer: all for one, one for all in data processing systems”.

One Page Summary

Aleksey Charapko

·

Mar 12, 2017

Many distributed computation platforms and programming frameworks exist today, and new ones constantly popping out from the industry and academia. Some platforms are domain specific, such as TensorFlow for machine learning. Others, like Hadoop and Naiad are more general, and this generality allows for sophisticated and specialized programming abstractions to be built on top. So…
Read More
One Page Summary: “Slicer: Auto-Sharding for Datacenter Applications”

One Page Summary

Aleksey Charapko

·

Mar 8, 2017

One of the questions engineers of large distributed system must answer is “where to compute”. This is a big and important question, as we do not want to send a request originating in the US to some server in Australia. It simply makes no sense to incur the communication overhead if there are resources available…
Read More

Aleksey Charapko

Aleksey Charapko

Do not Blame (only) Network for Your Paxos Scalability Issues. (PPM Part 1)

Trace Synchronization with HLC

One Page Summary: Flease – Lease Coordination without a Lock Server

One Page Summary: “milliScope: a Fine-Grained Monitoring Framework for Performance Debugging of n-Tier Web Services”

Sonification of Distributed Systems with RQL

Retroscoping Zookeeper Staleness

Why Government IT is Expensive and Archaic

The First Datastore-driven Vehicle

One Page Summary: “Musketeer: all for one, one for all in data processing systems”.

One Page Summary: “Slicer: Auto-Sharding for Datacenter Applications”

Search

Recent Posts

Categories