July 02, 2012
Last Friday, in what has become a not-so-unusual occurrence for the cloud services provider, Amazon's Elastic Compute Cloud went dark. The event, lasting two hours, happened just weeks after a power outage knocked the company's US-EAST-1 region offline for roughly six hours.
Both failures originated from Amazon operations based in Northern Virginia. Ars Technica detailed the earlier event, which was the result of primary power, primary backup power, and secondary backup power failures. Amazon explained that the issue began with a cable fault, disconnecting their primary power source. Shortly thereafter, the primary backup generator failed due to a faulty cooling fan. The Secondary backup was also inoperable due to an incorrectly configured circuit breaker.
Amazon had since promised that circuit breaker configuration would become part of their auditing process, but the message was met with some valid skepticism by Ars.
So, the breakers are fixed, but it's hard to imagine there won't be other problems in the future.
Surely enough, another power related event knocked out the US-EAST-1 region. This time, operations were affected by a major storm that left roughly 400,000 people without electricity. The issue took down websites Instagram, Pinterest, Heroku and Netflix for 2-3 hours.
Netflix is a prominent user of Amazon Web Services and is fully aware that the cloud provider is not infallible. Last April, their website famously stayed online during a major EC2 outage that took down Reddit, Quora, Hootsuite and Foursquare among others. Following that event, Netflix explained how their service stayed online during the Amazon outage.
Why were some websites impacted while others were not? For Netflix, the short answer is that our systems are designed explicitly for these sorts of failures. When we re-designed for the cloud this Amazon failure was exactly the sort of issue that we wanted to be resilient to.
Unfortunately, Friday's event took down the video streaming site as well. As of now, the Netflix tech blog has not posted a breakdown of the event. PC Mag received a vague explanation from a Netflix representative, saying the downtime was the result of a "rare technical issue that our engineers fixed."
The recent events demonstrate how fragile some portions of the Internet can be. They also act as a wake-up call to services relying on Amazon. HP, Microsoft and most recently Google, have entered the public cloud game, offering alternatives to EC2. If reliability continues to hinder the cloud giant, these competitors may be more than willing to tempt some of its current customers away.
10/30/2013 | Cray, DDN, Mellanox, NetApp, ScaleMP, Supermicro, Xyratex | Creating data is easy… the challenge is getting it to the right place to make use of it. This paper discusses fresh solutions that can directly increase I/O efficiency, and the applications of these solutions to current, and new technology infrastructures.
10/01/2013 | IBM | A new trend is developing in the HPC space that is also affecting enterprise computing productivity with the arrival of “ultra-dense” hyper-scale servers.
Ken Claffey, SVP and General Manager at Xyratex, presents ClusterStor at the Vendor Showdown at ISC13 in Leipzig, Germany.
Join HPCwire Editor Nicole Hemsoth and Dr. David Bader from Georgia Tech as they take center stage on opening night at Atlanta's first Big Data Kick Off Week, filmed in front of a live audience. Nicole and David look at the evolution of HPC, today's big data challenges, discuss real world solutions, and reveal their predictions. Exactly what does the future holds for HPC?