Big Data in the cloud – avoiding cloud lock-in

In our previous article, we looked at different approaches to introducing Big Data technology in your business – either as a generic or specific solution deployed on-premises or in the cloud. Cloud obviously offers great flexibility for experimenting in the early stages, when you need quick iterations of trying and failing before you find the use case and solution that best fit your business needs.

Photo: Pexels

Cloud lock-in

However, in your try-and-fail iterations, you will need to focus on not falling into another pitfall – the cloud vendor lock-in, or simply cloud lock-in. By cloud lock-in, we mean using vendor-specific implementations, which are only provided by a particular cloud supplier. A good example here could be Amazon Kinesis or Google BigQuery. Using this specialized functionality may seem like a quick way to implement and deliver your business value; however, if your cloud provider phases out support for that functionality, your business may be forced to reimplement parts of, or the whole of, the system that depends on it. A good strategy against lock-in is particularly important for established businesses, although for startups with a relatively thin software stack, this isn’t such a big deal since the switching costs are usually still low.

Open source to the rescue

Open-source software has a strong track record of providing effective solutions to reduce vendor lock-in. It has helped fight vendor lock-in for decades. In particular, within operating systems, Linux has played an important role in fighting vendor locking. Taking this into the Big Data world, it does not take long to see that automation, particularly open-source automation tools, plays an important role in avoiding cloud lock-in. This could, for instance, be achieved by deploying and running the same complete Big Data stack on-premise and in the cloud.

Using automation tools such as Chef, Puppet, and Ansible Tower is one strategy for avoiding vendor lock-in and quickly moving between cloud providers. Also, container technologies like Docker or OpenShift Container Platform make it possible to deploy the same Big Data stack, whether it is Hortonworks, Cloudera, or MapR, across different cloud providers, making it easier to swap or even use multiple cloud setups in parallel to diversify operational risks.

What about Open Source lock-in?

Listening to Paul Cormier at Red Hat Forum 2016 (Oslo) last week, one could quickly get the impression that cloud lock-in can simply be avoided by promoting Open Source tools such as Ansible Tower or OpenShift Container Platform. These solutions effectively help turn the IaaS and PaaS resources offered by the Big Three (Amazon, Google, and Microsoft) and other cloud providers into a commodity. On the other hand, critics of Open Source could argue that by using this kind of solution, you actually end up with another kind of lock-in. However, the immense success of Open Source software over the last 15 years shows that lock-in in the case of an Open Source system is at most hypothetical. It is easy to find a similar alternative or, in the absolute worst-case scenario, to maintain the software yourself. Open Source, by its very nature, removes barriers to competitive advantage, and new ideas and features can be easily copied by anyone, anywhere, and at almost no cost.

All content on this site, excluding the photos and pictures, is licensed under a Creative Commons Attribution 4.0 International License.

Creative Commons License

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.