Since I have a background and experience in cloud operations, I started
my research by looking at the tools available from the major cloud service
providers who compete with IBM. I reviewed Big Data service offerings
and tools from Alibaba, Amazon, Google, and Microsoft.
Alibaba has a solution called Hologres, a real-time data warehouse
compatible with Postgres SQL. They also offer an analytics product called
MaxCompute, enabling one to query and analyze data with low latency
and high concurrency. One can utilize MaxCompute to analyze logs, e-
commerce transactions, exploration of user profile data and behavior, and
other data analytics projects (What is MaxCompute? 2022). Alibaba also
has a product named QuickBI to perform data analytics, exploration, and
reporting.
Amazon offers several different data analytics services that support
computing with a large volume of unstructured and structured data with
low latency. One appealing offering is Amazon EMR. This offering
enables one to run Big Data workloads such as Apache Spark, Hive, and
Presto using complementary Amazon services such as Amazon EC2
clusters (Amazon EMR, n.d.).
Since we have not studied Presto in this class, further research revealed
that Presto is an open-source distributed SQL query engine. Presto is a
community-owned project by the Presto Foundation. (What is Presto?
n.d.). I found it interesting that one can use Presto to query data where it
resides (such as Cassandra, Hive, and traditional SQL databases), and
Facebook also uses Presto to run queries against its 300 PB (petabyte)
data warehouse (What is Presto? n.d.). From a glance at the Presto
documentation, it looks like one can use this tool as an alternative to
Apache Pig and Hive (Use Cases, n.d.). One benefit of using this tool
rather than Apache Pig or Hive is that one is not limited to just HDFS
queries. Even though this tool originated at Facebook, the documentation
indicates that Facebook and the open-source community continue
developing the tool (Use Cases, n.d.).
There are other computing service offerings that one can leverage in Big
Data analytics projects, such as Amazon Athena (to query Amazon S3
storage), Amazon RedShift (Data Warehousing), and Amazon Kinesis
(real-time video and stream analysis). Amazon QuickSight (business
analytics), Amazon OpenSearch (text and unstructured data search), and
AWS Glue DataBrew (data cleaning and normalization) round out their
data analytics offering (Analytics on AWS, n.d.).
I also noticed that Amazon provides several tutorials on using the service
that one could find helpful since a big part of Big Data analytics is
getting the data ingested into a data store to analyze the data. I liked that
Amazon's documentation appears to be thorough, mature, and full of
examples where one can take advantage of their service offerings.
Amazon also provides use cases and customer success stories to help
explain how their service offerings fit together to provide a complete Data
Analytics stack. As an example, Amazon explains how Pearson used
Amazon OpenSearch to analyze and gain insights from log data that it
collects to improve the company's learning platform (Pearson Boosts
Security and Productivity Using Amazon OpenSearch Service, 2020).
Not surprisingly, Google also has many service offerings that provide
compute capacity to handle a large volume of unstructured and structured
data with low latency. The centerpiece of their Big Data service offering
seems to be Google BigQuery, which Google claims to have a 26 – 34%
lower Total Cost of Ownership over three years than cloud data
warehouse alternatives (BigQuery, n.d.). BigQuery also has various client
libraries available. A data analyst can leverage familiar programming
languages such as Python, Java, Javascript, and Go to write and execute
queries to gain valuable insights from Big Data (What is BigQuery? n.d.).
One complimentary offering from Google that sounds interesting is their
BigQuery Omni product, which enables data analysis across clouds such
as Microsoft Azure and Amazon Web Services. One tremendous benefit to
Microsoft and Amazon customers is that they do not have to move their
data to Google Cloud to take advantage of Google BigQuery (BigQuery
Omni, n.d.). As a former Microsoft Azure administrator, I can certainly
appreciate Google's interest in serving customers to operate resources and
services in one or more public clouds.
Last but not least, Microsoft has several service offerings to support
computing with a large volume of unstructured and structured data with
low latency. It seems like their service offerings focus on either Azure
HDInsight or Azure Databricks. Azure HDInsight essentially enables one
to leverage a customizable environment of Hadoop components distributed
in Microsoft's cloud (What is Azure HDInsight? 2021). I found this
offering appealing since one can scale a workload up and down and only
pay for what one uses (similar to IBM's Watson Studio cloud offering).
Depending on business need, one can also select from any number of
cluster types to deploy in Microsoft Azure. For example, one could create
an Apache Spark cluster to support a Big Data analytics project where in-
memory processing is vital for performance reasons. Alternatively, one
could create an Apache Storm cluster if large data streams need to be
ingested and analyzed quickly.
I found it challenging to review some of these cloud-based service
offerings because it is hard to understand what Big Data technology is
behind each offering. It seems like some of the cloud vendors have
repurposed open-source tools for the cloud to make it easy for customers
to consume capacity on a pay-as-you-go basis, while others have built
their own proprietary tools to handle Big Data needs.
I think one needs to be naturally curious and determined to learn more to
understand the underpinnings of each one. One could spend hours reading
about the different features and experimenting with capabilities. For
example, I spent some time reading about Amazon Kinesis which provides
some exciting real-time data analysis capabilities with other Amazon
service offerings. I learned that Zillow uses Kinesis Data Streams to
collect public data and Multiple Listing Service (MLS) services which
helps the company keep buyers and sellers updated on estimated home
values (Amazon Kinesis, n.d.). This capability sounds like an Apache
Kafka or Flume feature, so I spent some time reading the Kinesis Data
Streams developer documentation to see what one can do with the service.
It isn't easy to render an informed opinion on which one is best with all
the available tools choices. My favorite answer, as a former consultant, is
- "it depends". I think one needs to carefully assess the business
requirements for any Big Data analytics project and then evaluate the
ability of various Big Data platforms, tools, and service providers to meet
the requirements. All kinds of questions come to mind: Does one need to
analyze data in real-time or using batch processing? How much disk
storage is required to store the data now and in the future? Who needs to
access the data, and what is the data governance strategy to ensure the
data is properly secured? Does the data need to be encrypted at rest? How
responsive do queries need to be? What kind of reports will one need to
generate and distribute? Is the data structured, semi-structured,
unstructured, or some combination of all three? Will a data warehouse
suffice to store the data, or is a more extensible data lake with data
originating from disparate data sources required? The list goes on and on,
but those are a few questions one may need to answer before selecting
Big Data tools, platforms, and a cloud service provider.
Lastly, my experience with product/tool/platform evaluations has been that
some vendors will practically fall over themselves to earn your business
and trust. For example, our agency was interested in determining how
well one vendor's machine learning platform would identify fraudulent
transactions in a dataset we had available. We invited the vendor to work
with us side by side in a sort of "hackathon" where we hacked through
the business problem and developed a machine learning model using the
vendor's platform. That exercise gave us a good understanding of the
vendor's capabilities and a baseline against which we could compare
products from other vendors. When it comes to vendors wanting to show
off shiny new toys coming out of their respective research labs, there is
typically no shortage of enthusiasm from the vendor's product managers.
I would be interested in hearing what tools others in our class have
discovered for computing with a large volume of unstructured and
structured with low latency.
Very Respectfully,
Michael Goddard
References:
"Amazon EMR" (n.d.). Amazon.com. Retrieved from
https://aws.amazon.com/emr/?c=a&sec=srv
"Amazon Kinesis" (n.d.). Amazon.com. Retrieved from
https://aws.amazon.com/kinesis/?c=a&sec=srv
"BigQuery Omni" (n.d.) Google.com. Retrieved from
https://cloud.google.com/bigquery-omni/docs/introduction
"Pearson Boosts Security and Productivity Using Amazon OpenSearch
Service" (2020). Amazon.com. Retrieved from
https://aws.amazon.com/solutions/case-studies/pearson-elasticsearch-case-
study/
"Use Cases" (n.d.). Presto.io. Retrieved from https://prestodb.io/
"What is Azure HDInsight?" (2021, November 18). Microsoft.com.
Retrieved from https://docs.microsoft.com/en-us/azure/hdinsight/hdinsight-
overview
"What is BigQuery?" (n.d.). Google.com Retrieved from
https://cloud.google.com/bigquery/docs/introduction
"What is MaxCompute?" (2022, January 28). Alibaba. Retrieved from
https://www.alibabacloud.com/help/en/doc-detail/27800.htm
"What is Presto?" (n.d.). Presto.io. Retrieved from https://prestodb.io/