Skip to main content

Strata + Hadoop World for you and for me

Big Data has a big marketing problem: it’s been sold as a universal game-changer, but is perceived by some in the general public as overblown hype or a toy for the exclusive benefit of those few companies that can already afford it.
The situation hasn’t been helped by the recent revelations of widespread NSA communications monitoring, which have added to the not-unreasonable fears that always existed around mass-scale data collection.
Industry insiders see the immense progress being made on the technology front and understand the transformative power those shifts hold for society. But society, on the whole, may harbor some doubts.
It’s refreshing, then, to see that this year’s Strata + Hadoop World conference is taking these concerns seriously and making a concerted effort to address them.
Industry ethics, data security, and privacy issues are among the main focuses of this year’s conference. Perhaps more importantly, the event, already a dependable weathervane for big data development, embodies the internal recognition that the industry, for the sake of its own growth as much as society’s, needs to make meaningful results accessible to the broader market.
From the official guest speakers list, the conference reflects the diverse present-day landscape of big data innovation and makes looking forward to a more inclusive future one of its main objectives. Corporate executives will appear alongside startup CTOs now producing the insights that make much of big enterprise data solutions possible. Both will offer their perspectives in a setting shared by journalists, historians, and university experts invited to contextualize the industry’s achievements and frame its challenges moving forward.
One of the main consensuses likely to emerge from the integrated lineup of workshops, panel discussions, presentations, and demonstrations is that the industry’s upside lies in looking inward. Even as fundamental tools like Hadoop, Cassandra, Storm, Spark/Shark, and DrillPublic, extend the broader possibilities of big data further and further, smaller, more specialized services have emerged to deliver usable insights for businesses, governments, and organizations in general. Looking ahead to the next few years, it’s that new bottom-up dynamic that holds some of the industry’s strongest promise.
As a result, a recent IDC report predicts big data, as an industry, will ride 27% compounded annual growth to reach $32.4 billion globally by 2017. According to O’Reilly statistics, data science job posts have already jumped 89% year-over-year, with data engineering openings rising by 38%. Gartner placed “advanced, pervasive, invisible analytics” at #4 on its list of top strategic IT trends for 2015, a prediction informed by the increasing ubiquity of mobile computing devices and standardization of built-in analytics within the mobile app industry. In addition, Gartner noted that the era of the smart machine is upon us, and predicted that it will be the most disruptive in the history of IT.
Strata + Hadoop world is the place to understand this process and its far-reaching implications. Big tech brands were to be expected, but even at a first glance, the sheer range of industries in attendance speaks volumes. Iconic industry names from banking, manufacturing, energy , utilities, telecom have all registered for the conference, eager to make connections with the smaller software services also on display and learn what the newly stratified face of big data could mean in their respective fields.
Big data is finally beginning to incorporate the innovation that carries huge impact for the world and that carries huge impact for you and for me. So, let’s gear up to hear what the Hadoop fraternity has to say about it and how the world responds to it.

About the author:

Sundeep Sanghavi is the CEO and Co-Founder of DataRPM which is an award-winning, industry pioneer in smart machine analytics for big data.


Popular posts from this blog

In-memory data model with Apache Gora

Open source in-memory data model and persistence for big data framework Apache Gora™ version 0.3, was released in May 2013. The 0.3 release offers significant improvements and changes to a number of modules including a number of bug fixes. However, what may be of significant interest to the DynamoDB community will be the addition of a gora-dynamodb datastore for mapping and persisting objects to Amazon's DynamoDB. Additionally the release includes various improvements to the gora-core and gora-cassandra modules as well as a new Web Services API implementation which enables users to extend Gora to any cloud storage platform of their choice. This 2-part post provides commentary on all of the above and a whole lot more, expanding to cover where Gora fits in within the NoSQL and Big Data space, the development challenges and features which have been baked into Gora 0.3 and finally what we have on the road map for the 0.4 development drive.
Introducing Apache Gora Although there are var…

Data deduplication tactics with HDFS and MapReduce

As the amount of data continues to grow exponentially, there has been increased focus on stored data reduction methods. Data compression, single instance store and data deduplication are among the common techniques employed for stored data reduction.
Deduplication often refers to elimination of redundant subfiles (also known as chunks, blocks, or extents). Unlike compression, data is not changed and eliminates storage capacity for identical data. Data deduplication offers significant advantage in terms of reduction in storage, network bandwidth and promises increased scalability.
From a simplistic use case perspective, we can see application in removing duplicates in Call Detail Record (CDR) for a Telecom carrier. Similarly, we may apply the technique to optimize on network traffic carrying the same data packets.
Some of the common methods for data deduplication in storage architecture include hashing, binary comparison and delta differencing. In this post, we focus on how MapReduce and…

Amazon DynamoDB datastore for Gora

What was initially suggested during causal conversation at ApacheCon2011 in November 2011 as a “neat idea”, would soon become prime ground for Gora's first taste of participation within Google's Summer of Code program. Initially, the project, titled Amazon DynamoDB datastore for Gora, merely aimed to extend the Gora framework to Amazon DynamoDB. However, it seem became obvious that the issue would include much more than that simple vision.

The Gora 0.3 Toolbox We briefly digress to discuss some other noticeable additions to Gora in 0.3, namely: Modification of the Query interface: The Query interface was amended from Query<K, T> to Query<K, T extends Persistent> to be more precise and explicit for developers. Consequently all implementors and users of the Query interface can only pass object's of Persistent type. Logging improvements for data store mappings: A key aspect of using Gora well is the establishment and accurate definitio…