Last week, Apache Tez graduated to become a top level project within the Apache Software Foundation (ASF). This represents a major step forward for the project and is representative of its momentum that has been built by a broad community of developers from not only Hortonworks but Cloudera, Facebook, LinkedIn, Microsoft, NASA JPL, Twitter, and Yahoo as well.What is Apache Tez and why is it useful?
The Hortonworks Blog
- Business Values of Hadoop
- Why Hortonworks
- Industry Verticals
- Industry Happenings
- Deployment Options
- Types of Data
We know the saying, “what happens in Vegas, stays in Vegas,” but we can’t keep this all to ourselves. We’re super excited to be among 12 of HP AllianceOne Partners to receive the HP Partner of the Year award at HP Discover 2014.
The HP AllianceOne Award for ConvergedSystems recognizes Hortonworks for its strategic relationship and reseller agreement for the Hortonworks Data Platform (HDP) and for delivering solutions that drive meaningful business for our shared customers.…
Apache Ambari has always provided an operator the ability to provision an Apache Hadoop cluster using an intuitive Cluster Install Wizard web interface, guiding the user through a series of steps:
- confirming the list of hosts
- assigning master, slave, and client components to configuring services, and
- installing, starting and testing the cluster.
With Ambari Blueprints, system administrators and dev-ops engineers can expedite the process of provisioning a cluster. Once defined, Blueprints can be re-used, which facilitates easy configuration and automation for each successive cluster creation.…
Last week’s release of HDP 2.1 was packed with countless new features for enterprise Hadoop. These included new processing capabilities with Tez and Hive on YARN, Solr and Storm, to operations with Ambari, governance with Falcon and security with Knox.
To guide you through these capabilities, Hortonworks is hosting a new series of webinars beginning on May 8 and running to June 26.
You can join any or all of the webinars listed below, and we’ve provided a simple way of signing up for all 7.…
Hortonworks Data Platform 2.1 for Windows is the 100% open source data management platform based on Apache Hadoop and available for the Microsoft Windows Server platform. I have built a helper tool that automates the process of deploying a multi-node Hadoop cluster – utilizing the MSI available in HDP 2.1 for Windows.
HDP on Windows installation package comes in the format of MSI, Microsoft’s MSI format utilizes the installation and configuration service provided with Windows called Windows Installer.…
The Apache Hive community has voted on and released version 0.13 today. This is a significant release that represents a major effort from over 70 members who worked diligently to close out over 1080 JIRA tickets.
Hive 0.13 also delivers the third and final phase of the Stinger Initiative, a broad community based initiative to drive the future of Apache Hive, delivering 100x performance improvements at petabyte scale with familiar SQL semantics.…
As enterprises build new applications with the data they cost effectively capture and process with Apache Hadoop it is important for the platform to facilitate the app dev processes. That’s why we are excited to announce that we’ve expanded our partnership with Concurrent, Inc. to simplify and accelerate application development on Hadoop.
There are two components to this expanded partnership.
As the Red Hat Summit shifts to the west coast in San Francisco this year Hortonworks and Red Hat will be demonstrating the progress of our engineering efforts. Our engineers have been hard at work in the factories and in the communities deeply integrating our open source offerings to create a comprehensive platform for new analytic applications. As a reminder in February Red Hat and Hortonworks announced a comprehensive open source initiative to deliver infrastructure solutions to bring 100-percent open source Hadoop to the hybrid cloud.…
If there’s one thing my interactions with our customers has taught me, it’s that Apache Hadoop didn’t disrupt the datacenter, the data did. The explosion of new types of data in recent years has put tremendous pressure on the datacenter, both technically and financially, and an architectural shift is underway where Enterprise Hadoop is playing a key role in the resulting modern data architecture.
Due to the flourish of Apache Software Foundation projects that have emerged in recent years in and around the Apache Hadoop project, a common question I get from mainstream enterprises is: What is the definition of Hadoop?
This question goes beyond the Apache Hadoop project itself, since most folks know that it’s an open source technology borne out of the experience of web scale consumer companies such as Yahoo!, Facebook and others who were confronted with the need to store and process massive quantities of data.…
We love to hear examples from the ecosystem of how organizations are benefiting from Hadoop and today Hortonworks partner Microsoft posted a great detailed case study on how one of their partners – Ascribe – is using Microsoft’s HDInsight Service, their cloud based 100% Apache Hadoop service to transform healthcare in the UK.
Ascribe is a UK based company focused on solutions for the healthcare industry and was an early adopter of HDInsight which is built using the Hortonworks Data Platform.…
We are delighted to host this is a guest blog from John Schitka at SAP.
Join us on March 12 to learn how SAP HANA and Hortonworks Data Platform combine to help you achieve Instant Insight and Infinite Scale — Register Here
Big Data is changing our world – enabling previously impossible insights and transforming the way we do business, work with others, and live our lives. To be competitive you need to lever Big Data and the business value it brings.…
Elasticsearch’s engine integrates with Hortonworks Data Platform 2.0 and YARN to provide real-time search and access to information in Hadoop.
See it in action: register for the Hortonworks and Elasticsearch webinar on March 5th 2014 at 10 am PST/1pm EST to see the demo and an outline for best practices when integrating Elasticsearch and HDP 2.0 to extract maximum insights from your data. Click here to register for this exciting and informative webinar!…
This is the fifth in our series on modern data architectures across industry verticals. Others in the series are:
- Modern Healthcare Architectures Built with Hadoop
- Modern Manufacturing Architectures Built with Hadoop
- Modern Telecom Architectures Built with Hadoop
- Modern Retail Architectures Built with Hadoop
Consumers have never generated so much data on how they research, discuss and buy products. This new data is valuable for shaping and promoting a brand or product, but it doesn’t line up neatly to fit in pre-defined, tabular formats.…
Hadoop can be a great complement to existing data warehouse platforms, such as Teradata, as it naturally helps to address two key storage challenges:
- Managing large volumes of historical or archival data.
- Handling data from non-standard or un-structured sources
The purpose of this article is to detail some of the key integration points and to show how data can be easily exchanged for enrichment between the two platforms.
As a data integrator who is familiar with RDBMS systems and is new to the Hadoop platform, I was looking for a simple way (i.e.…