There will be 4 tracks, Operations, Features and Internals, Ecosystem and Case Studies. The keynotes will include speakers from Cloudera who is the event host, Google BigTable team as a follow up to their ‘06 BigTable paper, Salesforce on their experience with HBase operations and use cases and Facebook on their strongly consistent multi data center replication scheme.…
The Hortonworks Blog
As the Red Hat Summit shifts to the west coast in San Francisco this year Hortonworks and Red Hat will be demonstrating the progress of our engineering efforts. Our engineers have been hard at work in the factories and in the communities deeply integrating our open source offerings to create a comprehensive platform for new analytic applications. As a reminder in February Red Hat and Hortonworks announced a comprehensive open source initiative to deliver infrastructure solutions to bring 100-percent open source Hadoop to the hybrid cloud.…
The community fixed 411 JIRAs for 2.4.0 (on top of the 511 JIRAs resolved for 2.3.0). Of the 411 fixes:
- 50 are in Hadoop Common,
- 171 are in HDFS,
- 160 are in YARN and
- 30 went into MapReduce
Hadoop 2.4.0 is the second Hadoop release in 2014, following Hadoop 2.3.0’s February release and its key enhancements to HDFS such as Support for Heterogeneous Storage and In-Memory Cache.…
Apache Tez is an alternative to MapReduce that provides a powerful framework for executing a complex topology of tasks for data access in Hadoop. Version 0.4 incorporates the feedback from extensive testing of Tez 0.3, released just last month.
This release is especially meaningful because it coincides with completion of the Stinger Initiative (a collaborative community effort involving 145 developers across 44 companies) and the upcoming release of Apache Hive 0.13.…
Securing any system requires you to implement layers of protection. Access Control Lists (ACLs) are typically applied to data to restrict access to data to approved entities. Application of ACLs at every layer of access for data is critical to secure a system. The layers for hadoop are depicted in this diagram and in this post we will cover the lowest level of access… ACLs for HDFS.
This is part of the HDFS Developer Trail series. …
Today we are proud to announce that the formation of a terrific partnership with LucidWorks to bring search to the Hortonworks Data Platform. LucidWorks delivers an enterprise-grade search development platform built atop the power of Apache Solr.
Shared Vision and New Scenarios
Both LucidWorks and Hortonworks have a shared vision of innovating in open source and delivering it to customers in an enterprise grade platform.
As part of our continuing mission to build the a completely open, versatile enterprise data platform across many data processing scenarios then Solr offers a simple, yet powerful interface providing advanced search capabilities.…
If you’re excited to get started with the new features in Hortonworks Data Platform 2.1, then we’ve included 4 tutorials for you try out – Sandbox-style.
You can download the HDP 2.1 Technical Preview here, and then get stuck into these great tutorials.
Interactive Query with Apache Hive and Apache Tez
OK, so you’re not going to get huge performance out of a one-node VM, but you can try out Hive on Tez, and see the performance gains versus MapReduce, and also try out features such as Vectorized Query, and the host of new SQL features.…
The pace of innovation within the Apache Hadoop community is truly remarkable, enabling us to announce the availability of Hortonworks Data Platform 2.1, incorporating the very latest innovations from the Hadoop community in an integrated, tested, and completely open enterprise data platform.
What’s In Hortonworks Data Platform 2.1?
The advancements in HDP 2.1 span every aspect of Enterprise Hadoop: from data management, data access, integration & governance, security and operations. …
Apache Falcon is a data governance engine that defines, schedules, and monitors data management policies. Falcon allows Hadoop administrators to centrally define their data pipelines, and then Falcon uses those definitions to auto-generate workflows in Apache Oozie.
InMobi is one of the largest Hadoop users in the world, and their team began the project 2 years ago. At the time, InMobi was processing billions of ad-server events in Hadoop every day.…
We are excited to welcome Blackrock and Passport Capital as Hortonworks investors who today led a $100M round of funding with continued participation from all existing investors.
This latest round of funding will allow us to double-down on our founding strategy: to make open source Apache Hadoop a true enterprise data platform. To that end we are focused in two areas:…
1. Lead the innovation of Hadoop. In open source, for everyone.
In February 2014, the Apache Storm community released Storm version 0.9.1. Storm is a distributed, fault-tolerant, and high-performance real-time computation system that provides strong guarantees on the processing of data. Hortonworks is already supporting customers using this important project today.
Many organizations have already used Storm, including our partner Yahoo! This version of Apache Storm (version 0.9.1) is:
- Highly scalable. Like Hadoop, Storm scales linearly
- Fault-tolerant. Automatically reassigns tasks if a node fails
LDAP provides a central source for maintaining users and groups within an enterprise. There are two ways to use LDAP groups within Hadoop. The first is to use OS level configuration to read LDAP groups. The second is to explicitly configure Hadoop to use LDAP-based group mapping.
Here is an overview of steps to configure Hadoop explicitly to use groups stored in LDAP.
- Create Hadoop service accounts in LDAP
- Shutdown HDFS NameNode & YARN ResourceManager
- Modify core-site.xml to point to LDAP for group mapping
- Re-start HDFS NameNode & YARN ResourceManager
- Verify LDAP based group mapping
Prerequisites: Access to LDAP and the connection details are available.…
Hortonworks would like to congratulate Leslie Lamport on winning the 2013 Turing Award given by the Association of Computing Machinery. This award is essentially the equivalent of the Nobel Prize for computer science. Among Lamport’s many and varied contributions to the field computer science are: TLA (Temporal Logic for Actions), LaTeX and PAXOS.
Due to the flourish of Apache Software Foundation projects that have emerged in recent years in and around the Apache Hadoop project, a common question I get from mainstream enterprises is: What is the definition of Hadoop?
This question goes beyond the Apache Hadoop project itself, since most folks know that it’s an open source technology borne out of the experience of web scale consumer companies such as Yahoo!, Facebook and others who were confronted with the need to store and process massive quantities of data.…
The HBase BlockCache is an important structure for enabling low latency reads. As of HBase 0.96.0, there are no less than three different BlockCache implementations to choose from. But how to know when to use one over the other? There’s a little bit of guidance floating around out there, but nothing concrete.…