Today we are proud to announce that the formation of a terrific partnership with LucidWorks to bring search to the Hortonworks Data Platform. LucidWorks delivers an enterprise-grade search development platform built atop the power of Apache Solr.Presentation & Applications Enable both existing and new applications to provide value to the organization. Enterprise Management & Security Empower existing operations and security tools to manage Hadoop. Governance Integration Data Workflow, Lifecycle & Governance
The Hortonworks Blog
If you’re excited to get started with the new features in Hortonworks Data Platform 2.1, then we’ve included 4 tutorials for you try out – Sandbox-style.
You can download the HDP 2.1 Technical Preview here, and then get stuck into these great tutorials.Interactive Query with Apache Hive and Apache Tez
OK, so you’re not going to get huge performance out of a one-node VM, but you can try out Hive on Tez, and see the performance gains versus MapReduce, and also try out features such as Vectorized Query, and the host of new SQL features.…
The pace of innovation within the Apache Hadoop community is truly remarkable, enabling us to announce the availability of Hortonworks Data Platform 2.1, incorporating the very latest innovations from the Hadoop community in an integrated, tested, and completely open enterprise data platform.
There is no doubt that enterprises recognize how Big Data is crucial to monetizing their business. The information contained in the volumes of data collected can offer key insights into product, customer and competitive trends. There are a variety of sophisticated tools for Big Data analytics and processing but most big data implementations are based on rudimentary technologies like FTP based scripts for data collection and distribution.
Although FTP is a widely used protocol, there is an inherent lack of reliability in this approach. …
Apache Falcon is a data governance engine that defines, schedules, and monitors data management policies. Falcon allows Hadoop administrators to centrally define their data pipelines, and then Falcon uses those definitions to auto-generate workflows in Apache Oozie.
InMobi is one of the largest Hadoop users in the world, and their team began the project 2 years ago. At the time, InMobi was processing billions of ad-server events in Hadoop every day.…
We are excited to welcome Blackrock and Passport Capital as Hortonworks investors who today led a $100M round of funding with continued participation from all existing investors.
This latest round of funding will allow us to double-down on our founding strategy: to make open source Apache Hadoop a true enterprise data platform. To that end we are focused in two areas:1. Lead the innovation of Hadoop. In open source, for everyone.…
In February 2014, the Apache Storm community released Storm version 0.9.1. Storm is a distributed, fault-tolerant, and high-performance real-time computation system that provides strong guarantees on the processing of data. Hortonworks is already supporting customers using this important project today.
Many organizations have already used Storm, including our partner Yahoo! This version of Apache Storm (version 0.9.1) is:
- Highly scalable. Like Hadoop, Storm scales linearly
- Fault-tolerant. Automatically reassigns tasks if a node fails
LDAP provides a central source for maintaining users and groups within an enterprise. There are two ways to use LDAP groups within Hadoop. The first is to use OS level configuration to read LDAP groups. The second is to explicitly configure Hadoop to use LDAP-based group mapping.
Here is an overview of steps to configure Hadoop explicitly to use groups stored in LDAP.
- Create Hadoop service accounts in LDAP
- Shutdown HDFS NameNode & YARN ResourceManager
- Modify core-site.xml to point to LDAP for group mapping
- Re-start HDFS NameNode & YARN ResourceManager
- Verify LDAP based group mapping
Prerequisites: Access to LDAP and the connection details are available.…
Hortonworks would like to congratulate Leslie Lamport on winning the 2013 Turing Award given by the Association of Computing Machinery. This award is essentially the equivalent of the Nobel Prize for computer science. Among Lamport’s many and varied contributions to the field computer science are: TLA (Temporal Logic for Actions), LaTeX and PAXOS.
We’ve said many times that in order for Apache Hadoop to succeed as the next generation data platform and power the Modern Data Architecture there must be a thriving ecosystem of enterprise vendors around it. It’s been core to our strategy from day one to foster these ecosystem vendors and work collaboratively to make them successful.
Our strategy of delivering enterprise Hadoop as 100% open source has resulted in close alignment and tight partnerships with a broad Hadoop ecosystem with vendors large and small including major data management leaders like SAP and Teradata.…
If there’s one thing my interactions with our customers has taught me, it’s that Apache Hadoop didn’t disrupt the datacenter, the data did. The explosion of new types of data in recent years has put tremendous pressure on the datacenter, both technically and financially, and an architectural shift is underway where Enterprise Hadoop is playing a key role in the resulting modern data architecture.
Due to the flourish of Apache Software Foundation projects that have emerged in recent years in and around the Apache Hadoop project, a common question I get from mainstream enterprises is: What is the definition of Hadoop?
This question goes beyond the Apache Hadoop project itself, since most folks know that it’s an open source technology borne out of the experience of web scale consumer companies such as Yahoo!, Facebook and others who were confronted with the need to store and process massive quantities of data.…
The HBase BlockCache is an important structure for enabling low latency reads. As of HBase 0.96.0, there are no less than three different BlockCache implementations to choose from. But how to know when to use one over the other? There’s a little bit of guidance floating around out there, but nothing concrete.…
This is the seventh in our series on modern data architectures across industry verticals. Others in the series are:
- Modern Healthcare Architectures Built with Hadoop
- Modern Manufacturing Architectures Built with Hadoop
- Modern Telecom Architectures Built with Hadoop
- Modern Retail Architectures Built with Hadoop
- Modern Advertising Architectures Built with Hadoop
- Modern Oil & Gas Architectures Built with Hadoop
Any financial services business cares about minimizing risk and maximizing opportunity. Banks weigh the risk of opening accounts versus the opportunity to hold deposits.…
Luminar is one of Hortonworks’ original customers. Apache Hadoop is a pillar of their modern data architecture, and since choosing Hortonworks in 2012, the Luminar team became expert users of Hortonworks Data Platform version 1.
They were eager to migrate to HDP2 after it launched in October 2013.
I recently spoke with Juan Manuel Alonso, Luminar’s Manager of Insights. Juan Manuel worked with the Hortonworks professional services team to plan and execute the migration from HDP1 to HDP2.…