Table Of ContentElasticsearch for Hadoop
Integrate Elasticsearch into Hadoop to effectively
visualize and analyze your data
Vishal Shukla
BIRMINGHAM - MUMBAI
Elasticsearch for Hadoop
Copyright © 2015 Packt Publishing
All rights reserved. No part of this book may be reproduced, stored in a retrieval
system, or transmitted in any form or by any means, without the prior written
permission of the publisher, except in the case of brief quotations embedded in
critical articles or reviews.
Every effort has been made in the preparation of this book to ensure the accuracy
of the information presented. However, the information contained in this book is
sold without warranty, either express or implied. Neither the author, nor Packt
Publishing, and its dealers and distributors will be held liable for any damages
caused or alleged to be caused directly or indirectly by this book.
Packt Publishing has endeavored to provide trademark information about all of the
companies and products mentioned in this book by the appropriate use of capitals.
However, Packt Publishing cannot guarantee the accuracy of this information.
First published: October 2015
Production reference: 1201015
Published by Packt Publishing Ltd.
Livery Place
35 Livery Street
Birmingham B3 2PB, UK.
ISBN 978-1-78528-899-9
www.packtpub.com
Credits
Author Project Coordinator
Vishal Shukla Suzanne Coutinho
Reviewers Proofreader
Vincent Behar Safis Editing
Elias Abou Haydar
Yi Wang Indexer
Mariammal Chettiyar
Acquisition Editors
Vivek Anantharaman Graphics
Disha Haria
Larissa Pinto
Production Coordinator
Content Development Editor
Nilesh R. Mohite
Pooja Mhapsekar
Cover Work
Technical Editor
Nilesh R. Mohite
Siddhesh Ghadi
Copy Editor
Relin Hedly
About the Author
Vishal Shukla is the CEO of Brevitaz Systems (http://brevitaz.com) and a
technology evangelist at heart. He is a passionate software scientist and a big data
expert. Vishal has extensive experience in designing modular enterprise systems.
Since his college days (more than 11 years), Vishal has enjoyed coding in JVM-based
languages. He also embraces design thinking and sustainable software development.
He has vast experience in architecting enterprise systems in various domains. Vishal
is deeply interested in technologies related to big data engineering, analytics, and
machine learning.
He set up Brevitaz Systems. This company delivers massively scalable and
sustainable big data and analytics-based enterprise applications to their global
clientele. With varied expertise in big data technologies and architectural acumen,
the Brevitaz team successfully developed and re-engineered a number of legacy
systems to state-of-the-art scalable systems. Brevitaz has imbibed in its culture agile
practices, such as scrum, test-driven development, continuous integration, and
continuous delivery, to deliver high-quality products to its clients.
Vishal is a music and art lover. He loves to sing, play musical instruments, draw
portraits, and play sports, such as cricket, table tennis, and pool, in his free time.
You can contact Vishal at [email protected] and on LinkedIn at
https://in.linkedin.com/in/vishalshu. You can also follow Vishal on Twitter
at @vishal1shukla2.
I would like to express special thanks to my beloved wife, Sweta
Bhatt Shukla, and my awaited baby for always encouraging me
to go ahead with the book, giving me company during late nights
throughout the write-up, and not complaining about my lack of
time. My hearty thanks to my sister Krishna Meet Bhavesh Shah and
Arpit Panchal for their detailed reviews, invaluable suggestions, and
assistance. Heartfelt thanks to my brother and idol, Pranav Shukla,
and my family and friends for their continued support and guidance.
I would also like to express my gratitude to my mentors and
colleagues, whose interaction transformed me into what I am today.
Though everyone's contribution is vital, it is not possible to name
them all, so to name a few, I would like to thank Thomas Hirsch,
Sven Boeckelmann, Abhay Chrungoo, Nikunj Parmar, Kuntal Shah,
Vinit Yadav, Kruti Shukla, Brett Connor, and Lovato Claiton.
About the Reviewers
Vincent Behar is a passionate software developer. He has worked on a search
engine, indexing 16 billion web pages. In this big data environment, the tools
he used were Hadoop, MapReduce, and Cascading. Vincent has also worked
with Elasticsearch in a large multitenant setup, with both ELK stack and specific
indexing/searching requirements. Therefore, bringing these two technologies
together, along with new frameworks such as Spark, was the next natural step.
Elias Abou Haydar is a data scientist at iGraal in Paris, France. He obtained an
MSc in computer science, specializing in distributed systems and algorithms, from
the University of Paris Denis Diderot. He was a research intern at LIAFA, CNRS,
working on distributed graph algorithms in image segmentation applications.
He discovered Elasticsearch during his end-of-course internship and has been
passionate about it ever since.
Yi Wang is currently a lead software engineer at Trendalytics, a data analytics start-
up. He is responsible for specifying, designing, and implementing data collection,
visualization, and analysis pipelines. He holds a master's degree in computer science
from Columbia University and a master's degree in physics from Peking University,
with a mixed academic background in math, chemistry, and biology.
www.PacktPub.com
Support files, eBooks, discount offers, and more
For support files and downloads related to your book, please visit www.PacktPub.com.
Did you know that Packt offers eBook versions of every book published, with PDF
and ePub files available? You can upgrade to the eBook version at www.PacktPub.
com and as a print book customer, you are entitled to a discount on the eBook copy.
Get in touch with us at [email protected] for more details.
At www.PacktPub.com, you can also read a collection of free technical articles, sign
up for a range of free newsletters and receive exclusive discounts and offers on Packt
books and eBooks.
TM
https://www2.packtpub.com/books/subscription/packtlib
Do you need instant solutions to your IT questions? PacktLib is Packt's online digital
book library. Here, you can search, access, and read Packt's entire library of books.
Why subscribe?
• Fully searchable across every book published by Packt
• Copy and paste, print, and bookmark content
• On demand and accessible via a web browser
Free access for Packt account holders
If you have an account with Packt at www.PacktPub.com, you can use this to access
PacktLib today and view 9 entirely free books. Simply use your login credentials for
immediate access.
Table of Contents
Preface ix
Chapter 1: Setting Up Environment 1
Setting up Hadoop for Elasticsearch 1
Setting up Java 2
Setting up a dedicated user 2
Installing SSH and setting up the certificate 3
Downloading Hadoop 3
Setting up environment variables 4
Configuring Hadoop 5
Configuring core-site.xml 5
Configuring hdfs-site.xml 6
Configuring yarn-site.xml 6
Configuring mapred-site.xml 6
The format distributed filesystem 7
Starting Hadoop daemons 7
Setting up Elasticsearch 8
Downloading Elasticsearch 8
Configuring Elasticsearch 8
Installing Elasticsearch's Head plugin 10
Installing the Marvel plugin 10
Running and testing 11
Running the WordCount example 12
Getting the examples and building the job JAR file 12
Importing the test file to HDFS 12
Running our first job 13
Exploring data in Head and Marvel 15
Viewing data in Head 15
Using the Marvel dashboard 17
Exploring the data in Sense 18
Summary 19
[ i ]