Hadoop Developer Resume
SUMMARY:
- Highly acumen and experienced IT professional with 5+ years Hadoop Developer in Big Data/Hadoop technology development.
- Experience working with Cloudera & Hortonworks Distribution of Hadoop.
- Expertise in HDFS, MapReduce, Hive, Pig, Sqoop, HBase, Flume, Zookeeper and hadoop ecosystem.
- Experience in analyzing the different types of data that flow from data lakes to Hadoop Clusters.
- Worked with Apache Hadoop along enterprise version of Cloudera and Hortonworks. Good knowledge on MAPR distribution.
- Hands on experience with Big Data Hadoop core and Eco - System components (HDFS, MR1, MR2, Yarn, Hive, Impala, Beeline, Sqoop, Flume, Oozie, Hbase, Zookeeper and Pig).
- Experience in manipulating the streaming data to clusters through Flume.
- Proficient in working with NoSQL database like MongoDB and HBase.
- Experience in partitioning the Big Data according the business requirements using Hive Indexing, partitioning and Bucketing.
- Working with data extraction, transformation and load in Hive, Pig and HBase.
- Working with data transformation from HDFS, HIVE, PIG, HBase, and MySQL.
- Experience in creating UDF's, UDAF's for Hive and Pig.
- Optimized streaming log files with no time latency using Flume and more importantly operating the data down stream flow to Hadoop ecosystems and it analysis segments.
- Profound experience in working with Cloudera CDH 5.x on multi-node cluster.
- Acumen in choosing an efficient ecosystem in Hadoop and providing the best solutions to Big Data problems.
- Good Knowledge in Spark and Scala.
- Developed POC in Spark and Scala.
- Experience in reporting analyzed data in vivid formats using reporting tool Tableau.
- Prolific in generating the splendid and informative dashboards for Business Intelligence teams.
- Experience in using design pattern, Java, Servlets, JSP, JavaScript, HTML, JQuery, Angular JS, Mobile JQuery, XML, Web Logic, JBOSS 4.2.3, SQL, PL/SQL, JUnit, and Apache-Tomcat, Linux.
- Expertise in relational databases like Oracle, My SQL and SQL Server.
- Experience in Agile methodologies.
- Proficient communication skills with an ability to lead a team & keep them motivated.
- Extensive experience with Java complaint IDE's like Eclipse.
- Adept in handling the team in untoward situations and capable of sailing the team to deliver the quality output.
- Highly motivated and versatile team player with the ability to work independently & adapt quickly to new emerging technologies.
TECHNICAL SKILLS:
Big Data: Hadoop, Storm, Trident, Hbase, Hive, Flume, CassandraKafka,Storm, Sqoop, Oozie, PIG, SparkMapreduce,Zookeeper, Yarn.
Operating Systems: UNIX, Mac, Linux, Windows 2000 / NT / XP / VistaAndroid
Programming Languages: Java (JDK 5/JDK 6), C/C++, Mat lab, R, HTML, SQLPL/SQL
Frameworks: Hibernate 2.x/3.x, Spring 2.x/3.x,Struts 1.x/2.x and JPA
Web Services: WSDL, SOAP, Apache CXF/XFire, Apache Axis, RESTJersey
Databases/technologies: Oracle 8i/9i/10g, Microsoft SQL Server, DB2 & MySQL 4.x/5.x
Middleware Technologies: Web sphere Message Queue, Web sphere Message Broker, XML gateway, JMS
Web Technologies: J2EE, Soap & REST Web Services, JSP, Servlets, EJBJavaScript, Struts, Spring,Web works, Direct Web remoting, HTML, XML,JMS, JSF, Ajax.
Testing Frameworks: Mockito, PowerMock, EasyMock.
Web/Application Servers: IBM Web sphere Application server, Jboss, Apache Tomcat.
Others: SoftwareBorland Star team, Clear case, Junit, ANTMaven, Android Platform,Microsoft Office, SQL DeveloperDB2 control center, Microsoft Visio, Hudson, Subversion, GIT, Nexus, Artifactory and Trac.
Developement Strategies: Agile, Lean Agile, Pair Programming, Water-Fall and Test Driven Development
PROFESSIONAL EXPERIENCE:
Hadoop Developer
Confidential
Responsibilities:
- Importing and exporting data into HDFS and Hive using Sqoop and Kafka.
- Develop different components of system like Hadoop process that involves Map Reduce, and Hive.
- Developed interface for validating incoming data into HDFS before kicking off Hadoop process.
- Written hive queries using optimized ways like user-defined functions, customizing Hadoop shuffle & sort parameters.
- Worked on tuning Hive and Pig to improve performance and solve performance related issues in Hive and Pig scripts with good understanding of Joins, Group and aggregation and how it does Map Reduce jobs.
- Developing map reduce programs for different types of Files using Combiners with UDF's and UDAF's.
- Experience working on multiple node cluster tool which offer several commands to return HBase usage.
- Experience in creating tables, dropping and altered at run time without blocking updates and queries using HBase and Hive.
- Experience on pre-processing the logs and semi structured content stored on HDFS using PIG.
- Experience in structured data imports and exports into Hive warehouse which enables business analysts to write Hive queries.
- Experience in managing and reviewing Hadoop log files.
- Experience on Unix shell scripts for business process and loading data from different interfaces to HDFS.
- Involved in creating Hive tables, Pig tables, and loading data and writing hive queries and pig scripts.
- Hands on experience in eclipse, Putty, winSCP, VNCviewer, etc.
Environment: Linux 6.7, CDH5.5.2, MapReduce, Hive 1.1, PIG, HBase, Shell Script, SQOOP 1.4.3, Eclipse, Java 1.8.
Hadoop Developer
Confidential
Responsibilities:
- Worked on analyzing Hadoop stack and different big data analytic tools including Pig,
- Hive, Hbase database and Sqoop.
- Experienced to implement Hortonworks distribution system (HDP 2.1, HDP 2.2 and HDP 2.3).
- Developed Map Reduce programs for some refined queries on big data.
- Experienced in working with Elastic MapReduce (EMR).
- Creating Hive tables and working on them for data analysis to cope up with the requirements.
- Developed a frame work to handle loading and transform large sets of unstructured data from UNIX system to HIVE tables.
- Worked with business team in creating Hive queries for ad hoc access.
- In depth understanding of Classic Map Reduce and YARN architectures.
- Implemented Hive Generic UDF's to implement business logic.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Installed and configured Pig for ETL jobs.
- Developed Pig UDF's to pre-process the data for analysis.
- Deployed Cloudera Hadoop Cluster on AWS for Big Data Analytics
- Analyzed the data by performing Hive queries, ran Pig scripts, Spark SQL and Spark
- Streaming.
- Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
- Used Apache NiFi to copy the data from local file system to HDFS.
- Developed Spark Streaming script which consumes topics from distributed messaging source Kafka and periodically pushes batch of data to Spark for real time processing.
- Extracted files from Cassandra through Sqoop and placed in HDFS for further processing.
- Involved in creating generic Sqoop import script for loading data into Hive tables from
- RDBMS.
- Involved in continuous monitoring of operations using Storm.
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
- Implemented indexing for logs from Oozie to Elastic Search.
- Design, develop, unit test, and support ETL mappings and scripts for data marts using Talend.
Environment: Hortonworks, Hadoop, Map Reduce, HDFS, Hive, Pig, Sqoop, Apache Kafka, Apache Storm, Oozie, SQL, Flume, Spark, Hbase, Cassandra, Informatica, Java, Github.
Hadoop Developer
Confidential
Responsibilities:
- Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop the application.
- Ingested the data from external data sources like MySQL using Sqoop and loaded data into HDFS.
- Exported the analyzed data to the relational databases using Sqoop for further visualization and to generate reports for the BI team.
- Worked on Linux shell scripts for business processes and with loading the data from different systems to the HDFS.
- Configured Pig and also designed Pig Latin scripts to process the data into a universal data model.
- Used Pig Latin scripts to convert data from JSON, XML and other formats to Avro file format.
- Wrote Custom UDFs in PIG to process and perform business intelligence on the data.
- Created partitioned and bucketed tables in Hive based on the hierarchy of the dataset.
- Involved in creating Hive internal and external tables, loading them with data and writing hive queries which require multiple join scenarios.
- Analyzed, transformed, filtered, Co-Grouped and aggregated data with HiveQL.
- Imported the transformed data into Talend to generate visualizations for further analysis by Business analysts.
- Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
- Developed Python scripts to analyze the data in the HDFS.
- Distributed snapshot Jar's across the cluster in the distributed cache to increase the performance and efficiency of the algorithms by reducing the shuffling of data.
- Used Cassandra node tool to manage Cassandra cluster.
- Worked on to configure a Size Tiered Compaction Strategy (STCS) compaction strategy for Cassandra.
- Processed the data with map reduce and staged end results to Cassandra.
- Analyzed the performance of the cluster and related resources to identify bottlenecks.
- Performed unit testing to meet the functional, technical and business requirements.
- Performed Integration testing of software and firmware to ensure that the product health is optimal.
Environment: Scala, MapReduce, HDFS, Sqoop, Hive, YARN, PIG, Oozie, Python, Shell Scripting, Git, Linux, Maven, DataStax Cassandra, AVRO.
Hadoop Developer
Confidential
Responsibilities:
- Evaluated Spark's performance vs Impala on transactional data.
- Used Spark transformations and aggregations using Python and Scala to perform min,max and average on transactional data.
- Experienced in migrating data from HiveQL to SparkSQL using Scala.
- Knowledge in using Spark Data-frames to load data in Spark Data-frames.
- Knowledge on handling Hive queries using Spark SQL that integrate with Sparkenvironment.
- Used java to develop Restful API for database Utility Project.
- Responsible for performing extensive data validation using Hive.
- Designed a data model in Cassandra (POC) for storing server performance data.
- Implemented a Data service as a rest API project to retrieve server utilization data from this Cassandra Table.
- Implemented Python script to call the Cassandra Rest API, performed transformations and loaded the data into Hive.
- Designed data model to ingest transactional data with and without URIs into Cassandra.
- Implemented shell script to call python script to perform min, max and average on utilization data of 1000s hosts and compared the performance on various levels of summarization.
- Involved in creating Oozie workflow and Coordinator jobs for Hive jobs to kick off the jobs on time for data availability.
- Generated reports from this hive table for visualization purpose.
- Migrated HiveQL to SparkSQL to validate Spark's performance with Hive's.
- Implemented Proof of concept for Dynamo DB, Redshift and EMR
- Proactively researched on Microsoft Azure.
- Presented Demo on Microsoft Azure, an overview of cloud computing with Azure.
Environment: Hadoop, Azure, AWS, HDFS, Hive, Hue, Oozie, Java, Linux, Cassandra, Python,Open TSDB, Scala
