Hadoop Developer Resume
Houston, TX
SUMMARY:
- Around 8 years of overall IT development experience including 4 years of experience exclusively on BIG DATA ECOSYSTEM using HADOOP framework and related technologies such as HDFS, MapReduce, HIVE, PIG, HBASE, FLUME, OOZIE, SQOOP, and ZOOKEEPER.
- Excellent knowledge on distributed storages (HDFS) and distributed processing (MapReduce, Yarn) for real - time streaming and batch processing.
- Experience in developing Map-Reduce programs to perform Data Transformation in Java.
- Experience in writing Custom MapReduce programs in java and also extending Hive and Pig core functionality by writing custom UDFs.
- Extensive experience with big data query tools like Pig Latin and HiveQL.
- Experience in extracting the data from RDBMS into HDFS using Sqoop.
- Experience in collecting the logs from log collector into HDFS using Flume.
- Good understanding of NoSQL databases such as HBase, Cassandra and Mongo DB.
- Experience in analyzing data in HDFS through MapReduce, Hive and Pig.
- Clear knowledge of rack awareness topology in the Hadoop cluster.
- Experience in job workflow scheduling and monitoring tools like Oozie and Zookeeper.
- Knowledge on Hadoop administration activities such as installation, configuration and management of clusters using Cloudera Manager and Apache Ambari.
- Good knowledge on Hadoop HDFS architecture and MapReduce framework.
- Hands-on experience on Scala programming language.
- Hands on experience in loading unstructured data (Log files, Xml data) into HDFS using Flume.
- Good knowledge on Apache Spark, Storm, Kafka, Splunk and BI tools such as Pentaho and Talend.
- Hands on experience on performing ETL by using Talend.
- Experience in tuning the performances by using Partitioning, Bucketing and Indexing in HIVE.
- Experience in job/workflow scheduling and monitoring tools like Oozie and Zookeeper.
- Hands-on experience with test frameworks for Hadoop using MRUnit framework.
- Experience in writing Complex SQL Queries involving multiple tables inner and outer joins.
- Flexible with Unix/Linux and Windows Environments working with Operating Systems like Centos, Ubuntu.
TECHNICAL SKILLS:
Hadoop: HDFS, MapReduce, PIG, Hive, Sqoop, Zookeeper, Flume, Oozie
NoSQL: HBase, MongoDB, Cassandra
Java Technologies and Frameworks: J2EE, JSTL, JDBC, JSP, Java Servlets, Struts, Spring, Hibernate
Languages: C, C++, Java, Python
Web Services: XML, SOAP, REST
Web Technologies: JavaScript, CSS, CSS3,HTML, HTML5, Bootstrap, XHTML, JQUERY, PHP
Databases: Oracle, DB2, MS-SQL Server, MySQL, MS-Access
Web Servers: Web Logic, Web Sphere, Apache Tomcat.
Modeling Tools: UML on Rational Rose, Rational Clear Case, Enterprise Architect, Microsoft Visio
IDE Development Tools: Eclipse, Net Beans, IntelliJ
Build Tools: Maven, Scala Build Tool(SBT), Ant
Operating systems: Linux (Red Hat, Ubuntu, Centos).
PROFESSIONAL EXPERIENCE:
Confidential, Houston, TX
Hadoop Developer
Responsibilities:
- Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.
- Built scalable distributed data solutions using Hadoop.
- Developed Simple to complex MapReduce Jobs using Hive and Pig.
- Experienced in defining job flows using Oozie
- Managed data coming from different sources and application
- Imported/exported data from RDBMS to HDFS using Data Ingestion tools like Sqoop.
- Optimized Map Reduce Jobs to use HDFS efficiently by using various compression mechanisms
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS.
- Imported the data from different sources like HDFS/HBase into Spark RDD.
- Experienced in implementing Spark RDD transformations, actions to implement business analysis.
- Migrated Hive QL queries on structured into Spark QL to improve performance.
- Installed Oozie workflow engine to run multiple Hive and Pig jobs which run independently.
- Worked on Kafka while dealing with raw data, by transforming into new Kafka topics for further consumption.
- Worked on large datasets to generate insights by using Splunk.
- Developed Splunk queries and dashboards targeted at understanding application performance and capacity analysis.
- Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
- Worked on NoSQL databases including HBase, MongoDB, and Cassandra.
- Worked on the Ad hoc queries, Indexing, Replication, Load balancing, Aggregation in MongoDB.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Managed and reviewed Hadoop log files.
- Involved in creating Hive tables, loading with data and writing hive queries which run internally in MapReduce.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Installed and configured Pig and also written Pig Latin scripts.
- Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
Environment: Hadoop, MapReduce, HDFS, Hive, Pig, Java, SQL, Spark,Splunk, Kafka, MongoDB, Cassandra, Sqoop, Nagios, Ganglia, Zookeeper,Java.
Confidential, St.Louis, MO
Hadoop Developer
Responsibilities:
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
- Written multiple Map Reduce programs in Java for data extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV and other compressed file formats.
- Involved in designing schema, writing CQL's and loading data using Cassandra
- Good experience with CQL Data manipulation commands and CQL clauses
- Worked with CQL collections
- Installed and configured Hive and also written Hive UDFs.
- Involved and experienced with Datastax.
- Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way
- Developed Map Reduce jobs to automate transfer of data from/to Hbase
- Assisted with the addition of Hadoop processing to the IT infrastructure
- Worked in creation of indexes and increased the search results very faster
- Used flume to collect the entire web log from the online ad-servers and push into HDFS
- Implemented Map/Reduce job and execute the MapReduce job to process the log data from the ad-servers.
- Wrote efficient map reduce code to aggregate the log data from the Ad-server
- Mainly worked on Hive queries to categorize data of different claims.
- Integrated the Hive warehouse with HBase.
- Involved in loading data from LINUX file system to HDFS.
- Written customized Hive UDFs in Java.
- Implemented Partitioning, Dynamic Partitions, Buckets in Hive.
- Responsible to manage the test data coming from different sources Reviewing peer table creation in Hive, data loading and queries.
- Monitored System health and logs and respond accordingly to any warning or failure conditions.
- Gained experience in managing and reviewing Hadoop log files.
- Involved in scheduling Oozie workflow engine to run multiple Hive and Pig jobs involved unit testing, interface testing, system testing and user acceptance testing of the workflow tool.
- Used Spark API over Hadoop to perform analytics on data in Hive
- Explored with Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark context, Spark-SQL, Data Frame, Spark YARN
- Imported and exported data into HDFS and Hive using Sqoop.
- Developed Spark Programs for Batch and Real time Processing
- Developed Spark Streaming applications for Real time Processing
- Worked under UNIX environment in development of application using Python and familiar with all of its commands
- Performed transformations, cleaning and filtering on imported data using Hive, Map Reduce, and loaded final data into HDFS.
Environment: Hadoop, HDFS, Hive, Map Reduce, Java, Pig, Oracle, MySQL,Spark, Flume,Spark Streaming, Sqoop, Cassandra
Confidential
Hadoop Developer
Responsibilities:
- Written MapReduce code for processing and parsing the data from various sources and storing parsed data into HBase and Hive using HBase-Hive Integration.
- Worked with HBase and Hive scripts to extract, transform and load data into HBase and Hive.
- Worked on moving all log files generated from various sources to HDFS for further processing.
- Developed workflows using custom MapReduce, Pig, Hive, and Sqoop.
- Tuned the cluster for optimal performance to process these large data sets.
- Worked hands on with ETL process. Handled importing data from various data sources, performed transformations.
- Built reusable Hive UDF libraries for business requirements which enabled users to use these UDFs in Hive Querying.
- Loaded the created HFiles into HBase for faster access of large customer base without taking performance hit.
- Written Hive UDF to sort Structure fields and return complex data type.
- Responsible for loading data from UNIX file system to HDFS.
- Developed suit of Unit Test Cases for Mapper, Reducer and Driver classes using MR Testing library.
- Used Maven extensively for building jar files of MapReduce programs and deployed to Cluster.
- Modelled Hive partitions extensively for data separation and faster data processing and followed.
- Work with network and Linux system engineers to define optimum network configurations, server hardware and operating system.
- Evaluate and propose new tools and technologies to meet the needs of the organization.
- Production support responsibilities include cluster maintenance.
- Pig and Hive best practices for tuning.
- Gained good experience with NOSQL database.
Environment: CDH, Hive, MySQL, HBase, HDFS, HIVE, Eclipse, Hadoop, Oracle, PL/SQL, SQL*PLUS, Toad 9.6, Flume, PIG, Sqoop, UNIX.
Confidential
Java Developer
Responsibilities:
- Develop GUI related changes using JSP, HTML and client validations using JavaScript.
- Designed and developed front end using HTML, JSP and Servlets
- Implemented client side validation using JavaScript
- Used Hibernate in persistence layer of the application
- Created UML class diagrams that depict the code’s design and its compliance with the functional requirements.
- Developed user interface using JSP to simplify the complexities of the application.
- Developed the Web Interface using Servlets, Java Server Pages, HTML and CSS.
- Extensively used the JDBC Prepared Statement to embed the SQL queries into the java code.
- Implemented the DAO pattern.
- Developed business logic using Stateless session beans for calculating asset depreciation on Straight line and written down value approaches.
- Involved coding SQL Queries, Stored Procedures and Triggers.
- Created java classes to communicate with database using JDBC.
- Developed user interface using JSP, JSP Tag libraries, and JavaScript to simplify the complexities of the application.
- Provided request and reports and Supporting Client Services for all customer requests being an SME of the application Experience developing for Unix/Linux based systems Development of Tools and Value adds to assist Performance Testing.
- Followed agile methodology and SCRUM meetings to track, optimize and tailored features to customer needs
Environment: Java 1.4, Servlets, JSP, EJB, J2EE 1.4, XML, XSLT, Java Script, SQL, PL/SQL, MS Visio, Eclipse, JDBC, Win CVS, Windows XP.
