We provide IT Staff Augmentation Services!

Sr. Hadoop Developer/administrator Resume

4.00/5 (Submit Your Rating)

Chicago, IllinoiS

SUMMARY:

  • Over 8+ years of experience in Hadoop/Big Data technologies such as in Hadoop, Pig, Hive, HBase, Oozie, Zookeeper, Sqoop, Storm, Flink, Flume, Impala, Tez, Kafka and Spark with hands on experience in writing Map Reduce/YARN and Spark/Scala jobs.
  • Expertise in working and designing of Row keys &Schema design with NOSQLdatabases like Mongo DB 3.0.1, HBase, Cassandra and DynamoDB (AWS).
  • Deep expertise in Analysis, Design, Development and Testing phases of Enterprise Data Warehousing solutions.
  • Expertise in Tableau BI reporting tools &Tableau Dashboards Developments &Server Administration.
  • Around 1 year of experience in Business Objects Desktop Intelligence, Web Intelligence, Universe Designer, Crystal Reports and Central Management Console.
  • Experience in data ingress and egress using Sqoop from HDFS to Relational Database Systems and vice - versa. Good knowledge of Log4j for error handling.
  • Expert knowledge in real time data analytics using Apache Storm.
  • Expertise in Java/J2EE technologies such as Core Java, spring, Hibernate, JDBC, JSON, HTML, Struts, Servlets, JSP, JBOSS and JavaScript.
  • Experience in designing both time driven and data driven automated workflows using Oozie.
  • Experience in migrating the data using Sqoop from Hadoop to Relational Database System and vice-versa.
  • Worked in Agile methodology of software development process as a Scrum Master.
  • Work experience inETLprocesses consisting of data sourcing, data transformation, mapping and loading of data from multiple source systems into Data Warehouse using Informatic Power Center.
  • Strong foundation in Programming, debugging skills, developed modules which have met with client requirements & targets.
  • Expertise in Hadoop administration such as managing cluster, reviewing Hadoop log files.
  • Expertise in Data warehousing concepts, Dimensional Modeling and Data Modeling systems.
  • Extensive Experience in developing test cases, performing Unit Testing and Integration Testing using source code management tools such as GIT, SVN, Perforce
  • Experience in using PL/SQL to write Stored Procedures, Functions and Triggers.
  • Have Experience of using integrated development environment like Eclipse, Net beans, JDeveloper, My Eclipse.
  • Good Experience in writing complex SQL queries with databases like DB2, Oracle 10g, MySQL, SQL Server and MS SQL Server 2005/2008.
  • Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs.
  • Extensive experience in writing Pig scripts to transform raw data from several data sources into forming baseline

TECHNICAL SKILLS:

Business Tools: Tableau8.X, Business Objects XI R2, Informatica PowerCenter 8.x, OLAP/OLTP, Dimension Modeling, Data Modeling

Big Data: Hadoop Map Reduce 1.0/2.0, Pig, Hive, HBase, Sqoop, Oozie, Zookeeper, Avro, Kafka, Spark, Flume, Storm, Impala, Scala, Mahout, Hue, Flink, Tez

Web development: HTML, Java Script, XML, PHP, JSP, Servlets, JavaScript

Databases: DB2, MySQL, MS Access, MS SQL server, Teradata, Vertica, SSAS, Oracle, Oracle Essbase, Cassandra, MongoDB, Amazon DynamoDB, Redis

Languages: Java / J2EE, Scala, Python HTML, SQL,Spring, Hibernate, JDBC, JSON, JavaScript

Operating Systems: MacOS, Unix, Linux (Various Versions)Windows 2003/7/8/8.1/XP/Vista

Web/Application server: Apache Tomcat,WebLogic, WebSphere Tools Eclipse

Version Control: Git, SVN, Perforce, Eclipse, NetBeans, JDeveloper

PROFESSIONAL EXPERIENCE:

Confidential - Chicago, Illinois

Sr. Hadoop Developer/Administrator

Responsibilities:

  • Installed and configured Spark ecosystem components (Spark SQL, Spark Streaming, MLlib or GraphX)
  • Developed high integrity programs used in systems where predictable and highly reliable operation is essential using Spark
  • Designed Columnar families in Cassandra and Ingested data from RDBMS, performed data transformations, and then exported the transformed data to Cassandra as per the business requirement.
  • Tested the cluster Performance using Cassandra-stress tool to measure and improve the Read/Writes.
  • Diagnosed Cassandra problems by setting Log4J Debug mode for detailed tracing and analyzing Cassandra deferred reads and writes.
  • Configured internode communication between Cassandra nodes and client using SSL encryption.
  • Worked on tuning Bloom filters and configured compaction strategy based on the use case.
  • Closely associated with Cassandra DBA in implementing Cassandra data model in application environment to ensure solution is not affecting existing business as usual
  • Cloudera Hadoop installation & configuration of multiple nodes using Cloudera Manager and CDH 4.X/5. X.
  • Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
  • Prepared low-level Design document and estimated efforts for the project.
  • Build/Tune/Maintain Hive QL and Pig Scripts for user reporting.
  • Developed the PIG code for loading, filtering and storing the data.
  • Developed a real-time analytics project using Kafka and Spark in Scala for optimizing ad placements on Confidential .com website.
  • Used Kafka HDFS connector to export data from Kafka topics to HDFS files in a variety of formats and integrates with Apache Hive to make data immediately available for querying with HiveQL.
  • Developed Oozie workflow for scheduling and orchestrating the ETL process.
  • Involved in bug fixing.
  • Extracted the data from MySQL into HDFS using Sqoop.
  • Implemented HBASE for creating tabular data.
  • Developed Map Reduce Programs using MRv1 and MRv2 (YARN).
  • Involved in loading data from UNIX to HDFS.
  • Installed and configured Hive.
  • Coordinate and communicate with Onsite team and preparing technical design documents.
  • Involved in creating Hive tables, loading data, and writing Hive queries.
  • Involved in running Hadoop jobs for processing millions of records of text data for batch and online processes by using Tuned/Modified SQL.
  • Understanding the business Requirements and Technical Requirements.
  • Designed published workbooks and dashboards usingTableau Dashboard/Server 6.X/7.X
  • Developed data pipeline using Flume, Sqoop, Pig and Java map reduce to ingest customer behavioral data and financial histories into HDFS for analysis.

Environment: Hadoop (HDFS) multi-node installation, Map Reduce, Spark, Kafka Hive, Impala, flume, Storm, Zookeeper, Oozie, Java, Scala, JDK, UNIX Shell Scripting,TestNG, MySQL, Eclipse, Toad, Tableau 8.X/9.X and HP Vertica 6.X/7.X

Confidential - LosAngeles, CA

Sr. Hadoop Developer

Responsibilities:

  • Converting the existing relational database model to Hadoop ecosystem
  • Wrote JUNIT programs to test the data
  • Performed integration testing and unit testing for the data processed using various big data components.
  • Developed high integrity programs used in systems where predictable and highly reliable operation is essential using Spark.
  • Designed Columnar families in Cassandra and Ingested data from RDBMS, performed data transformations, and then exported the transformed data to Cassandra as per the business requirement.
  • Used DataStax Spark-Cassandra connector to load data into Cassandra and used CQL to analyze data from Cassandra tables for quick searching, sorting and grouping
  • Generated Datasets to load them into HADOOP Ecosystem.
  • GoldenGateKafka adapters are used to write data to Kafka clusters.
  • Configured Kafka to read and write messages from external programs and handle real time data.
  • Worked with Linux systems and RDBMS database on a regular basis to ingest data using Sqoop.
  • Worked with Spark to create structured data from the pool of unstructured data received.
  • Managed and reviewed Hadoop and HBase log files.
  • Creating Hive tables and working on them using Hive QL.
  • Involved in loading data from UNIX file system and FTP to HDFS.
  • Designed and implemented HIVE queries and functions for evaluation, filtering, loading and storing of data.
  • Responsible to manage data coming from different sources.
  • Developed Hive queries to analyze the output data,
  • Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
  • Responsible to do the cluster co-ordination services through Zookeeper.
  • Collected the logs data from web servers and integrated into HDFS using Flume.
  • Used Hive to do transformations, event joins and some pre-aggregations before storing the data onto HDFS.
  • Supported the existing MapReduce Programs those are running on the cluster,
  • Wrote the shell scripts to monitor the health check of Hadoop daemon services and respond accordingly for warning and failure conditions.
  • Involved in Hadoop cluster task like adding and removing Nodes without any effect to running jobs and data.
  • Followed agile methodology for the entire project.
  • Installed and configured Apache Hadoop, Hive and Pig environment.

Environment: Linux-ubuntu, Hadoop pseudo distributed mode 1.2.1, HDFS, Hive 0.1.2, Flume, Horton works,Spark,Flume,Hive.

Confidential - New York, NY

Hadoop Developer

Responsibilities:

  • Launching and setup of Hadoop cluster which includes configuring different components of Hadoop.
  • Hands on experience in loading data fromUNIX file system to HDFS.
  • Wrote the MapReduce Jobs to parse the web logs which are stored in HDFS
  • Developed simple to complex MapReduce jobs using Hive and Pig.
  • DevelopedMap Reduce jobs in PIG and Hive for data cleaning and pre-processing.
  • Cluster coordination services through Zookeeper.
  • Designed and implemented Hive queries and functions for evaluation, filtering, loading and storing of data.
  • Used Oozie workflow engine to run multiple Hive and pig jobs which run independently with time and data availability.
  • Designing and development of technical architecture, requirements and statistical models using R,
  • Used storm to analyze large amounts of non-unique data points with low latency and high throughput.
  • Expertise in Partitions, bucketing concepts in Hive and analyzed the data using the HiveQL
  • Worked on migrating PIG scripts and MapReduce programs to Spark Data frames API and Spark SQL to improve performance.
  • Experienced with performing CURD operations in HBase.
  • Involved in writing optimized PIG script along with involved in developing and testing PIG Latin scripts,
  • Created MapReduce programs for some refined queries on big data.
  • Working knowledge in writing PIG’s load and store functions.
  • Worked with NoSQL databases like HBase in creating HBase tables to load large sets of semi-structured.

Environment: Apache, Hadoop1.0.1, MapReduce, HDFS, CentOS, Zookeeper, Sqoop, Cassandra, Hive, PIG, Oozie, Java, Eclipse, Amazon, EC2, JSP, servlets.

Confidential - Chicago,IL

Hadoop Developer

Responsibilities:

  • Pulled the data from data warehouse using Sqoop and placed in HDFS.
  • Responsible for Installing, Configuring, testing Hadoop Ecosystem Components
  • Wrote MapReduce jobs to join data from multiple tables and convert it to CSV files.
  • Worked with Play Framework to design the frontend of the application.
  • Worked programs on Scala to support the play framework and act as code behind for the frontend application.
  • Wrote programs in Java and at times Scala to implement intermediate functionalities like events or record count from the HBase.
  • Configured multiple remote worker nodes and Master nodes from scratch to as per the software requirement specifications.
  • Also wrote some Pig scripts to do ETL transformations on the MapReduce processed data.
  • Involved in review of functional and non-functional requirements.
  • Responsible to manage data coming from different sources.
  • Wrote shell scripts to pull the necessary fields from huge files generated by MapReduce jobs,
  • Converted ORC data from hive into flat file using MapReduce jobs.
  • Creating Hive tables and working on them using Hive QL.
  • Supported the existing MapReduce Programs those are running on the cluster.
  • Followed agile methodology for the entire project.
  • Preparing technical design and detailed design documents

Environment: Linux Ubuntu, Hadoop Pseudo distributed mode 1.2.1, HDFS, Hive, Hortonworks, Flume, Hive.

Confidential

Hadoop Developer

Responsibilities:

  • Responsible for loading the customer's data and event logs from Kafka into HBase using REST API.
  • Experienced in managing and reviewing Hadoop Logfiles.
  • Converting the existing relational database model to Hadoop ecosystem
  • Responsible for Cluster maintenance, adding and removing cluster nodes, Cluster Monitoring and Troubleshooting, Manage and review data backups and log files.
  • Worked on debugging, performance tuning and Analyzing data using Hadoop components Hive & Pig.
  • Created Hive tables from JSON data using data serialization framework like AVRO.Implemented generic export framework for moving data from HDFS to RDBMS and vice-versa.
  • Worked on loading data from LINUX file system to HDFS.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Responsible for processing unstructured data using Pig and Hive.
  • Adding nodes into the clusters & decommission nodes for maintenance.
  • Extensive experience in managing and reviewing Hadoop log files.
  • Developed Pig Latin scripts for extracting data.
  • Extensively used Pig for data cleansing and HIVE queries for the analysts.
  • Created PIG script jobs in maintaining minimal query optimization.
  • Very good understanding of Partitions, bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
  • Created WebIreport with multiple data providers and synchronized the data using Merge Dimensions.
  • Developed WebI Reports as per the business requirements.
  • Developed WebI Reports (on demand, AdHoc Reports, Frequency Reports, Summary Reports, Sub Reports, Drill-Down and Cross-Tab).
  • Used Business Object Reporting functionalities such as Slice and Dice, Master/detail, User Response function and different Formulas.

Environment: Hadoop, HDFS, Pig, Hive, MapReduce, Sqoop, Oozie, Spark, Hue, LINUX, Teradata, Java APIs, Java collection, SQL Business Objects XI R2.

Confidential

Data warehouse Consultant

Responsibilities:

  • Involved in the design and development of Data Warehouse.
  • We developed this project from initial scratch to final production.
  • Extensively used SQL and PL/SQL for development of Procedures, Functions, Packages and Triggers.
  • Developed and Supported Map Reduce Programs those are running on the cluster.
  • Involved in using Pig Latin to analyze the large-scale data.
  • Involved in loading data from UNIX file system to HDFS.
  • Used Informatica Power Center 9.5/8.6.1 as ETL tool for developing the project.
  • Gathered the requirements from the client for the ETL Objects Implementation Designed jobs to FTP the data, using FTP stage, from flat file source systems onto the Informatica UNIX server.
  • Interacted with business users on regular basis to consolidate and analyze the requirements and presented them with design results.
  • Involved in data visualization and provided the files required for the team by analyzing the data in Hive and developed Pig scripts for advanced analytics on the data
  • Created many user-defined routines, functions, before/after subroutines which facilitated in implementing some of the complex logical solutions.
  • Worked on improving the performance by using various performance tuning strategies.
  • Managed the evaluation of ETL and OLAP tools and recommended the most suitable solutions depending on business needs.
  • Migrated jobs from development to test and production environments.
  • Used Shell Scripts for loading, unloading, validating and records auditing purposes.
  • Used Teradata Aster bulk load feature to bulk load flat files to Aster.
  • Shell Scripts are also used for file validating, records auditing purposes.
  • Used Aster UDFs to unload data from staging tables and client data for SCD which resided on Aster database.

Environment: Informatica 8.X/9.X, Oracle 10g, Java, SQL, PL/SQL, Unix Shell Scripting, XML, Teradata Aster, Hive, Pig, Hadoop, MapReduce, Clear Case, HP Unix, Windows XP professional.

We'd love your feedback!