Sr. Hadoop Developer/administrator Resume
Chicago, IllinoiS
SUMMARY:
- Over 8+ years of experience in Hadoop/Big Data technologies such as in Hadoop, Pig, Hive, HBase, Oozie, Zookeeper, Sqoop, Storm, Flink, Flume, Impala, Tez, Kafka and Spark with hands on experience in writing Map Reduce/YARN and Spark/Scala jobs.
- Expertise in working and designing of Row keys &Schema design with NOSQLdatabases like Mongo DB 3.0.1, HBase, Cassandra and DynamoDB (AWS).
- Deep expertise in Analysis, Design, Development and Testing phases of Enterprise Data Warehousing solutions.
- Expertise in Tableau BI reporting tools &Tableau Dashboards Developments &Server Administration.
- Around 1 year of experience in Business Objects Desktop Intelligence, Web Intelligence, Universe Designer, Crystal Reports and Central Management Console.
- Experience in data ingress and egress using Sqoop from HDFS to Relational Database Systems and vice - versa. Good knowledge of Log4j for error handling.
- Expert knowledge in real time data analytics using Apache Storm.
- Expertise in Java/J2EE technologies such as Core Java, spring, Hibernate, JDBC, JSON, HTML, Struts, Servlets, JSP, JBOSS and JavaScript.
- Experience in designing both time driven and data driven automated workflows using Oozie.
- Experience in migrating the data using Sqoop from Hadoop to Relational Database System and vice-versa.
- Worked in Agile methodology of software development process as a Scrum Master.
- Work experience inETLprocesses consisting of data sourcing, data transformation, mapping and loading of data from multiple source systems into Data Warehouse using Informatic Power Center.
- Strong foundation in Programming, debugging skills, developed modules which have met with client requirements & targets.
- Expertise in Hadoop administration such as managing cluster, reviewing Hadoop log files.
- Expertise in Data warehousing concepts, Dimensional Modeling and Data Modeling systems.
- Extensive Experience in developing test cases, performing Unit Testing and Integration Testing using source code management tools such as GIT, SVN, Perforce
- Experience in using PL/SQL to write Stored Procedures, Functions and Triggers.
- Have Experience of using integrated development environment like Eclipse, Net beans, JDeveloper, My Eclipse.
- Good Experience in writing complex SQL queries with databases like DB2, Oracle 10g, MySQL, SQL Server and MS SQL Server 2005/2008.
- Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs.
- Extensive experience in writing Pig scripts to transform raw data from several data sources into forming baseline
TECHNICAL SKILLS:
Business Tools: Tableau8.X, Business Objects XI R2, Informatica PowerCenter 8.x, OLAP/OLTP, Dimension Modeling, Data Modeling
Big Data: Hadoop Map Reduce 1.0/2.0, Pig, Hive, HBase, Sqoop, Oozie, Zookeeper, Avro, Kafka, Spark, Flume, Storm, Impala, Scala, Mahout, Hue, Flink, Tez
Web development: HTML, Java Script, XML, PHP, JSP, Servlets, JavaScript
Databases: DB2, MySQL, MS Access, MS SQL server, Teradata, Vertica, SSAS, Oracle, Oracle Essbase, Cassandra, MongoDB, Amazon DynamoDB, Redis
Languages: Java / J2EE, Scala, Python HTML, SQL,Spring, Hibernate, JDBC, JSON, JavaScript
Operating Systems: MacOS, Unix, Linux (Various Versions)Windows 2003/7/8/8.1/XP/Vista
Web/Application server: Apache Tomcat,WebLogic, WebSphere Tools Eclipse
Version Control: Git, SVN, Perforce, Eclipse, NetBeans, JDeveloper
PROFESSIONAL EXPERIENCE:
Confidential - Chicago, Illinois
Sr. Hadoop Developer/Administrator
Responsibilities:
- Installed and configured Spark ecosystem components (Spark SQL, Spark Streaming, MLlib or GraphX)
- Developed high integrity programs used in systems where predictable and highly reliable operation is essential using Spark
- Designed Columnar families in Cassandra and Ingested data from RDBMS, performed data transformations, and then exported the transformed data to Cassandra as per the business requirement.
- Tested the cluster Performance using Cassandra-stress tool to measure and improve the Read/Writes.
- Diagnosed Cassandra problems by setting Log4J Debug mode for detailed tracing and analyzing Cassandra deferred reads and writes.
- Configured internode communication between Cassandra nodes and client using SSL encryption.
- Worked on tuning Bloom filters and configured compaction strategy based on the use case.
- Closely associated with Cassandra DBA in implementing Cassandra data model in application environment to ensure solution is not affecting existing business as usual
- Cloudera Hadoop installation & configuration of multiple nodes using Cloudera Manager and CDH 4.X/5. X.
- Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
- Prepared low-level Design document and estimated efforts for the project.
- Build/Tune/Maintain Hive QL and Pig Scripts for user reporting.
- Developed the PIG code for loading, filtering and storing the data.
- Developed a real-time analytics project using Kafka and Spark in Scala for optimizing ad placements on Confidential .com website.
- Used Kafka HDFS connector to export data from Kafka topics to HDFS files in a variety of formats and integrates with Apache Hive to make data immediately available for querying with HiveQL.
- Developed Oozie workflow for scheduling and orchestrating the ETL process.
- Involved in bug fixing.
- Extracted the data from MySQL into HDFS using Sqoop.
- Implemented HBASE for creating tabular data.
- Developed Map Reduce Programs using MRv1 and MRv2 (YARN).
- Involved in loading data from UNIX to HDFS.
- Installed and configured Hive.
- Coordinate and communicate with Onsite team and preparing technical design documents.
- Involved in creating Hive tables, loading data, and writing Hive queries.
- Involved in running Hadoop jobs for processing millions of records of text data for batch and online processes by using Tuned/Modified SQL.
- Understanding the business Requirements and Technical Requirements.
- Designed published workbooks and dashboards usingTableau Dashboard/Server 6.X/7.X
- Developed data pipeline using Flume, Sqoop, Pig and Java map reduce to ingest customer behavioral data and financial histories into HDFS for analysis.
Environment: Hadoop (HDFS) multi-node installation, Map Reduce, Spark, Kafka Hive, Impala, flume, Storm, Zookeeper, Oozie, Java, Scala, JDK, UNIX Shell Scripting,TestNG, MySQL, Eclipse, Toad, Tableau 8.X/9.X and HP Vertica 6.X/7.X
Confidential - LosAngeles, CA
Sr. Hadoop Developer
Responsibilities:
- Converting the existing relational database model to Hadoop ecosystem
- Wrote JUNIT programs to test the data
- Performed integration testing and unit testing for the data processed using various big data components.
- Developed high integrity programs used in systems where predictable and highly reliable operation is essential using Spark.
- Designed Columnar families in Cassandra and Ingested data from RDBMS, performed data transformations, and then exported the transformed data to Cassandra as per the business requirement.
- Used DataStax Spark-Cassandra connector to load data into Cassandra and used CQL to analyze data from Cassandra tables for quick searching, sorting and grouping
- Generated Datasets to load them into HADOOP Ecosystem.
- GoldenGateKafka adapters are used to write data to Kafka clusters.
- Configured Kafka to read and write messages from external programs and handle real time data.
- Worked with Linux systems and RDBMS database on a regular basis to ingest data using Sqoop.
- Worked with Spark to create structured data from the pool of unstructured data received.
- Managed and reviewed Hadoop and HBase log files.
- Creating Hive tables and working on them using Hive QL.
- Involved in loading data from UNIX file system and FTP to HDFS.
- Designed and implemented HIVE queries and functions for evaluation, filtering, loading and storing of data.
- Responsible to manage data coming from different sources.
- Developed Hive queries to analyze the output data,
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
- Responsible to do the cluster co-ordination services through Zookeeper.
- Collected the logs data from web servers and integrated into HDFS using Flume.
- Used Hive to do transformations, event joins and some pre-aggregations before storing the data onto HDFS.
- Supported the existing MapReduce Programs those are running on the cluster,
- Wrote the shell scripts to monitor the health check of Hadoop daemon services and respond accordingly for warning and failure conditions.
- Involved in Hadoop cluster task like adding and removing Nodes without any effect to running jobs and data.
- Followed agile methodology for the entire project.
- Installed and configured Apache Hadoop, Hive and Pig environment.
Environment: Linux-ubuntu, Hadoop pseudo distributed mode 1.2.1, HDFS, Hive 0.1.2, Flume, Horton works,Spark,Flume,Hive.
Confidential - New York, NY
Hadoop Developer
Responsibilities:
- Launching and setup of Hadoop cluster which includes configuring different components of Hadoop.
- Hands on experience in loading data fromUNIX file system to HDFS.
- Wrote the MapReduce Jobs to parse the web logs which are stored in HDFS
- Developed simple to complex MapReduce jobs using Hive and Pig.
- DevelopedMap Reduce jobs in PIG and Hive for data cleaning and pre-processing.
- Cluster coordination services through Zookeeper.
- Designed and implemented Hive queries and functions for evaluation, filtering, loading and storing of data.
- Used Oozie workflow engine to run multiple Hive and pig jobs which run independently with time and data availability.
- Designing and development of technical architecture, requirements and statistical models using R,
- Used storm to analyze large amounts of non-unique data points with low latency and high throughput.
- Expertise in Partitions, bucketing concepts in Hive and analyzed the data using the HiveQL
- Worked on migrating PIG scripts and MapReduce programs to Spark Data frames API and Spark SQL to improve performance.
- Experienced with performing CURD operations in HBase.
- Involved in writing optimized PIG script along with involved in developing and testing PIG Latin scripts,
- Created MapReduce programs for some refined queries on big data.
- Working knowledge in writing PIG’s load and store functions.
- Worked with NoSQL databases like HBase in creating HBase tables to load large sets of semi-structured.
Environment: Apache, Hadoop1.0.1, MapReduce, HDFS, CentOS, Zookeeper, Sqoop, Cassandra, Hive, PIG, Oozie, Java, Eclipse, Amazon, EC2, JSP, servlets.
Confidential - Chicago,IL
Hadoop Developer
Responsibilities:
- Pulled the data from data warehouse using Sqoop and placed in HDFS.
- Responsible for Installing, Configuring, testing Hadoop Ecosystem Components
- Wrote MapReduce jobs to join data from multiple tables and convert it to CSV files.
- Worked with Play Framework to design the frontend of the application.
- Worked programs on Scala to support the play framework and act as code behind for the frontend application.
- Wrote programs in Java and at times Scala to implement intermediate functionalities like events or record count from the HBase.
- Configured multiple remote worker nodes and Master nodes from scratch to as per the software requirement specifications.
- Also wrote some Pig scripts to do ETL transformations on the MapReduce processed data.
- Involved in review of functional and non-functional requirements.
- Responsible to manage data coming from different sources.
- Wrote shell scripts to pull the necessary fields from huge files generated by MapReduce jobs,
- Converted ORC data from hive into flat file using MapReduce jobs.
- Creating Hive tables and working on them using Hive QL.
- Supported the existing MapReduce Programs those are running on the cluster.
- Followed agile methodology for the entire project.
- Preparing technical design and detailed design documents
Environment: Linux Ubuntu, Hadoop Pseudo distributed mode 1.2.1, HDFS, Hive, Hortonworks, Flume, Hive.
Confidential
Hadoop Developer
Responsibilities:
- Responsible for loading the customer's data and event logs from Kafka into HBase using REST API.
- Experienced in managing and reviewing Hadoop Logfiles.
- Converting the existing relational database model to Hadoop ecosystem
- Responsible for Cluster maintenance, adding and removing cluster nodes, Cluster Monitoring and Troubleshooting, Manage and review data backups and log files.
- Worked on debugging, performance tuning and Analyzing data using Hadoop components Hive & Pig.
- Created Hive tables from JSON data using data serialization framework like AVRO.Implemented generic export framework for moving data from HDFS to RDBMS and vice-versa.
- Worked on loading data from LINUX file system to HDFS.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Responsible for processing unstructured data using Pig and Hive.
- Adding nodes into the clusters & decommission nodes for maintenance.
- Extensive experience in managing and reviewing Hadoop log files.
- Developed Pig Latin scripts for extracting data.
- Extensively used Pig for data cleansing and HIVE queries for the analysts.
- Created PIG script jobs in maintaining minimal query optimization.
- Very good understanding of Partitions, bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
- Created WebIreport with multiple data providers and synchronized the data using Merge Dimensions.
- Developed WebI Reports as per the business requirements.
- Developed WebI Reports (on demand, AdHoc Reports, Frequency Reports, Summary Reports, Sub Reports, Drill-Down and Cross-Tab).
- Used Business Object Reporting functionalities such as Slice and Dice, Master/detail, User Response function and different Formulas.
Environment: Hadoop, HDFS, Pig, Hive, MapReduce, Sqoop, Oozie, Spark, Hue, LINUX, Teradata, Java APIs, Java collection, SQL Business Objects XI R2.
Confidential
Data warehouse Consultant
Responsibilities:
- Involved in the design and development of Data Warehouse.
- We developed this project from initial scratch to final production.
- Extensively used SQL and PL/SQL for development of Procedures, Functions, Packages and Triggers.
- Developed and Supported Map Reduce Programs those are running on the cluster.
- Involved in using Pig Latin to analyze the large-scale data.
- Involved in loading data from UNIX file system to HDFS.
- Used Informatica Power Center 9.5/8.6.1 as ETL tool for developing the project.
- Gathered the requirements from the client for the ETL Objects Implementation Designed jobs to FTP the data, using FTP stage, from flat file source systems onto the Informatica UNIX server.
- Interacted with business users on regular basis to consolidate and analyze the requirements and presented them with design results.
- Involved in data visualization and provided the files required for the team by analyzing the data in Hive and developed Pig scripts for advanced analytics on the data
- Created many user-defined routines, functions, before/after subroutines which facilitated in implementing some of the complex logical solutions.
- Worked on improving the performance by using various performance tuning strategies.
- Managed the evaluation of ETL and OLAP tools and recommended the most suitable solutions depending on business needs.
- Migrated jobs from development to test and production environments.
- Used Shell Scripts for loading, unloading, validating and records auditing purposes.
- Used Teradata Aster bulk load feature to bulk load flat files to Aster.
- Shell Scripts are also used for file validating, records auditing purposes.
- Used Aster UDFs to unload data from staging tables and client data for SCD which resided on Aster database.
Environment: Informatica 8.X/9.X, Oracle 10g, Java, SQL, PL/SQL, Unix Shell Scripting, XML, Teradata Aster, Hive, Pig, Hadoop, MapReduce, Clear Case, HP Unix, Windows XP professional.
