Hadoop/ Big Data Developer Resume
Omaha, NE
SUMMARY:
- Over 4+ years of experience in software development with experience in phases of Hadoop and HDFS development.
- 3 years of exclusive experience in Big Data technologies and Hadoop ecosystem.
- Strong experience working with HDFS, Spark, MapReduce, Hive, Pig, YARN, HDFS, Oozie, Sqoop, Flume, Kafka and NoSQL Databases like HBase and Cassandra.
- Strong understanding of Distributed systems design, HDFS architecture, internal working details of MapReduce and Spark processing frameworks.
- Worked with Talend BigData tool for the development of Mapreduce and spark jobs.
- Solid experience developing Spark Applications for performing high scalable data transformations using RDD, DataFrames and Spark - SQL.
- Experienced in working with structured data using HiveQL, JOIN operations, writing custom UDFs and optimizing Hive queries.
- Experienced working with various Hadoop Distributions (Cloudera, Hortonworks, MapR, Amazon EMR) to fully implement and leverage new Hadoop features.
- Experienced in working with Apache Flume and Kafka to collect, aggregate and move huge chunks of real-time data from various sources such as web server, telnet sources etc.
- Developed custom Kafka Producers and Consumers in Java using Kafka High level consumer API.
- Extensive experience in importing/ exporting data to / from RDBMS and HDFS using Apache Sqoop.
- Strong experience working with big data services on the cloud especially with AWS
- Strong experience in working with UNIX/ LINUX environments, writing shell scripts.
- Strong understanding of real time streaming technologies Spark and Kafka.
- Very good knowledge of Partitions, bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
- Extensive experiences in working with semi/unstructured data by implementing complex MapReduce. Programs.
- Used different file formats like RCFile, ORC and Parquet formats.
- Extensive experience in building dashboards with Tableau.
- Strong experience of developing, implementing and maintaining application systems under UNIX Operating System using SQL, PL/SQL, SQL Server, UNIX Shell Script.
- Demonstrated technical expertise, organization and client service skills in various projects undertaken.
TECHNICAL SKILLS:
Programming Languages & Scripts: Python, Java
Bigdata/ Hadoop Technologies: Hadoop, HDFS, YARN, MapReduce, Hive, Pig, Impala, Sqoop, Flume, Spark, Kafka, Storm, Drill, Zookeeper, and Oozie
NO SQL Databases: Cassandra, HBase, MongoDB.
Hadoop Distributions: Cloudera Enterprise, Horton Works.
Operating Systems: Linux Red Hat/Ubuntu/CentOS, Windows 10/8.1/7/XP.
Databases: Oracle, SQL Server, PostgreSQL.
Software Tools: Talend, Tableau.
EXPERIENCE:
Confidential, Omaha, NE
Hadoop/ Big Data Developer
Responsibilities:
- Processed data into HDFS by developing solutions and analyzed the data using Map Reduce PIG, and Hive to produce summary results from Hadoop to downstream systems.
- Responsible for installing Talend on multiple environments, creating projects, setting up user roles, setting up job servers, configure TAC options, adding Talend jobs, job failures, on-call support and scheduling etc.
- Build servers using AWS: Importing volumes, launching EC2, creating security groups, auto-scaling, load balancers, Route53, SES and SNS in the defined virtual private connection.
- Written Map Reduce code to process and parsing the data from various sources and storing parsed data into HBase and Hive using HBase-Hive Integration.
- Streamed AWS log group into Lambda function to create service now incident.
- Developed Spark code by using Scala and Spark-SQL for faster processing and testing and performed complex HiveQL queries on Hive tables.
- Used Hive to perform data validation on the data ingested using scoop and flume and the cleansed data set is pushed into HBase.
- Worked on S3 buckets on AWS to store networking log files.
- Enabled load balancer for impala to distribute data load on all Impala daemons across the cluster.
- Designing ETL Data Pipeline flow to ingest the data from RDBMS source to Hadoop using a shell script, Sqoop, package,and MySQL.
- Implementing Hadoop with the AWSEC2 system using a few instances in gathering and analyzing data log files.
- Involved in Spark and Spark Streaming creating RDD's, applying operations -Transformation and Actions.
- Involved in Cluster maintenance, Cluster Monitoring, and Troubleshooting, Manage and review data backups and log files.
- Involved in the development of Talend Jobs and preparation of design documents, technical specification documents.
- Analyzing the Hadoop cluster and different BigData analytic tools including Pig, Hive, HBase and Sqoop.
- Improved the Performance by tuning of HIVE and map reduce.
Environment: Hadoop, AWS, Spark, Scala, Python, Kafka, Hive, Sqoop, Pyspark, Apache Hue, Talend, Oozie, HBase, QlikSense, Jenkins, HortonWorks.
Confidential, Wilton, CT
Hadoop Developer
Responsibilities:
- Installed and configured Hadoop clusters for application development and Hadoop tools like Hive, Pig, Sqoop, HBase, Flume and Zookeeper.
- Worked on developing ETL processes to load data into HDFS using Sqoop and export the results back to RDBMS.
- Effectively used Sqoop to transfer data between databases and HDFS.
- Designed workflow by scheduling Hive processes for Log file data, which is streamed into HDFS using Flume.
- Involved in creating Hive tables, and loading and analyzing data using hive queries .
- Developed Pig Latin scripts to extract the data from the mainframes output files to load into HDFS.
- Developed Map-Reduce programs to cleanse the data in HDFS obtained from heterogeneous data sources to make it suitable for ingestion into Hive schema for analysis.
- Used Talend ETL tool to extract data from oracle to HDFS.
- Used Talend Admin Console Job conductor to schedule ETL Jobs on daily, weekly, monthly and yearly basis.
- Followed the organization defined Naming conventions for naming the Flat file structure, Talend Jobs and daily batches for executing the Talend Jobs.
- Built code for real-time data ingestion using Java, MapR-Streams (Kafka) and STORM.
- Written Hive queries for analysation and reporting purposes of different streams in the company.
- Processed the source data to structured data and store in NoSQL database Cassandra.
- Created alter, insert and delete queries involving lists, sets and maps in Cassandra.
- Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce, Hive and Sqoop as well as system specific jobs.
- Use Avro serialization technique to serialize data. Applied transformations and standardizations and loaded into HBase for further data processing.
- Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
- Moved data from Third party system to Hadoop File System (HDFS) vice versa using shell commands hosted on an AWS cluster.
- Working with data delivery teams to setup new Hadoop users. This job includes setting up Linux users, setting up Kerberos principals and testing MFS, Hive.
- Exported the analyzed data to the relational databases using Sqoop for virtualization and to generate reports for the BI team.
- Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
- Developed Shell Scripts for automation of routine tasks and implementing file sharing on the network by configuring Hadoop on the system to share essential resources.
- Worked in AWS environment for development and deployment of Custom Hadoop Applications.
- Developed a data pipeline using Kafka and Storm to store data into HDFS.
- Knowledge on handling Hive queries using Spark SQL that integrates with Spark environment.
- Consuming the data from HBASE and producing to the Apache SOLR.
- Loading the data from HBase to Solr for performance cache.
- Documented all the requirements, code and implementation methodologies for reviewing and analysation purposes.
Environment: Java 7, Talend, Eclipse, Hadoop, Pig, Hive, MapReduce, Hbase, Sqoop, Flume, AWS, Spark, Solr, Oozie, Kafka, Storm Oracle11g.
