We provide IT Staff Augmentation Services!

Hadoop/ Big Data Developer Resume

5.00/5 (Submit Your Rating)

Omaha, NE

SUMMARY:

  • Over 4+ years of experience in software development with experience in phases of Hadoop and HDFS development.
  • 3 years of exclusive experience in Big Data technologies and Hadoop ecosystem.
  • Strong experience working with HDFS, Spark, MapReduce, Hive, Pig, YARN, HDFS, Oozie, Sqoop, Flume, Kafka and NoSQL Databases like HBase and Cassandra.
  • Strong understanding of Distributed systems design, HDFS architecture, internal working details of MapReduce and Spark processing frameworks.
  • Worked with Talend BigData tool for the development of Mapreduce and spark jobs.
  • Solid experience developing Spark Applications for performing high scalable data transformations using RDD, DataFrames and Spark - SQL.
  • Experienced in working with structured data using HiveQL, JOIN operations, writing custom UDFs and optimizing Hive queries.
  • Experienced working with various Hadoop Distributions (Cloudera, Hortonworks, MapR, Amazon EMR) to fully implement and leverage new Hadoop features.
  • Experienced in working with Apache Flume and Kafka to collect, aggregate and move huge chunks of real-time data from various sources such as web server, telnet sources etc.
  • Developed custom Kafka Producers and Consumers in Java using Kafka High level consumer API.
  • Extensive experience in importing/ exporting data to / from RDBMS and HDFS using Apache Sqoop.
  • Strong experience working with big data services on the cloud especially with AWS
  • Strong experience in working with UNIX/ LINUX environments, writing shell scripts.
  • Strong understanding of real time streaming technologies Spark and Kafka.
  • Very good knowledge of Partitions, bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
  • Extensive experiences in working with semi/unstructured data by implementing complex MapReduce. Programs.
  • Used different file formats like RCFile, ORC and Parquet formats.
  • Extensive experience in building dashboards with Tableau.
  • Strong experience of developing, implementing and maintaining application systems under UNIX Operating System using SQL, PL/SQL, SQL Server, UNIX Shell Script.
  • Demonstrated technical expertise, organization and client service skills in various projects undertaken.

TECHNICAL SKILLS:

Programming Languages & Scripts: Python, Java

Bigdata/ Hadoop Technologies: Hadoop, HDFS, YARN, MapReduce, Hive, Pig, Impala, Sqoop, Flume, Spark, Kafka, Storm, Drill, Zookeeper, and Oozie

NO SQL Databases: Cassandra, HBase, MongoDB.

Hadoop Distributions: Cloudera Enterprise, Horton Works.

Operating Systems: Linux Red Hat/Ubuntu/CentOS, Windows 10/8.1/7/XP.

Databases: Oracle, SQL Server, PostgreSQL.

Software Tools: Talend, Tableau.

EXPERIENCE:

Confidential, Omaha, NE

Hadoop/ Big Data Developer

Responsibilities:

  • Processed data into HDFS by developing solutions and analyzed the data using Map Reduce PIG, and Hive to produce summary results from Hadoop to downstream systems.
  • Responsible for installing Talend on multiple environments, creating projects, setting up user roles, setting up job servers, configure TAC options, adding Talend jobs, job failures, on-call support and scheduling etc.
  • Build servers using AWS: Importing volumes, launching EC2, creating security groups, auto-scaling, load balancers, Route53, SES and SNS in the defined virtual private connection.
  • Written Map Reduce code to process and parsing the data from various sources and storing parsed data into HBase and Hive using HBase-Hive Integration.
  • Streamed AWS log group into Lambda function to create service now incident.
  • Developed Spark code by using Scala and Spark-SQL for faster processing and testing and performed complex HiveQL queries on Hive tables.
  • Used Hive to perform data validation on the data ingested using scoop and flume and the cleansed data set is pushed into HBase.
  • Worked on S3 buckets on AWS to store networking log files.
  • Enabled load balancer for impala to distribute data load on all Impala daemons across the cluster.
  • Designing ETL Data Pipeline flow to ingest the data from RDBMS source to Hadoop using a shell script, Sqoop, package,and MySQL.
  • Implementing Hadoop with the AWSEC2 system using a few instances in gathering and analyzing data log files.
  • Involved in Spark and Spark Streaming creating RDD's, applying operations -Transformation and Actions.
  • Involved in Cluster maintenance, Cluster Monitoring, and Troubleshooting, Manage and review data backups and log files.
  • Involved in the development of Talend Jobs and preparation of design documents, technical specification documents.
  • Analyzing the Hadoop cluster and different BigData analytic tools including Pig, Hive, HBase and Sqoop.
  • Improved the Performance by tuning of HIVE and map reduce.

Environment: Hadoop, AWS, Spark, Scala, Python, Kafka, Hive, Sqoop, Pyspark, Apache Hue, Talend, Oozie, HBase, QlikSense, Jenkins, HortonWorks.

Confidential, Wilton, CT

Hadoop Developer

Responsibilities:

  • Installed and configured Hadoop clusters for application development and Hadoop tools like Hive, Pig, Sqoop, HBase, Flume and Zookeeper.
  • Worked on developing ETL processes to load data into HDFS using Sqoop and export the results back to RDBMS.
  • Effectively used Sqoop to transfer data between databases and HDFS.
  • Designed workflow by scheduling Hive processes for Log file data, which is streamed into HDFS using Flume.
  • Involved in creating Hive tables, and loading and analyzing data using hive queries .
  • Developed Pig Latin scripts to extract the data from the mainframes output files to load into HDFS.
  • Developed Map-Reduce programs to cleanse the data in HDFS obtained from heterogeneous data sources to make it suitable for ingestion into Hive schema for analysis.
  • Used Talend ETL tool to extract data from oracle to HDFS.
  • Used Talend Admin Console Job conductor to schedule ETL Jobs on daily, weekly, monthly and yearly basis.
  • Followed the organization defined Naming conventions for naming the Flat file structure, Talend Jobs and daily batches for executing the Talend Jobs.
  • Built code for real-time data ingestion using Java, MapR-Streams (Kafka) and STORM.
  • Written Hive queries for analysation and reporting purposes of different streams in the company.
  • Processed the source data to structured data and store in NoSQL database Cassandra.
  • Created alter, insert and delete queries involving lists, sets and maps in Cassandra.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce, Hive and Sqoop as well as system specific jobs.
  • Use Avro serialization technique to serialize data. Applied transformations and standardizations and loaded into HBase for further data processing.
  • Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
  • Moved data from Third party system to Hadoop File System (HDFS) vice versa using shell commands hosted on an AWS cluster.
  • Working with data delivery teams to setup new Hadoop users. This job includes setting up Linux users, setting up Kerberos principals and testing MFS, Hive.
  • Exported the analyzed data to the relational databases using Sqoop for virtualization and to generate reports for the BI team.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
  • Developed Shell Scripts for automation of routine tasks and implementing file sharing on the network by configuring Hadoop on the system to share essential resources.
  • Worked in AWS environment for development and deployment of Custom Hadoop Applications.
  • Developed a data pipeline using Kafka and Storm to store data into HDFS.
  • Knowledge on handling Hive queries using Spark SQL that integrates with Spark environment.
  • Consuming the data from HBASE and producing to the Apache SOLR.
  • Loading the data from HBase to Solr for performance cache.
  • Documented all the requirements, code and implementation methodologies for reviewing and analysation purposes.

Environment: Java 7, Talend, Eclipse, Hadoop, Pig, Hive, MapReduce, Hbase, Sqoop, Flume, AWS, Spark, Solr, Oozie, Kafka, Storm Oracle11g.

We'd love your feedback!