We provide IT Staff Augmentation Services!

Hadoop Developer Resume

5.00/5 (Submit Your Rating)

Northville, MI

SUMMARY:

  • 8 years of experience in computer engineering and IT industry andthis includes up to three years as Hadoop and Big Data Developer/Engineer.
  • Expertise in configuring Hadoop ecosystem (1.x and 2.x) components such as HDFS, MapReduce, Yarn, Pig, Hive, Impala, Hbase, Oozie, Zookeeper, Ambari, Sqoop, Kafka and Flume.
  • Good understanding of Hadoop architecture, HDFS daemons such asJob Tracker, Task Tracker, Name Node, Data Node, Secondary Name Node and YARN daemons like Resource Manager, Node Manager, Application Master and Containers.
  • Knowledge of working and modifying with Hadoop configuration files for cluster setup.
  • Hands on experience in writing complex Map Reduce jobs(Mapper, Reducer, Partioners, Combiners) in Java.
  • Experience in writing data transformations, data cleansing and processing usingPIG operation s and custom UDFs using Java and Maven dependencies.
  • ImplementingHive scriptsand creatingcustom UDFs using Java and Maven dependencies.
  • Hands on work experience with Oozie Workflow Engine for running jobs in Hadoop ecosystem.
  • Experience in importing and exporting data from different databases like MySQL, MongoDB, Cassandra, Oracle and Teradata into HDFS and vice - versa using Sqoop.
  • Working knowledge of creating real time data streaming solutions using Apache Spark/ Spark Streaming, Kafka and Flume.
  • Experience of handling different file formats like JSON, AVRO, Parquet, CSV and SequenceFile.
  • In-depth knowledge of machine learning concepts like clustering and classifiers and implementing them using Spark MLLib and NumPy.
  • Knowledge of Business Intelligence tools like Tableau and familiar with data warehousing concepts.
  • Proficient in writing Maven builds scripts to automate the application build and deployment.
  • Hands-on experience with VPN, Putty and WinScpand version control tool like Git.
  • Knowledge of Agile and Scrum methodology for project management.
  • Excellent in creating reports, technical presentations and displaying results and findings using Microsoft PowerPoint and Excel.
  • A great team player with ability to effectively communicate with all levels of the organization such as technical, management and customers.

TECHNICAL SKILLS:

Big Data Technologies:: Apache Hadoop, MapReduce, HDFS, Hive, Pig, Impala, HBase, Sqoop, Flume, Zookeeper, Ambari,Oozie, Kafka, YARN, Spark, MongoDB and Cassandra.

Hadoop Distribution:: Cloudera, Hortonworks.

Databases: Oracle, MySQL, Teradata, Microsoft SQL Server, MS Access,DB2 and NoSQL.

Programming Languages: C, C++, Java, Scala, SQL, Python, MPI, OpenMP, CUDA, Visual Basic and Unix Shell Scripts.

Tools: Eclipse IDE, IntelliJ IDEA, Maven and ANT, Tableau, MATLAB, Cygwin, Microsoft Visual Studio.

Operating Systems & others: Linux(CentOS, Ubuntu), Unix, Windows, Putty, WinSCP, VMWare, Oracle VirtualBox, AWS and Microsoft Office Suite.

PROFESSIONAL EXPERIENCE:

Confidential, Northville, MI

Hadoop Developer

  • Exporting rawfinancialdata (stock tickers) in different formats (Avro, XML, CSV) from client in batches (yearly and quarterly) into HDFS environment to create Data Lakes using Sqoop.
  • Performing data cleaning and pre-processing of the highly granular data using PIGscripts in grunt shell.
  • Implementing and deploying PIG UDFs using Java and Maven dependencies for filtering, joining/ aggregating and structuring the data for efficient querying.
  • Creating Hive tables, loading data using Incremental imports and writing queries to parse the data.
  • Collaborating with the BI team to understand the project requirements and creating customHive UDFs using Java and Maven dependencies to compute various metrics for reporting and analysis.
  • Using Spark for fast processing of data to ease iterative and interactive querying by creating RDDs by determining effective partitioning.
  • Utilizing Spark MLLib and NumPy for developing machine learning algorithms like clustering, adaptive prediction and classifiers to analyze the data.
  • Deploying Oozie Workflow Engine and writing shell scripts to manage and automate interdependent Sqoop, Pig and Hive and Spark jobs .
  • Installed and configured Apache Hadoop clustersusingAmbari for application development and Hadoop tools like Hive, Pig, Hbase, HDFS and YARN in Linux (Cent OS).
  • Designing and given technical presentations and presenting results on Hadoop architecture, Hadoop ecosystem and Big Data applications and projects.

Environment: Hadoop, Python, HDFS, Spark, MapReduce, Pig, Hive, Sqoop, Oozie, Scala, Java, SQL Scripting, Linux Shell, Zookeeper, MySQL.

Confidential, South Plainfield, NJ

Hadoop/Big Data Engineer

  • Experience in importing and exporting tera bytes of data using Sqoop from HDFS to RDBMS (MySQL) and real time data streams from servers to HDFS using Apache Flume for IoT applications.
  • Created applications requiring shared memory for processing large datasets using distributed computing principles, parallel programming using MPI and OpenMP.
  • Using Java and Maven dependencies to create Map Reduce jobs and custom Hive UDFs for querying and data extraction.
  • Experience in analyzing large datasets and finding patterns and insights within structured data using k-means clusteringalgorithm developed in Python and NumPy .
  • Deployed Hadoop ecosystem components (Apache Hadoop, Hive, Pig, Sqoop, Flume, Spark and ZooKeeper) and multimode clusters in Linux system (CentOS) using virtual machine (Oracle VM VirtualBox).
  • Used SQL querying on MySQL database for storing and updating table, creating views and extracting required data for processing and analysis.
  • Responsibilities involve performing requirement analysis, planning, application development and testing.

Environment: Hive, HDFS, Spark, Spark MLLib, MapReduce, Pig, Hive, Sqoop, Java, Linux Shell, Zookeeper, MySQL, SQL/PL, IntelliJ IDEA.

Confidential, Durham, NC

Hadoop Developer

  • Exporting raw unstructured clinical data from healthcare provider in batches into HDFS environment tocreate Data Lakes using Sqoop.
  • Performed data cleaning and pre-processing of the unstructured data using PIG scripts in grunt shell .
  • Wrote PIG UDFs using Java and Maven dependencies for filtering, joining/ aggregatingand structuring the data for efficient querying.
  • Implemented Hive internal and external tables and utilized partitioning, bucketing and joining concepts to prepare data for fast querying.
  • Created custom Hive UDFs using Java and Maven dependencies for querying and extracting data for reporting and analysis.
  • Created Spark RDDs and used Scala (functional programming) for performing data cleaning, transformation and querying.
  • Monitored and managed Hadoop clusters using Apache Ambari.
  • Created presentations using Microsoft PowerPoint and Excel for reporting results and collaborated with client’s BI team to provide data for analysis.

Environment: MPI, OpenMP, Oracle VM VirtualBox, Linux, Windows, Java, Maven, Eclipse IDE, Hadoop, Map Reduce, Hive.

Confidential

Big Data Developer/Analyst

  • Used Sqoop to export structured data batches received from clients into HDFS for storage and processing.
  • Handled the cleaning, sorting and aggregating data from large datasets for analysis and processing .
  • Supported code/design analysis, strategy development and project planning.
  • Developed multiple MapReduce jobs in Java in IntelliJ IDEA using Maven dependencies for data cleaning and pre-processing.
  • Created and utilized containers and partitioners for Map Reduce applications to improve performance and reduce communication and computation time.
  • Implemented Spark scripts by using Scala shell commands for data transformation and processing as per the project requirement.
  • Implemented text classification and recognition using classifiers and machine learning algorithms fromSpark MLlib and NumPy in Python.

Environment: Sqoop, Map Reduce, Java, Maven, IntelliJ IDE, Spark, Scala, Spark MLLib, NumPy, Python.

We'd love your feedback!