We provide IT Staff Augmentation Services!

Hadoop Developer Resume

3.00/5 (Submit Your Rating)

Weston, FloridA

PROFESSIONAL SUMMARY:

  • 5 years of total IT experience in software development life cycle, around 3 years of experience in Hadoop and Big Data Eco System.
  • Good experience in Hadoop ecosystem like Hadoop MapReduce, HDFS, NIFI, Oozie, Hive, Sqoop, Pig, Zookeeper, Flume, Spark streaming, Spark SQL,HBase and Cassandra.
  • Expertise in Hadoop 2.0 and YARN architecture.
  • Experience in usingCloudera’s CDH, Horton works HDP.
  • Expertise in writing Hadoop Jobs for analyzing, scripting of data using MapReduce, Hive and Pig.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems (RDBMS) and vice - versa.
  • Developed and implemented Apache NIFI across various environments, written QA scripts in Python for tracking files.
  • Expertise in writing custom UDF’s and UDAF’s for extending Hive and Pig core functionalities.
  • Experience in implementation of various Hadoop file-formats and compression techniques like Sequence, Parquet, ORC, Avro, Z-Zip and Text file.
  • Experienced in using NoSQL data bases like HBase, Cassandra, MongoDB.
  • Experience in working with different Databases like Oracle, MySQL, MS SQL.
  • Experience in writing UNIX, SHELL and BASH scripts.
  • Good experience in implementing advanced procedures like text analytics and processing the in-memory computing capabilities with Apache Impala, Scala.
  • Implemented Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data.
  • Converted existing MapReduce jobs into Spark transformations and actions using Spark RDD’s, Data frames and Spark SQL APIs.
  • Developed Spark jobs and Hive Jobs to summarize and transform data.
  • Experience in Writing Producers/Consumers and creating messaging centric applications using Apache Kafka.
  • Hands on experience in Amazon Web Services (AWS) provisioning tools likeEC2, Simple Storage Service (S3), Elastic Map Reduce.
  • Extensive Experience in Java development skills using J2SE, J2EE technologies like Servlets,Spring Hibernate, JSP, JDBC.
  • Experienced in Java components like Frame work collection, Exception handling, Multithreading and I/O system.
  • Experience in SOA using Soap and Restful.
  • Experience in working with Agile development methodology.
  • Proficiency in developing secure enterprise Java applications using technologies such as X-Servlets, Maven, Hibernate, XML, HTML, CSS Version Control Systems.
  • Ability to learn and adapt quickly to new tools and environment with strong communication and analytical skills.

TECHNICAL SKILLS:

Big Data Eco Systems: Hadoop (HDFS & Map Reduce), PIG, HIVE, HBASE, ZooKeeper, Sqoop, Flume, Kafka, Apache Spark, Impala, Oozie.

Databases: Oracle, SQL server, My SQL.

No SQL Databases: HBase, Cassandra, Mongo DB.

Hadoop Distributions: Cloudera, Horton works.

Cloud: AWS.

Languages: Java, Java SE, Java J2EE, Scala, Python, C.

Web Technologies: JavaScript, J-Query, Boot Strap, AJAX,XML,CSS, HTML, AngularJS.

Web Services: REST, SOAP, JAX-WS, JAX-RPC, JAX-RS, WSDL, Axis2, Apache HTTP, CVS, SVN.

IDE: Eclipse, Net beans, IntelliJ.

Operating Systems: MacOS, Linux, Windows.

PROFESSIONAL EXPERIENCE:

Hadoop Developer

Confidential, Weston, Florida

Responsibilities:

  • Worked with Hadoop Ecosystem components like HBase, Sqoop, Zookeeper, Oozie, Hive and Pig with MapRHadoop distribution.
  • Wrote Pig Scripts for sorting, joining, filtering and grouping the data.
  • Developed programs in Spark based on the application for faster data processing than standard MapReduce programs.
  • Developed spark programs using Scala, involved in creating Spark SQL Queries and Developed Oozie workflow for spark jobs.
  • Developed the Oozie workflows with Sqoop actions to migrate the data from relational databases like Oracle, Teradata to HDFS.
  • Used Hadoop FS actions to move the data from upstream location to local data locations.
  • Written extensive Hive queries to do transformations on the data to be used by downstream models.
  • Developed map reduce programs as a part of predictive analytical model development.
  • Developed Hive queries to do analysis of the data and to generate the end reports to be used by business users.
  • Worked on scalable distributed computing systems, software architecture, data structures and algorithms using Hadoop, Apache Spark and ingested streaming data into Hadoop using Spark Framework and Scala.
  • Extensively used GIT as a code repository and Version One for managing day agile project development process and to keep track of the issues and blockers.
  • Used Pyspark for model integration layer.
  • Implemented Spark using Scala, Java and utilizing Data frames and Spark SQL API for faster processing of data.
  • Developed Spark code and Spark-SQL/Streaming for faster testing and processing of data.
  • Written new spark jobs in Scala to analyze the data of the customers and sales history.
  • Used AWS SDK for connection to Amazon S3 buckets as it is used as the object storage service to store and retrieve the media files related to the application.
  • Developed API for using AWS Lambda to manage the servers and run the code in AWS.
  • Using Jenkins AWS Code Deploy plugin to deploy into AWS.
  • Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data.
  • Developed a data pipeline using Kafka, HBase, Mesos Spark and Hive to ingest, transform and analyzing customer behavioral data.

Environment: MapR, LINUX, Hadoop, HBase, Hive, AWS, Impala, Oracle, Spark, Scala, Python, Pig, Sqoop, Teradata, Zookeeper, Oozie, MongoDB, Map Reduce, GitHub.

HADOOP DEVELOPER

Confidential, Plano, TX

Responsibilities:

  • Developed data pipeline using Spark, Hive, Pig, python, Impala and HBase to ingest customer behavioral data and financial histories into Hadoop cluster for analysis.
  • Responsible for implementing a generic framework to handle different data collection methodologies from the client primary data sources, validate transform using spark and load into S3
  • Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregation on the fly to build the common learner data model and persists the data in HDFS.
  • Explored the usage of Spark for improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark SQL and Spark Yarn.
  • Developed Spark Code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
  • Involved in converting Hive/SQL queries into Spark Transformations using Spark RDDs and Scala.
  • Worked on the Spark SQL and Spark Streaming modules of Spark and used Scala and Python to write code for all Spark use cases.
  • Explored the Spark to improve the performance and optimization of the existing algorithms in Hadoop using Spark-Context, Spark-SQL, Data Frame and Pair RDD's.
  • Migrated historical data to S3 and developed a reliable mechanism for processing the incremental updates.
  • Used Oozie workflow engine to manage independent Hadoop jobs and to automate several types of Hadoop such as java MapReduce, Hive and Sqoop as well as system specific jobs
  • Used to monitor and debug Hadoop jobs/applications running in production.
  • Worked on providing user support and application support on HadoopEcosystem.

Environment: Spark, Hive, Pig, Spark SQL, Spark Streaming, HBase, Sqoop, Kafka, AWS EC2, S3, Cloudera, Scala IDE (Eclipse), Scala, Linux Shell Scripting, HDFS.

Hadoop Developer

Confidential, Wilmington, DE

Responsibilities:

  • Importing data using Sqoop into HDFS vice versa.
  • Worked on loading and transformation of large sets of structured, semi structured and unstructured data into Hadoop System.
  • Responsible to manage data coming from different data sources.
  • Developed simple and complex MapReduce programs in Java for Data Analysis.
  • Load data from various data sources into HDFS using Flume.
  • Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
  • Developed Java MapReduce programs for the analysis of sample log file stored in cluster.
  • Used Hive and created Hive tables and involved in data loading and writing Hive UDFs.
  • Responsible for spooling data from DB2 sources to HDFS using Sqoop.
  • Created HIVE tables and provided analytical queries for business user analysis
  • Extensive knowledge on PIG scripts using bags and tuples.
  • Created tables in HIVE by partitioning and bucketing for granularity and optimization of HIVEQL.
  • Involved in identifying job dependencies to design workflow for Oozie and resource management for YARN.
  • Capturing data from existing databases that provide SQL interfaces using Sqoop.
  • Involved in loading data from UNIX file system to HDFS.
  • Installed and configured Pig, Hive and written Pig and Hive UDFs.
  • Involved in creating Hive tables, loading with data and writing Hive queries which will run internally in map way.

Environment: Cloudera,HBase, Java, Hive, Pig,Sqoop,Oozie, Oracle, SVN, Kafka, GitHub, JIRA, Talend.

JAVA DEVELOPER

Confidential

Responsibilities:

  • Extensively involved in different stages of Agile Development Cycle including Detailed Analysis, Design, Develop and Test.
  • Implemented the Back-End Business Logic using Core Java technologies including Collections, Generics, Exception Handling, Java Reflection and Java I/O.
  • Wrote and specified Spring Annotation Configuration to define Beans and View Resolutions to configure Spring beans, dependencies and the services needed by beans.
  • Used Spring IC to implement dynamic dependency injection and Spring AOP to implement crosscutting concerns such as transaction management.
  • Wrote Mapping Configuration files to implement ORM Mappings in the Persistence Layer.
  • Using Hibernate DAO support extended Dao Implementation.
  • Hibernate Configuration files were written to connect Oracle database and fetch data.
  • The Hibernate Query Cache was implemented using EhCache to improve the performance.
  • Implemented web services with RESTful standards with the support of JAX-RS APIs.
  • Confirmation of registration and monthly statements are sent to users by integrating and implementing JavaMail API.
  • Manipulated database data with SQL queries, including setting up stored procedures and triggers.
  • Implemented front-end developments such as webpages design, data binding, Single-Page Applications using HTML/CSS, JavaScript, jQuery and AJAX.
  • Used jQuery libraries to simplify the frontend programming works. Performed users' input validation using JavaScript and jQuery.
  • Utilized Node.js and MongoDB to generate tendency charts of the application for Payment History.
  • Performed JUnit test cases to test the service layers of the application.
  • Used JIRA to track the projects and GIT to ensure version control.

Environment: Java, Spring, JavaMail, JavaScript, HTML, CSS, AJAX, Jquery,Junit, JIRA, Oracle DB,MongoDB,GIT.

We'd love your feedback!