We provide IT Staff Augmentation Services!

Sr. Hadoop Developer Resume

5.00/5 (Submit Your Rating)

Hudson, OhiO

PROFESSIONAL SUMMARY:

  • Around 8 years of professional experience with 4 years of experience in developing, implementing, and configuring Bigdata technologies like Hadoop ecosystem and development of various web applications using Java, J2EE.
  • Hands on experience in Big Data Eco system like HDFS, Map Reduce, Hive, Pig, HBase, Sqoop, YARN, Spark, Scala, Oozie, Kafka and Zoo - Keeper.
  • Work experience with different Hadoop distributions like Horton Works and Cloudera.
  • Excellent understanding of Hadoop distributed File system and experienced in developing efficient MapReduce jobs to process large datasets.
  • Good working knowledge in using Sqoop and Flume for data ingestion.
  • Good knowledge in using apache NiFi to automate the data movement between different Hadoop systems.
  • Implemented Talend jobs to load data from different sources and integrated with Kafka.
  • Highly skilled in integrating Kafka with Spark streaming for high speed data processing.
  • Very good at loading data into spark schema RDD’s and querying them using Spark-SQL.
  • Good at writing custom RDD’s in Scala and implemented design patterns to improve the performance.
  • Experienced in using apache Hue and Ambari to manage and monitor the Hadoop clusters.
  • Experience in analysing large amounts of data using Pig and Hive scripts.
  • Sound knowledge in using Apache Solr to search against structured and un-structured data.
  • Worked with Azkaban and Oozie workflow schedulers to recurrently run Hadoop jobs.
  • Experience in writing Queries for moving data from HDFS to HIVE and analyzing the data using HIVE QL
  • Experience in implementing Kerberos authentication protocol in Hadoop for data security.
  • Experience in creating dash boards and generating reports using QlikSense.
  • Experience in using ORC, Parquet and Avro file formats.
  • Developed Spark code and Spark-SQl/Streaming for faster testing and processing of data.
  • Worked on NoSQL databases like HBase, Cassandra and MongoDB to store the processed data.
  • Good knowledge in cloud integration with Amazon Elastic MapReduce (EMR), Amazon Cloud Compute (EC2), Amazon's Simple Storage Service (S3) and Microsoft Azure.
  • Hands on experience on UNIX environment and shell scripting.
  • Experienced in using version control system GIT, build tool Maven and integration tool Jenkins.
  • Expertise in development of Web Applications using J2EE technologies like Servlets, JSP, Web Services, Spring, Hibernate, HTML, JQuery, Ajax etc.
  • Implemented design patterns to improve quality and performance of the applications.
  • Worked on Junit to test the functionality of java methods and used Python to do automation.
  • Good experience in using Relational databases Oracle, SQL Server, and PostgreSQL.
  • Experience in developing the J2EE applications using technologies like Java, JDBC and Servlets.
  • Worked with Waterfall, agile, Scrum and Sprint software development framework for managing product development.

AREAS OF EXPERTISE:

Hadoop Eco System: HDFS, MapReduce, Pig, Hive, Sqoop, Flume, Zookeeper, Oozie, Kafka, Storm, Talend, Spark, NiFi and Avro.

Programming languages: Java, Python, Scala.

No SQL Databases: HBase, Cassandra, MongoDB.

Databases: Oracle, SQL Server, PostgreSQL.

Web Technologies: HTML, JQuery, Ajax, CSS, JavaScript, JSON, XML.

Business Intelligence Tools: QlikSense, Jasper reports.

Testing: Hadoop Testing, Hive Testing, MRUnit.

Operating Systems: Linux Red Hat/Ubuntu/CentOS, Windows 10/8.1/7/XP.

Hadoop Distributions: Cloudera Enterprise, Horton Works.

Technologies and Tools: Servlets, JSP, Spring (Boot, MVC, Batch, Security), Web Services, Hibernate, Maven, GitHub.

Application Servers: Tomcat, JBoss.

IDE’s: Eclipse, Net Beans, IntelliJ.

PROFESSIONAL EXPERIENCE:

Confidential, Hudson, Ohio

Sr. Hadoop Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop cluster environment with HortonWorks distribution.
  • Used sqoop to load the data from relational databases.
  • Involved in converting Hive/SQL queries into spark transformations using Spark RDD’s.
  • Worked with CSV, Jason, Avro and Parquet file formats.
  • Implemented usage of Amazon EMR for processing Big Data across Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service(S3).
  • Worked on Kafka to collect and load the data on Hadoop file systems.
  • Used Hive to form an abstraction on top of structured data resides in HDFS and implemented Partitions , Buckets on HIVE tables.
  • Developed and implemented real-time data pipelines with Spark Streaming.
  • Designed, developed data integration programs in a Hadoop environment with NoSQL data store HBase for data access and analysis.
  • Worked with Python , to develop analytical jobs using PySpark API of spark.
  • Using Job management scheduler apache Oozie to execute the workflow.
  • Using Ambari to monitor node’s health, status of the jobs and to run the analytics jobs in Hadoop clusters.
  • Experience with pyspark for using spark libraries by using python scripting for data analysis.
  • Worked on Qliksense to build customized interactive reports, worksheets, and dashboards.
  • Implemented Kerberos for strong authentication to provide data security.
  • Involved in performance tuning of spark jobs using Cache and by utilizing complete advantage of cluster environment.

Environment: Hadoop , Spark, Scala, Python, Kafka, Hive, Sqoop, Pyspark, Apache Hue, Talend, Oozie, HBase, QlikSense, Jenkins, HortonWorks.

Confidential, Bowie, Maryland

Big Data Developer

Responsibilities:

  • Involved in loading data from UNIX file system to HDFS using Shell Scripting.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems and vice-versa.
  • Involved in creating Hive Tables, loading with data and writing Hive queries which will invoke and run Map Reduce jobs in the backend and generating tableau reports on top of it.
  • Created a Spark Streaming application to consume real-time data from Kafka sources and applied real-time data analysis models that we can update on new data in the stream as it arrives.
  • Worked on importing, transforming large sets of structured semi-structured and unstructured data.
  • Implemented the workflows using Azkaban to automate tasks. Used Zoo-keeper to co-ordinate cluster services.
  • Developed Hive queries for data analysis which are automated in Azkaban process.
  • Used Spark-Structured-Streaming to perform necessary transformations and data model which get’s the data from Kafka in real time and Persists into HDFS.
  • Worked extensively in Impala Hue to analyze the processed data and to generate the end reports.
  • Experience in both SQL Context and Spark Session.
  • Developed custom UDF’s for pig scripts for cleaning unstructured data and used different joins and groups whenever required to optimize the pig scripts.
  • Integrated Map Reduce with HBase to import bulk data using MR programs.
  • Worked on different file formats like Sequence files, XML files and Map files using Map Reduce Programs.
  • Used Tableau for building customized reports and dashboards.

Environment: HDFS, Map Reduce, Sqoop, Pig, Hive, Impala, Cassandra, Azkaban, Tableau, Java, Git and Shell Scripting.

Confidential, Atlanta, Georgia

Hadoop Developer

Responsibilities:

  • Design and develop analytic systems to extract meaningful data from large scale structured and unstructured health data with Cloudera distribution.
  • Created Sqoop jobs to populate data present in relational databases to hive tables.
  • Developed UDF’s in java for enhancing functionalities of Pig and Hive scripts.
  • Solved performance issues in Pig and Hive scripts with deep understanding in joins, groups and aggregations and how these jobs do translate into MapReduce jobs.
  • Involved in creating Hive external tables, loading data, and writing Hive queries.
  • Developed the processed data in Cassandra for faster querying and random access.
  • Defined job flow using Azkaban scheduler to automate the Hadoop jobs and installed zookeepers for automatic node failovers.
  • Managing and reviewing Hadoop log files to find the source for job failures and debugging the scripts for code optimization.
  • Developed complex MapReduce Programs to analyse data that exists on the cluster.
  • Developed the processes to load data from server logs into HDFS using Flume and loading from UNIX file system to HDFS.
  • Build a platform to query and display the analysis results in dashboard using QlikSense .
  • Used Cloud Manager web interface to monitor the Hadoop cluster and run the jobs.
  • Developed Shell scripts to automate routine DBA tasks (i.e. data refresh, backups)
  • Involved in the performance tuning for Pig Scripts and Hive Queries.

Environment: HDFS, Map Reduce, Sqoop, Pig, Hive, Flume, Cassandra, Azkaban, QlikSense, Java, Maven, Git, Horton works, Eclipse, Ambari and Shell Scripting.

Confidential

Java Developer

Responsibilities:

  • Designed and implemented the training and reports modules of the application using Servlets, JSP and Ajax.
  • Developed custom JSP tags for the application.
  • Writing queries for fetching and manipulating data using ORM software iBatis.
  • Used Quartz schedulers to run the jobs sequentially at given time.
  • Implemented design patterns like Filter, Cache Manager and Singleton to improve the performance of the application.
  • Implemented the reports module of the application using Jasper Reports to display dynamically generated reports for business intelligence.
  • Deployed the application in client’s location on Tomcat Server.

Environment: HTML, Java Script, Ajax, Java, Servlets, JSP, iBatis, Tomcat Server, SQL Server, Jasper Reports.

Confidential

Java Developer

Responsibilities:

  • Involved in designing and development of the project using java and J2EE technologies by following MVC architecture of which JSP’s are views and Servlets as controllers.
  • Using StarUML designed network and use case diagrams to monitor the work flow.
  • Wrote server-side programs to handle requests coming from different types of devices like iOS and using RESTful Web Services.
  • Implemented design patterns like Cache Manager and Factory classes to improve the performance of the application.
  • Used hibernate ORM tool to store and retrieve the data from PostgreSQL database.
  • Involved in writing test cases for the application using Junit.
  • Followed the Agile software development process to do this project and achieved the fast development.

Environment: JSP, Spring MVC, Spring Security, Servlets, Ajax, RESTful, Hibernate, Design Patterns, StarUML, Eclipse and PostgreSQL.

We'd love your feedback!