We provide IT Staff Augmentation Services!

Hadoop-spark Developer Resume

3.00/5 (Submit Your Rating)

Mason, OhiO

PROFESSIONAL SUMMARY:

  • Having overall 5 years of Experience as a Hadoop/Spark Developer with experience in all phases of Software Application requirement analysis, design, development and maintenance ofHadoop/Big Data application and web applications using java/J2EE technologies with specializing in Finance, Health care the Role of Hadoop/Spark developer in different developing methodologies like Agile and Waterfall.
  • Expertise in all components of Hadoop Ecosystem - Hive, Pig, HBase, Impala, Sqoop, HUE, Flume, Zookeeper, Oozie and Apache Spark.
  • In depth understanding/knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, and MapReduce concepts and experience in working with MapReduce programs using Apache Hadoop for working with Big Data to analyze large data sets efficiently.
  • Hands-on experience on YARN (MapReduce 2.0) architecture and components such as Resource Manager, Node Manager, Container and Application Master and execution of a MapReduce job.
  • Hands on Experience in designing and developing applications in Spark using Scala to compare the performance of Spark with Hive.
  • Experienced in integrating Kafka with Spark streaming for high speed data processing.
  • Developed SPARK CODE using SCALA and Spark-SQL for faster testing and processing of data.
  • Used Spark API over Cloudera Hadoop YARN to perform analytics on data.
  • Exposure in working with data frames.
  • Experience in collecting the log data from different sources (webservers and social media) using Flume, Kafka and storing in HDFS to perform the MapReduce jobs.
  • Worked on Data Serialization formats for converting Complex objects into sequence bits by using AVRO, PARQUET, CSV format
  • Strong knowledge of Pig and Hive's analytical functions, extending Hive and Pig core functionality by writing custom UDFs.
  • Expertise in developing PIG Latin Scripts and Hive Query Language for data Analytics.
  • Well-versed in and implemented Partitioning, Dynamic-Partitioning and bucketing concepts in Hive to compute data metrics.
  • Well-versed with Agile Development process tools like Jira.
  • Integrated BI tool like Tableau with Impala and analyzed the data.
  • Experience with NoSQL databases like HBase, MongoDB and Cassandra.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems/ Non-Relational Database Systems and vice-versa.
  • Used Oozie job scheduler to schedule MapReduce jobs and automate the job flows and Implemented cluster coordination services using Zookeeper.
  • Reviewed the HDFS usage and system design for future scalability and fault-tolerance.
  • Experienced in working with Amazon Web Services (AWS) using EC2 for computing and S3 as storage mechanism.
  • Knowledge in creating different visualizations using Bars, Lines and Pies, Maps, Scatter plots, Histograms, Highlight tables and application of local and global filters according to the end user requirement in Tableau.
  • Knowledge in designing and creating various analytical reports and Automated Dashboards to help users to identify critical KPIs and facilitate strategic planning in the organization.
  • Experience in working with different relational databases like MySQL and Oracle.
  • Strong experience in database design, writing complex SQL Queries and Stored Procedures
  • Expertise in various faces of Software Development including analysis, design, development and deployment of applications using Servlets, JSP, Java Beans, Struts, Spring Framework, JDBC.
  • Having Experience on Development applications like Eclipse, NetBeans etc.
  • Proficient in software documentation and technical report writing.
  • Versatile team player with good communication, analytical, presentation and inter-personal skills.

TECHNICAL EXPERTISE:

Big Data Ecosystem: HDFS and Map Reduce, Pig, Hive, Impala, YARN, HUE, Oozie, Zookeeper, ApacheSpark, Apache STORM, Apache Kafka, Sqoop, Flume.

Operating Systems: Windows, Ubuntu, RedHat Linux, Unix

Programming Languages: C, C++, Java, SCALA

Scripting Languages: Shell Scripting, Java Scripting

Databases: Oracle 11g/10g/9i, MySQL, DB2, MS-SQL Server, SQL, PL/SQL

NoSQL Databases: HBase, Cassandra, and MongoDB

Hadoop Distributions: Cloudera, Hortonworks

Build Tools: Maven, sbt

Development IDEs: NetBeans, Eclipse IDE

Web Servers: Web Logic, Web Sphere, Apache Tomcat 6

Version Control Tools: SVN, Git, GitHub

Packages: Microsoft Office, putty, MS Visual Studio

PROFESSIONAL EXPERIENCE:

Confidential, Mason, Ohio

Hadoop-Spark Developer

Responsibilities:

  • Extracted the data from RDBMS into HDFS using Sqoop.
  • Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing.
  • Used Apache Kafka for importing real time network log data into HDFS.
  • Used Flume to collect, aggregate and store the web log data from different sources like web servers, mobile and network devices and pushed into HDFS.
  • Created and worked Sqoop jobs with incremental load to populate Hive External tables in Big Data.
  • Implemented MapReduce programs on log data to transform into structured way to find user information.
  • Extensive experience in writing Pig scripts to transform raw data from several data sources into forming baseline data.
  • Developed UDF functions for Hive and wrote complex queries in Hive for data analysis.
  • Export the analyzed data to relational databases using Sqoop for visualizations and to generate reports for the BI team.
  • Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig in Big Data.
  • Involved in file movements between HDFS and AWS S3 and extensively worked with S3 bucket in AWS.
  • Extensively used ETL processes to load data from flat files into the target database by applying business logic on transformation mapping for inserting and updating records when loaded.
  • Created standard and best practices for Talend ETL components and jobs.
  • Created Talend jobs to copy the files from one server to another and utilized Talend FTP components. Created and managed Source to Target mapping documents for all Facts and Dimension tables.
  • Worked on data serialization formats for converting complex objects into sequence bits by using Avro and ORC file formats.
  • Used Hadoop with the AWS EC2 using certain instances to gather and analyzing the data log files.
  • Created data partitions on large data sets in S3. Analyzed the SQL scripts and designed the solution to implement using Scala.
  • Developed analytical component using Scala, Spark and Spark Stream.
  • Development of complex Hive scripts to transform raw data from the staging area in Big Data.
  • Designed and developed Hive tables to store staging and historical data.
  • Created Hive tables as per requirement, internal and external tables are defined with appropriate static and dynamic partitions, intended for efficiency in Big Data.
  • Processed large data sets utilizing Hadoop cluster. The data that are stored on HDFS were preprocessed/validated using Pig and then processed data was stored into Hive warehouse which enabled Business analysts to get the required data from Hive.
  • Experience in using ORC file format with Snappy compression for optimized storage of Hive tables.
  • Analyzing the source data to know the quality of data by using Talend Data Quality.
  • Solved performance issues in Hive scripts with understanding of Joins, Group and aggregation.
  • Developed Oozie workflow for scheduling and orchestrating the ETL process.
  • Involved in migrating MapReduce jobs into Spark jobs and used Spark SQL and Data Frames API to load structured and semi-structured data into Spark clusters.
  • Implemented Spark using Scala and Spark SQL for faster testing and processing of data.

Environment: Apache Hadoop, HDFS, MapReduce, Sqoop, Flume, Pig, Hive, HBase, Oozie, Scala, Spark, Spark Streaming, Kafka, Linux.

Confidential, CHICAGO, IL

Hadoop-Spark Developer

Responsibilities:

  • Responsible for Writing MapReduce jobs to perform operations like copying data on HDFS, load and transform large sets of structured, semi-structured and unstructured data in Big Data.
  • Developed a process for Scooping data from multiple sources like SQL Server, Oracle and Teradata.
  • Responsible for creation of mapping document from source fields to destination fields mapping.
  • Developed a shell script to create staging, landing tables with the same schema like the source and generate the properties which are used by Oozie jobs.
  • Developed Oozie workflow's for executing Sqoop and Hive actions.
  • Worked with NoSQL databases like HBase in creating HBase tables to load large sets of semi structured data coming from various sources in Big Data.
  • Experienced in using debug mode of Talend to debug a job to fix errors.
  • Developed Talend jobs to populate the claims data to data warehouse-star schema. Created workflows and tested mappings and workflows in development and production environment.
  • Worked with Parallel connectors for Parallel Processing to improve job performance while working with bulk data sources in Talend.
  • Handled insert and update Strategy using tmap. Used ETL methodologies and best practices to create Talend ETL jobs.
  • Performance optimizations on Spark/Scala. Diagnose and resolve performance issues.
  • Responsible for developing Python wrapper scripts which will extract specific date range using Sqoop by passing custom properties required for the workflow.
  • Developed scripts to run Oozie workflows, capture the logs of all jobs that run on cluster and create a metadata table which specifies the execution times of each job.
  • Developed Hive scripts for performing transformation logic and loading the data from staging zone to final landing zone in Big Data.
  • Designed and Implemented the ETL process using Talend Enterprise Big Data Edition to load the data from Source to Target Database.
  • Worked on Parquet File format to get a better storage and performance for publish tables.
  • Involved in loading transactional data into HDFS using Flume for Fraud Analytics.
  • Developed Python utility to validate HDFS tables with source tables.
  • Designed and developed UDF'S to extend the functionality in both PIG and HIVE in Big Data.
  • Import and Export of data using Sqoop between MySQL to HDFS on regular basis.
  • Responsible for developing multiple Kafka Producers and Consumers from scratch as per the software requirement specifications.
  • Involved in using CA7 tool to setup dependencies at each level (Table Data, File and Time).
  • Automated all the jobs for pulling data from FTP server to load data into Hive tables using Oozie workflows.
  • Involved in developing Spark code using Scala and Spark-SQL for faster testing and processing of data and exploring of optimizing it using Spark Context, Spark-SQL, Pair RDD's, Spark YARN.
  • Migrating the needed data from Oracle, MySQL in to HDFS using Sqoop and importing various formats of flat files in to HDFS.

Environment: MapReduce, HDFS, Hive, Java, Pig, Linux, XML. HBase, Zookeeper, Kafka, Sqoop, Flume, Oozie.

Confidential

Hadoop Developer

Responsibilities:

  • Designing technical architecture and developed various Big Data workflows using custom map Reduce, Pig, Hive, Cassandra and Sqoop.
  • Deployed on premise cluster and tuned the cluster for optimal performance for job execution needs and processes large data sets.
  • Built re-usable Hive UDF libraries for business requirements which enabled various business analysts to use these UDF’s in Hive querying.
  • Used FLUME to dump the application server logs into HDFS.
  • The logs that are stored on HDFS are analyzed and the cleaned data is imported into Hive warehouse which enabled end business analysts to write Hive queries.
  • Configured various big data workflows to run on the top of Hadoop using Oozie and these workflows comprise of heterogeneous jobs like Pig, Hive, Sqoop and MapReduce.
  • Experience in working with NoSQL database HBase in getting real time data analytics.
  • Used Maven extensively for building jar files of MapReduce programs and deployed to Cluster.
  • Assigned the tasks of resolving defects found in testing the new application and existing applications.
  • Analyzing the requirements, designing and developing solutions.
  • Managing Project team in achieving the project goals including resource allocation, resolving technical issues and mentoring the resources.
  • Used Linux (Ubuntu) machine for designing, developing and deploying of Java modules.

Environment: MapReduce, Pig, Hive, Sqoop, Kafka, FLUME, HBase, JDK 1.6, Maven, Linux.

Confidential

JAVA Developer

Responsibilities:

  • Core Javacoding and development using Multithreading and Design Patterns.
  • Designed dynamic user interfaces using AJAX and jQuery to retrieve data without reloading the page and send asynchronous request.
  • Developed presentation layer using JSP, HTML, CSS and client validation using JavaScript.
  • Developed Servlets and JSP based on MVC pattern using Struts framework.
  • Created Action Classes, Form Beans, and Model Objects for the application using Model View Controller (MVC) approach.
  • Involved in the integration of spring for implementing Dependency Injection.
  • Created connections to database using Hibernate Session Factory, using Hibernate APIs to retrieve and store data to the database with Hibernate transaction control.
  • Working on a banking application to create JSON webservice by calling a SOAP service.
  • Hands on experience with Web Services including REST and SOAP.
  • Optimized SQL queries used in batch processing.
  • Extensively written unit test cases using JUnit framework.
  • Used JIRA tool for tracking stories progress and follow agile methodology.
  • Developed the application using NetBeans as the IDE and used its features for editing, debugging, compiling, formatting, build automation and SVN.
  • Used Gradle tool for building and deploying the Web applications in Jboss.
  • For Bulk Order Processing, Implemented Functionality to Read Input Data from MS-Excel Files using Javaand JXL API.

Environment: Core Java, Multithreading, Jdk, JDBC, Servlets, JSP, Struts, Hibernate, Spring, Web Services, JSP, jQuery, JSON, AJAX, Html, CSS, JavaScript, log4j, SQL Server, Junit, Gradle, Jboss Server, GIT, NetBeans, DOJO, UNIX, Waterfall.

Confidential

JAVA Developer

Responsibilities:

  • Production support role for the application which was developed in JSP, Servlets to solve the sensitive issues.
  • Designed and deployed the required Stateful Session Beans to achieve various functionalities.
  • Used JavaBeans to handle the form data and the data from the back-end database.
  • Created and deployed Servlets, Session Beans on to WebLogic Server.
  • Extensively used XML to save and retrieve the user ps.
  • Used DOM parser for manipulating XML document.
  • Created dynamic web pages using JSP, static pages using HTML and developed business logic using EJB and XML.
  • Numerous XSL style sheets created for highly complex, graphically Presentations.

Environment: J2EE (EJB, JNDI, JDBC), Servlets, Log4J, Struts framework, JMS, JSP1.2, Applets, JDBC, HTML, DHTML, Ajax, jQuery JavaScript, CSS, ANT, UML, XML, WebSphere 5.0, CVS, Sybase.

We'd love your feedback!