Sr. Big Data/hadoop Developer Resume
Philadelphia, PA
SUMMARY:
- Overall 8+ years of Experience in IT industry in Designing, Developing and Maintaining Enterprise Applications using Big Data Technologies like Hadoop and Spark Ecosystems, Java/J2EE.
- Experience in installation, configuration and deployment of Big Data solutions.
- Experience in installing, customizing and testing the Big Data and Hadoop Eco Systems such as Hive, Pig, Sqoop, Spark, and Oozie.
- Extensive experience with advanced J2EE Frameworks such as spring, Struts, JSF and Hibernate.
- Working on Bootstrap, AngularJS and NodeJS, knockout, ember, Java Persistence Architecture (JPA).
- Excellent understanding of Hadoop Architecture such as HDFS, Name Node, Data Node, Job Tracker, Task Tracker and Map Reduce Concepts.
- Hands on experience in various big data application phases like data ingestion, data analytics and data visualization
- In - depth understanding of Spark Architecture including Spark Core, Spark SQL, Spark Streaming.
- Experience in Developing Applications using Core Java Programming.
- Good experience in Core Java, JEE technologies as JDBC, Servlets, and JSP.
- Good experience in application development primarily using Hadoop, Java and worked on data analysis.
- Experience in installation, configuration, supporting and monitoring Hadoop clusters using Apache.
- Hands on experience in application development using Java, RDBMS.
- Experience with NoSQL databases including HBase, Cassandra.
- Experienced the integration of various data sources like RDBMS, Spreadsheets, and Text files.
- Working experience in creating complex data ingestion pipelines, data transformations, data management and data governance in a centralized enterprise data hub.
- Proficient in Core Java, Enterprise technologies such as EJB, Hibernate, Java Web Service, SOAP, REST Services, Java Thread, Java Socket, Java Servlet, JSP, JDBC
- Good exposure to Service Oriented Architectures (SOA) built on Web services (WSDL) using SOAP protocol.
- Written multiple MapReduce programs in Python for data extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV and other compressed file formats.
- Experience in working on the Hadoop Ecosystem, also have little experience on installing and configuring of the Hortonworks distribution and Cloudera distribution (CDH3 and CDH4).
- Experience in NoSQL database HBase, MongoDB and Cassandra.
- Good understanding of Hadoop architecture and hands on experience with Hadoop components such as Job Tracker, Task Tracker, Name Node, Data Node and MapReduce programming.
TECHNICAL SKILLS:
Hadoop/Big Data Technologies: Hadoop 3.0, HDFS, MapReduce, HBase 1.4, Apache Pig, Hive 2.3, Sqoop 1.4, Apache Impala 2.1, Oozie 4.3, Yarn, Apache Flume 1.8, Kafka 1.1, Zookeeper 3.5
Cloud Technologies: AWS, Azure
Programming Language: Java, Scala, Python 3.6, SQL, PL/SQL, Shell Scripting, Storm 1.0, JSP, Servlets
Frameworks: Spring 5.0.5, Hibernate 5.2, Struts 1.3, JSF, EJB, JMS
Web Technologies: HTML, CSS, JavaScript, JQuery 3.3, Bootstrap 4.1, XML, JSON, AJAX
Database Tools: TOAD, SQL PLUS, SQL Server
Operating Systems: Linux, Unix, Windows 10/8/7
IDE and Tools: Eclipse 4.7, NetBeans 8.2, IntelliJ, Maven
NoSQL Databases: HBase 1.4, Cassandra 3.2, MongoDB
Web/Application Server: Apache Tomcat 9.0.7, JBoss, Web Logic, Web Sphere
SDLC Methodologies: Agile, Waterfall, SDLC
Version Control: GIT, SVN, CVS
PROFESSIONAL EXPERIENCE:
Confidential - Philadelphia, PA
Sr. Big Data/Hadoop Developer
Responsibilities:
- As a Sr. Big Data Developer, I’ve implemented solutions for ingesting data from various sources and processing the Data-at-Rest utilizing Big Data technologies.
- Installed and configured Hadoop MapReduce, HDFS, developed multiple MapReduce jobs in Java and Nifi for data cleaning and preprocessing.
- Involved in Agile methodologies, daily scrum meetings, spring planning.
- Primarily involved in Data Migration process using Azure by integrating with GitHub repository and Jenkins.
- Imported and exported (Data Ingestion) data into HDFS from Oracle database and vice versa using Sqoop.
- Installed and configured Hadoop Cluster for major Hadoop distributions.
- Worked on importing data from various sources and performed transformations using MapReduce, Hive to load data into HDFS.
- Implemented multiple MapReduce Jobs in java for data cleansing and pre-processing.
- Imported the data from different sources like HDFS/HBase into Spark RDD and developed a data pipeline using Kafka and Storm to store data into HDFS.
- Used MapReduce to receive real time data from the Kafka and store the stream data to HDFS using Scala and NoSQL databases such as HBase and Cassandra.
- Involved in loading and transforming large sets of Structured, Semi-Structured and Unstructured data and analyzed them by running Hive queries.
- Worked on compression mechanisms to optimize MapReduce Jobs.
- Developed Big Data Solutions that enabled the business and technology teams to make data-driven decisions on the best ways to acquire customers and provide them business solutions.
- Used MapR frame work for Map Reduce scripts .
- Created scripts to automate the process of Data Ingestion.
- Performed joins, group by and other operations in MapReduce by using Java.
- Configured Sqoop jobs to import data from RDBMS into HDFS using Oozie workflows.
- Worked on analyzing Hadoop cluster and different big data analytic tools including MapReduce, Hive.
- Written multiple MapReduce programs for data extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV & other compressed file formats.
- Experience in working with Flume to load the log data from multiple sources directly into HDFS.
- Worked in the BI team in Big Data Hadoop cluster implementation and data integration in developing large-scale system software.
- Experienced in using Java Rest API to perform CURD operations on HBase data.
- Extensively used XSL as XML parsing mechanism for showing Dynamic Web Pages in HTML format.
- Implemented SOAP protocol to get the requests from the outside System.
- Used CVS as a source control for code changes.
- Implemented Security in Web Applications using Azure and deployed Web Applications to Azure.
- Implemented a distributed messaging queue to integrate with Cassandra using Apache Kafka and Zookeeper.
- Used ANT scripts to build the project and JUnit to develop unit test cases.
Environment: Hadoop 3.0, HDFS, MapReduce, Java, Agile, Jenkins 2.16, Azure, Hive 2.3, Pig 0.17, HBase1.0, HDFS, NoSQL, Cassandra 3.2, ANT, Apache Kafka 2.2, Zookeeper 3.5, JSON
Confidential - Menlo Park, CA
Sr. Hadoop/Spark Developer
Responsibilities:
- Worked on performance and optimization of existing algorithms in Hadoop using Spark context, Spark-SQL and Spark YARN using Scala.
- Involved in complete Big Data flow of the application data ingestion from upstream to HDFS, processing the data in HDFS and analyzing the data.
- Involved in various phases of development analysed and developed the system going through Agile Scrum methodology.
- Responsible for Importing and exporting data into HDFS and Hive using Sqoop from Oracle and MySQL.
- Developed Restful web services using Java, Spring Boot, NoSQL databases like MongoDB.
- Develop and maintain several batch jobs to run automatically depending on business requirements.
- Used AWS Data Pipeline to create complex data processing workloads that are fault tolerant, repeatable, and highly available.
- Responsible for building scalable distributed data solutions using Hadoop.
- Developed spark programs that integrate with AWS S3 to retrieve process and store data.
- Used AWS EMR for running spark jobs which provided dynamic resizing ability.
- Involved in loading data from edge node to HDFS using shell scripting.
- Implemented Partitioning, Dynamic Partitioning, Buckets in Hive.
- Involved in managing and reviewing Hadoop log files.
- Implemented test scripts to support test driven development and continuous integration.
- Used DynamoDB for cloud based storage.
- Worked on MongoDB by using CRUD (Create, Read, Update and Delete), Indexing, Replication and Sharding features.
- Exported the analyzed data to the relational databases like MySQL using Sqoop for visualization and to generate reports for the BI team.
- Created MapReduce programs using Java API that filter un-necessary records and find out unique records based on different criteria.
- Experienced in implementing POC's to migrate iterative map reduce programs into Spark transformations, actions using Scala.
- Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
- Worked in Test Driven Deployment (TDD) environment and used Confluence for documentation.
- Performed unit testing using MRunit.
- Installed Oozie workflow engine to run multiple Hive and pig jobs.
- Used Jira for project tracking, Bug tracking and Project Management.
Environment: Hadoop 3.0, HDFS, Agile, Spark 2.1, Oracle 12c, AWS, Hive 2.3, Sqoop, Java, Oozie 5.1, Flume, AWS, MongoDB
Confidential - Lowell, AR
Spark Developer
Responsibilities:
- Wrote Programs in Spark - Scala for Data quality check.
- Worked on Big Data infrastructure for batch processing and real time processing. Built scalable distributed data solutions using Hadoop.
- Using Python scripts, implemented algorithms.
- Developed Spark code and Spark-SQL/Streaming for faster testing and processing of data.
- Imported the data from different sources like HDFS into Spark RDD, developed a data pipeline using Kafka to store data into HDFS.
- Performed real time analysis on the incoming data.
- Worked extensively with Sqoop for importing and exporting the data from data Lake HDFS to Relational Database systems.
- Developed python scripts to collect data from source systems and store it on HDFS.
- Involved in converting Hive or SQL queries into Spark transformations using Python and Scala.
- Built Kafka Rest API to collect events from front end.
- Built real time pipeline for streaming data using Kafka and Spark Streaming.
- Worked on integrating Apache Kafka with Spark Streaming process to consume data from external sources and run custom functions
- Exploring with the Spark and improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, and Pair RDD's.
- Used Data Frame API for converting the distributed collection of data organized into named columns.
- Developed Spark jobs and Hive Jobs to summarize and transform data.
- Worked with spark core, Spark Streaming and spark SQL modules of Spark
- Developed multiple POCs using Spark and deployed on the Yarn cluster, compared the performance of Spark, with Hive.
- Developed Kafka producer and consumers, Cassandra clients and Spark along with components on HDFS, Hive.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, and Python.
- Developed the code for Importing and exporting data into HDFS and Hive using Sqoop.
- Used Spark-SQL to perform event enrichment and to prepare various levels of user behavioral summaries.
- Designed dynamic and browser compatible pages using HTML5, DHTML, CSS3 and JavaScript.
- Used AngularJS and Bootstrap to consume service and populated the page with the products and pricing returned.
- Used RESTFUL in conjunction with Ajax calls using JAX-RS and Jersey.
- Designed and developed the Application using spring and Hibernate framework.
- Developed Intranet Web Application using J2EE architecture, using JSP to design the user interfaces and Hibernate for database connectivity.
- Extensively Worked with Eclipse as the IDE to develop, test and deploy the complete application.
Environment: Spark 2.0, Scala 2.1, Python 3.7, Hadoop 3.0, HDFS, Kafka 2.2, Sqoop, python, Hive, HTML5, CSS3, JavaScript, AngularJS, Bootstrap 4.3, Ajax, Hibernate 5.4, J2EE, Eclipse 4.1, Cassandra 3.0
Confidential - Houston, TX
Java/J2ee Developer
Responsibilities:
- Experience using Agile methodology and Test-Driven Development (TDD) to build the web-based application.
- Using Spring MVC architecture and Spring Framework Implementation including Enterprise JavaBeans.
- Implemented efficient code using DAO, Singleton Pattern, Inversion of Control (IoC) and Dependency Injection (DC).
- Implemented Data Access Object (DAO) pattern to isolate the application layer from relational database.
- Used Inversion of control for event handling by the interface of the user.
- Used EJB stateless session bean for retrieving data to maintain employees profile.
- Created interactive user interface using JavaScript and JQuery libraries, AJAX validation.
- Involved RESTful web Services using java API JAX-RS for implementing annotation support.
- Used Hibernate to create persistence layer and Java Persistence API for Object Relation Mapping solution to confirm data to MySQL database.
- Developed Graphical User Interfaces by using JSF, JSP, HTML, CSS, and JavaScript.
- Integrated Hibernate with spring using Hibernate Template and uses provided methods to implement CRUD operations.
- Installation and configuration of Development Environment using Eclipse with JBoss Application server.
- Used Test-Driven Development (TTD) for improving design and the quality of the code.
- Used DTO and Service Locator, J2EE design Pattern
- Created Data Transfer Object (DTO) to minimize the number of methods calls and encapsulate the serialization mechanism for data transfer.
- Used JDBC for database connection and executing SQL queries.
- Built cross-browser compatibility using HTML, CSS, JSP, JavaScript and JQuery.
- Developed web-based application using JSP custom tag and designed Java Servlets and objects using J2EE standards.
- Implemented test cases using JUnit to perform test cases.
- Responsible for maintaining documentation and project structure as per J2EE Standards.
- Used Log4J for logging errors and debugging.
- Used Git as a version control system.
Environment: J2EE, Spring2.0, MVC, JavaScript, java, JQuery3.4, AJAX, Hibernate 5.4, MySQL, JSP 2.3, HTML5, CSS3, Eclipse, JBoss, JUnit 5.4
Confidential
Java Developer
Responsibilities:
- Developed Servlets, JSP pages, Beans, JavaScript.
- Involved in developing module for transformation of files across the remote systems using JSP and servlets.
- Designed and developed the front end using HTML, CSS and JavaScript with Ajax and tag libraries.
- Handled the client side and server side validations using Struts validation framework.
- Used Spring Core Annotations for Dependency Injection.
- Used Apache Ant for the build process.
- Involved in the development of the Java bean classes, JSPs, Servlets, and JDBC to access Oracle.
- Developed helper java classes needed for the application.
- Developed the building components of application such as JSPs, Servlets, using WebSphere Studio Application Developer.
- Generated Use case diagrams, Class diagrams, and Sequence diagrams using Rational Rose.
- Developed application using Struts Framework that leverages classical Model View Controller (MVC) architecture.
- Integrated Struts with spring by delegating Struts action management to Spring Framework which is UNIX operating system environment.
- Used Spring Framework for Dependency injection.
- Used complex queries like SQL statements and procedures to fetch the data from the database.
- Used JavaScript for client side validations and validation frame work for server side validations.
- Worked on front end JavaScript frameworks like AngularJS.
- Deployed the application on to WebSphere application server.
- Developed Maven based project structure having data layer, ORM, and Web module.
- Established a JSON contract to make a communication between the JSP and java classes.
- Used Restful Web Services to read the XML content from suppliers.
Environment: JavaScript, Ajax, Apache Ant, Struts 1.3, Java, MVC, UNIX, JavaScript, HTML 4.0, CSS2, Spring 2.0
