We provide IT Staff Augmentation Services!

Hadoop/spark Developer Resume

4.00/5 (Submit Your Rating)

Atlanta, GA

SUMMARY

  • Around 8 years of overall IT experience in all phases of Software Development Life Cycle (SDLC) with skills in data analysis, design, development, testing and deployment of software systems.
  • 3+ years of relevant experience in design and development of Big Data Analytics using Apache Hadoop ecosystem components Map Reduce, HDFS, HBase, Hive, Impala, Sqoop, Pig, Oozie, Zookeeper, Flume and Spark.
  • Good knowledge in Hadoop architecture and various components such as HDFS, Job tracker, Task tracker, Resource Manager, Name Node, Data Node and Map Reduce concepts.
  • Wrote Hive Queries, Pig Scripts for data analysis to meet the requirements.
  • Enhanced the functionalities of Hive and Pig by writing UDF’s.
  • Experience in importing and exporting data using Sqoop in between HDFS and RDBMS.
  • Worked on Map - Reduce programs in java for ETL operations with multiple file formats including XML, JSON, CSV and other compressed file formats.
  • Utilized Flume to analyze log files and write into HDFS.
  • Worked on NoSQL databases using Hbase and Cassandra to analyze data.
  • Experienced Hadoop job schedulers like Oozie, Control-M workflow engine.
  • Good understanding in concepts like java technologies such as Hibernate, JDBC, Servlet, JSP, JavaScript and Struts.
  • Good experience in working with data ingestion, storage, processing and analyzing the big data.
  • Used GitHub version control tool to push and pull functions to get the updated code from repository.
  • Worked in various AWS cloud services like EC2, S3 and RDS.
  • Worked on replacing the existing MR jobs to Spark data transform, actions for faster in memory operations.
  • Developed Spark SQL jobs on hive tables to load data into HDFS and run queries on top of that.
  • Developing Spark best practices like Partitions, Caching check pointing for performance and UDF’s.
  • Created calculated columns in Spark data streams.
  • Worked with YARN, MESOS and Spark default schedulers.

TECHNICAL SKILLS

Hadoop Eco-Systems: HDFS, MapReduce, Pig, Hive, Sqoop, Oozie, Flume, Kafka, Impala, ZooKeeper, CDH, Spark, NiFi.

Spark Components: Apache Spark, Data Frames, Spark SQL

Programming Languages: SQL, Java, Pig Latin, Hive QL, Cobol, Scala, Python

Databases: IBM DB2, VSAM, MySQL, Hbase, Cassandra 2.1, Oracle DB, AWS

Web Technologies: HTML, CSS, XML, Java script, Ajax

Operating Systems: Windows, UNIX, Linux Distributions

IDE's & other tools: Eclipse, Net Beans, AutoSys, Log4j, Toad, SQL Developer

PROFESSIONAL EXPERIENCE

Confidential, Atlanta, GA

Hadoop/Spark Developer

Responsibilities:

  • Worked on Hadoop YARN clusters for data processing and analysis using Spark core, Spark SQL, Sqoop, Pig, Hive, Impala and NoSQL databases.
  • Ingested data to Hadoop Data lake using Sqoop from different RDBMS databases.
  • Worked with different data ingestion teams to embed best practices.
  • Exported the information to RDBMS using Sqoop export from HDFS to accommodate data for BI team to analyze and generate reports.
  • Used Map Reduce programs to perform data validation checks like data type validation, NULL checks for the ingested data.
  • Developed Hive and Pig custom UDF’s to maintain unique date format across the HDFS.
  • Implemented partitioning, dynamic partitioning and buckets in Hive to overcome performance issues while querying historical tables in reference to batch FDB date.
  • Streamlined Hadoop jobs and workflows using Oozie workflow.
  • Performance tuning the Hive and Pig queries with Tez engine, partitioning.
  • Worked with SCD Type 1 and SCD Type 2 data. Used Hbase to handle SCD Type 2 data.
  • Load the data into Spark RDD and performed in-memory data computation to get faster output response.
  • Used Snappy compression technique to optimize the HDFS storage.
  • Worked on different data formats like Text file, Avro and Parquet.
  • Worked with Shell-Scripting to clean the base data.
  • Used an AWS services like S3, EC2 for smaller datasets.
  • Performed data validity checks on the imported data using Spark in Scala.
  • Implemented AWS S3 to migrate the data from internal servers.
  • Worked on POC for Apache Kafka and Spark Streaming.
  • Worked with spark eco system using Spark SQL queries on data formats like Text file, CSV file and XML files.
  • Replaced jobs with Spark SQL which were running earlier in Hive QL for performance.
  • Worked with Kafka message queue for Spark streaming.
  • Used Kerberos to enable security for databases also created secure passwords using jceks for flume.

Environment: Horton Works, Apache Hadoop, HDFS, AWS, Map Reduce, Eclipse, Hive, Pig, Sqoop, Spark, Flume, Hue, Oozie, Hbase, Cassandra, Control-M, Apache Kafka, Apache NiFi

Confidential, Phoenix, AZ

Hadoop Developer

Responsibilities:

  • Installed and configured Hadoop Environment.
  • Developed multiple Map-Reduce jobs in java for data cleaning and preprocessing.
  • Installed and configured Pig and also written Pig Latin scripts.
  • Used pig and map reduce to analyze XML files and log files.
  • Imported data using Sqoop to load data from IBM DB2 to HDFS on regular basis.
  • Written Hive queries for data analysis to meet the business requirements.
  • Creating Hive tables and working on them using Hive QL.
  • Importing and exporting data into HDFS and Hive using Sqoop from IBM DB2, Netezza Databases.
  • Used Oozie workflow to co-ordinate pig and hive scripts.
  • Used Impala for querying HDFS data to achieve better performance.
  • Designed and implemented Map-Reduce based large-scale parallel relation-learning system.
  • Setup and benchmarked Hadoop/Hbase clusters for internal use.
  • Developed UDF’s to pre-process the data and compute various metrics for reporting in both pig and hive.
  • Developed Map Reduce program to convert mainframe fixed length data to delimited data.
  • Data ingestion from various IBM DB2 tables to HDFS using Sqoop.
  • Automated Python scripts to pull and synchronize the code in GitHub environment.

Environment: Hadoop, CDH, Map Reduce, HDFS, Pig, Hive, Oozie, Java, UNIX, Flume, Impala, Hbase, Oracle, Map R AutoSys, Mainframes, JCL, IBM DB2, NDM.

Confidential

Java/J2EE Developer

Responsibilities:

  • Developed the web pages using JSP, JavaScript, HTML5, CSS and struts for the user Interface.
  • Multi-Threading tasks are been performed using Java Threading API.
  • Using Struts validation technique worked on UI validations.
  • Worked in Agile Methodology and played a role of SCRUM master.
  • Worked on limited access capabilities for various modules to make sure un-authorized users are taken care of.
  • Performed UAT and SIT on application modules.
  • Wrote SQL scripts to create and maintain the database, roles, users, tables, views, procedures and triggers.
  • Used JDBC connectors to make connection with external sources like payment gateway.
  • Performing Root Cause Analysis (RCA) on defects found and fixing them in order to make the application potentially strong.
  • Used hibernate framework to communicate with Oracle 11g database.
  • Performed unit testing using JUnit and functional testing.
  • Implemented application servers like Apache Tomcat, Web Sphere, Glassfish and Web Logic in project based on the requirement.
  • Worked with servlets to connect with database server and to fetch the data.
  • Involved in implementation of SOAP and REST based web services.
  • Achieved dependency injection using injection of spring services, spring controllers and DAOs.
  • Involved at the time of project implementation and make sure that business checkouts are successful.

Environment: HTML5, CSS, struts, JDBC, Windows Unix, Servlets, UML, Xml, SQL, JUnit, Apache Tomcat, Web Sphere, Glassfish, REST and SOAP web services, Spring servlets, Spring controllers.

Confidential

Java Developer

Responsibilities:

  • Implemented JQuery to make changes to different websites to update the page layout and content.
  • Used Adobe Flex and MVC framework to develop Web pages.
  • Maintained more number of stored procedures by converting many SQL statements to SP in order to have less number of DB accesses.
  • Developed business logic using Enterprise Java Beans.
  • Designed UML diagrams using UML and Rational Rose.
  • Used JUnit framework for UAT of application and Log4j to capture the run time exception logs.
  • Used JDBC to call stored procedures and JDBC to connect the SQL database.
  • Enhanced web pages for Single Sign On using JSP and implemented Hibernate for mapping and persist the data.
  • Worked and supported the creation of database schema objects (tables, stored
  • Procedures and triggers) using Oracle SQL/PLSQL.

Environment: JDBC, HTML, CSS, Ajax, SQL, Log4j, JSP, Eclipse IDE, Stored Procedures, IBM, Rational Rose, Java Script, Jquery.

Confidential

Programmer Analyst/SQL Developer

Responsibilities:

  • Responsible for the designing the advance SQL queries, procedure, cursor, triggers, scripts.
  • Maintained the documents and create reports for the business analysts and end-users.
  • Responsible for the management of the database performance, backup, replication, capacity and security.
  • Created & modified database objects like tables, views, procedures, functions, triggers, packages, indexes, synonyms, materialized views using Oracle tools like TOAD and SQL Navigator.
  • Developed SQL and PL/SQL scripts to transfer tables across the schemas and databases.
  • Updated procedures, functions, triggers and packages based on the change request from users.
  • Worked with testing teams; perform UAT testing with business users
  • Worked with release team for the staging & production move.
  • Going through the requirements. Developing flat file reports for transaction, card, merchant activity and merchant money settlement reporting Using PL/SQL and Informatica.
  • Loading Incremental Data to fact and dimension tables.
  • Registration of new forms, creation of concurrent executable and concurrent programs for reports, opening up descriptive flex fields was done.
  • Analyzed, developed, tuned, tested, debugged and documented processes using technologies SQL, PL/SQL, Informatica, UNIX, and Control-M.

Environment: Oracle 10g, TOAD, SQL Developer, Forms&Reports6i, PL/SQL, UNIX, Informatica 8.x/9.x, UNIX Linux, Windows, SVN, Control-M, Environment: Oracle, SQL Developer, TOAD, Windows 2000/XP, ASP.Net, Visual Studio.

We'd love your feedback!