Hadoop/spark Developer Resume
Atlanta, GA
SUMMARY
- Around 8 years of overall IT experience in all phases of Software Development Life Cycle (SDLC) with skills in data analysis, design, development, testing and deployment of software systems.
- 3+ years of relevant experience in design and development of Big Data Analytics using Apache Hadoop ecosystem components Map Reduce, HDFS, HBase, Hive, Impala, Sqoop, Pig, Oozie, Zookeeper, Flume and Spark.
- Good knowledge in Hadoop architecture and various components such as HDFS, Job tracker, Task tracker, Resource Manager, Name Node, Data Node and Map Reduce concepts.
- Wrote Hive Queries, Pig Scripts for data analysis to meet the requirements.
- Enhanced the functionalities of Hive and Pig by writing UDF’s.
- Experience in importing and exporting data using Sqoop in between HDFS and RDBMS.
- Worked on Map - Reduce programs in java for ETL operations with multiple file formats including XML, JSON, CSV and other compressed file formats.
- Utilized Flume to analyze log files and write into HDFS.
- Worked on NoSQL databases using Hbase and Cassandra to analyze data.
- Experienced Hadoop job schedulers like Oozie, Control-M workflow engine.
- Good understanding in concepts like java technologies such as Hibernate, JDBC, Servlet, JSP, JavaScript and Struts.
- Good experience in working with data ingestion, storage, processing and analyzing the big data.
- Used GitHub version control tool to push and pull functions to get the updated code from repository.
- Worked in various AWS cloud services like EC2, S3 and RDS.
- Worked on replacing the existing MR jobs to Spark data transform, actions for faster in memory operations.
- Developed Spark SQL jobs on hive tables to load data into HDFS and run queries on top of that.
- Developing Spark best practices like Partitions, Caching check pointing for performance and UDF’s.
- Created calculated columns in Spark data streams.
- Worked with YARN, MESOS and Spark default schedulers.
TECHNICAL SKILLS
Hadoop Eco-Systems: HDFS, MapReduce, Pig, Hive, Sqoop, Oozie, Flume, Kafka, Impala, ZooKeeper, CDH, Spark, NiFi.
Spark Components: Apache Spark, Data Frames, Spark SQL
Programming Languages: SQL, Java, Pig Latin, Hive QL, Cobol, Scala, Python
Databases: IBM DB2, VSAM, MySQL, Hbase, Cassandra 2.1, Oracle DB, AWS
Web Technologies: HTML, CSS, XML, Java script, Ajax
Operating Systems: Windows, UNIX, Linux Distributions
IDE's & other tools: Eclipse, Net Beans, AutoSys, Log4j, Toad, SQL Developer
PROFESSIONAL EXPERIENCE
Confidential, Atlanta, GA
Hadoop/Spark Developer
Responsibilities:
- Worked on Hadoop YARN clusters for data processing and analysis using Spark core, Spark SQL, Sqoop, Pig, Hive, Impala and NoSQL databases.
- Ingested data to Hadoop Data lake using Sqoop from different RDBMS databases.
- Worked with different data ingestion teams to embed best practices.
- Exported the information to RDBMS using Sqoop export from HDFS to accommodate data for BI team to analyze and generate reports.
- Used Map Reduce programs to perform data validation checks like data type validation, NULL checks for the ingested data.
- Developed Hive and Pig custom UDF’s to maintain unique date format across the HDFS.
- Implemented partitioning, dynamic partitioning and buckets in Hive to overcome performance issues while querying historical tables in reference to batch FDB date.
- Streamlined Hadoop jobs and workflows using Oozie workflow.
- Performance tuning the Hive and Pig queries with Tez engine, partitioning.
- Worked with SCD Type 1 and SCD Type 2 data. Used Hbase to handle SCD Type 2 data.
- Load the data into Spark RDD and performed in-memory data computation to get faster output response.
- Used Snappy compression technique to optimize the HDFS storage.
- Worked on different data formats like Text file, Avro and Parquet.
- Worked with Shell-Scripting to clean the base data.
- Used an AWS services like S3, EC2 for smaller datasets.
- Performed data validity checks on the imported data using Spark in Scala.
- Implemented AWS S3 to migrate the data from internal servers.
- Worked on POC for Apache Kafka and Spark Streaming.
- Worked with spark eco system using Spark SQL queries on data formats like Text file, CSV file and XML files.
- Replaced jobs with Spark SQL which were running earlier in Hive QL for performance.
- Worked with Kafka message queue for Spark streaming.
- Used Kerberos to enable security for databases also created secure passwords using jceks for flume.
Environment: Horton Works, Apache Hadoop, HDFS, AWS, Map Reduce, Eclipse, Hive, Pig, Sqoop, Spark, Flume, Hue, Oozie, Hbase, Cassandra, Control-M, Apache Kafka, Apache NiFi
Confidential, Phoenix, AZ
Hadoop Developer
Responsibilities:
- Installed and configured Hadoop Environment.
- Developed multiple Map-Reduce jobs in java for data cleaning and preprocessing.
- Installed and configured Pig and also written Pig Latin scripts.
- Used pig and map reduce to analyze XML files and log files.
- Imported data using Sqoop to load data from IBM DB2 to HDFS on regular basis.
- Written Hive queries for data analysis to meet the business requirements.
- Creating Hive tables and working on them using Hive QL.
- Importing and exporting data into HDFS and Hive using Sqoop from IBM DB2, Netezza Databases.
- Used Oozie workflow to co-ordinate pig and hive scripts.
- Used Impala for querying HDFS data to achieve better performance.
- Designed and implemented Map-Reduce based large-scale parallel relation-learning system.
- Setup and benchmarked Hadoop/Hbase clusters for internal use.
- Developed UDF’s to pre-process the data and compute various metrics for reporting in both pig and hive.
- Developed Map Reduce program to convert mainframe fixed length data to delimited data.
- Data ingestion from various IBM DB2 tables to HDFS using Sqoop.
- Automated Python scripts to pull and synchronize the code in GitHub environment.
Environment: Hadoop, CDH, Map Reduce, HDFS, Pig, Hive, Oozie, Java, UNIX, Flume, Impala, Hbase, Oracle, Map R AutoSys, Mainframes, JCL, IBM DB2, NDM.
Confidential
Java/J2EE Developer
Responsibilities:
- Developed the web pages using JSP, JavaScript, HTML5, CSS and struts for the user Interface.
- Multi-Threading tasks are been performed using Java Threading API.
- Using Struts validation technique worked on UI validations.
- Worked in Agile Methodology and played a role of SCRUM master.
- Worked on limited access capabilities for various modules to make sure un-authorized users are taken care of.
- Performed UAT and SIT on application modules.
- Wrote SQL scripts to create and maintain the database, roles, users, tables, views, procedures and triggers.
- Used JDBC connectors to make connection with external sources like payment gateway.
- Performing Root Cause Analysis (RCA) on defects found and fixing them in order to make the application potentially strong.
- Used hibernate framework to communicate with Oracle 11g database.
- Performed unit testing using JUnit and functional testing.
- Implemented application servers like Apache Tomcat, Web Sphere, Glassfish and Web Logic in project based on the requirement.
- Worked with servlets to connect with database server and to fetch the data.
- Involved in implementation of SOAP and REST based web services.
- Achieved dependency injection using injection of spring services, spring controllers and DAOs.
- Involved at the time of project implementation and make sure that business checkouts are successful.
Environment: HTML5, CSS, struts, JDBC, Windows Unix, Servlets, UML, Xml, SQL, JUnit, Apache Tomcat, Web Sphere, Glassfish, REST and SOAP web services, Spring servlets, Spring controllers.
Confidential
Java Developer
Responsibilities:
- Implemented JQuery to make changes to different websites to update the page layout and content.
- Used Adobe Flex and MVC framework to develop Web pages.
- Maintained more number of stored procedures by converting many SQL statements to SP in order to have less number of DB accesses.
- Developed business logic using Enterprise Java Beans.
- Designed UML diagrams using UML and Rational Rose.
- Used JUnit framework for UAT of application and Log4j to capture the run time exception logs.
- Used JDBC to call stored procedures and JDBC to connect the SQL database.
- Enhanced web pages for Single Sign On using JSP and implemented Hibernate for mapping and persist the data.
- Worked and supported the creation of database schema objects (tables, stored
- Procedures and triggers) using Oracle SQL/PLSQL.
Environment: JDBC, HTML, CSS, Ajax, SQL, Log4j, JSP, Eclipse IDE, Stored Procedures, IBM, Rational Rose, Java Script, Jquery.
Confidential
Programmer Analyst/SQL Developer
Responsibilities:
- Responsible for the designing the advance SQL queries, procedure, cursor, triggers, scripts.
- Maintained the documents and create reports for the business analysts and end-users.
- Responsible for the management of the database performance, backup, replication, capacity and security.
- Created & modified database objects like tables, views, procedures, functions, triggers, packages, indexes, synonyms, materialized views using Oracle tools like TOAD and SQL Navigator.
- Developed SQL and PL/SQL scripts to transfer tables across the schemas and databases.
- Updated procedures, functions, triggers and packages based on the change request from users.
- Worked with testing teams; perform UAT testing with business users
- Worked with release team for the staging & production move.
- Going through the requirements. Developing flat file reports for transaction, card, merchant activity and merchant money settlement reporting Using PL/SQL and Informatica.
- Loading Incremental Data to fact and dimension tables.
- Registration of new forms, creation of concurrent executable and concurrent programs for reports, opening up descriptive flex fields was done.
- Analyzed, developed, tuned, tested, debugged and documented processes using technologies SQL, PL/SQL, Informatica, UNIX, and Control-M.
Environment: Oracle 10g, TOAD, SQL Developer, Forms&Reports6i, PL/SQL, UNIX, Informatica 8.x/9.x, UNIX Linux, Windows, SVN, Control-M, Environment: Oracle, SQL Developer, TOAD, Windows 2000/XP, ASP.Net, Visual Studio.
