Spark /big Data Developer, Etl Developer Resume
4.00/5 (Submit Your Rating)
Objective
- To keep up with the cutting edge of technologies, while making a significant contribution to the success of the firm.
SUMMARY
- Total 7.1 years of experience in Data warehousing, includes around 2.5 years of experience in building Big Data applications including 1 Years of Experience in Kafka. using different frameworks like Hadoop, Hive, Sqoop, Spark and Cloud technologies like AWS Redshift, S3 and around 4 .5 years in ETL tools like Informatica PC, SSIS and Datastage.
- Holding extensive experience working on Spark SQL, Dataframes, Spark streaming, Scala, Hive, ETL tools and performance tuning of the Spark applications.
- Having experience in all stages of the project life cycle like requirements gathering, designing & documenting architecture, development and Testing, performance optimization and Production support.
TECHNICAL SKILLS
- Big Data Technologies
- ETL Tool
- NOSQL Databases
- Database
- Kafka
- Data Cleansing Tool
- HDFS
- YARN
- Hive
- Sqoop
- Spark
- Spark Streaming
- Informatica Power Center 9.6.1
- SSIS
- Datastage
- HBASE
- CASSANDRA
- Teradata
- Oracle 11g
- SQL Server 2008
- Informatica Data Quality(IDQ)
- C, Scala
- SQL
- HiveQL
- Unix
- Shell Scripting
PROFESSIONAL EXPERIENCE
Confidential
Spark /Big data Developer, ETL Developer
Responsibilities:
- To build an initial POC to process live data for customer and their related orders coming from Kafka cluster, ingest data using
- Spark Streaming to process, apply certain logic/transformation and feed then in Cassandra, use Apache zeppelin for further analysis and report generation. To understand more of Kafka, was also involved in fixing Kafka server failures.
- Performed source data analysis and data profiling once the data has been moved to Cassandra.
Confidential
Spark Developer
Responsibilities:
- Based on Mapping document, import data from Informatica and store data in Hive based on business logic.
- Query performance optimization, Update data using spark.
- Performance Analysis using spark web UI.
- Write alternative code to optimize Spark performance.
- Validating and testing the code.
- Walked through the the modified code to business
- Involved in building the technical design document and peer review of the code
Confidential
Big data Developer
Responsibilities:
- Designed the complete data pipeline to pull batch data from SQL Server, Amazon S3 bucket & internal SFTP and put them into
- AWS Redshift data in one staging table as Raw data for further analysis, designed Spark Jobs to achieve this.
- Worked as per given requirement document, mapping document and adhoc requests to load in different dimension and fact tables as part of Target table load, create Spark Jobs to process above data, which was further scheduled via AWS Lambda to run.
- Modularize the code
- Perform unit testing
- Design Technical design document
- Production Support
Confidential
ETL Developer
Responsibilities:
- Designing effective ETL architecture which matches the requirement and provides the required outputs.
- Developing and Troubleshooting of Mappings and Workflows for ETL Jobs based on the Low Level design documents and according to the standards of client.
- Work with client onshore Manager to understand and determine Business priorities and determine release dates for different phases/backlog items.
- Developed comprehensive data loading strategies using detailed Visio diagrams - communicating with various end user teams, and understanding their data needs.
- Completely involved right from the beginning of analysing the existing system and functionality requirements, review of all deliverables
- Code migrations among various environments- SIT, UAT and Production ensuring successful and timely delivery of the code for Production users.
Confidential
ETL Developer
Responsibilities:
- Responsible for the complete Design, Development, validation and testing of interfaces by ensuring specific standards and consistency.
- Preparation of Low Level design documents using Microsoft Visio based on the business logic and requirement.
- Dealt with high volume of data and hence leveraged the utilities of Teradata for performance enhancement.
- Migration of ETL jobs into testing and pre-production repository environments.
- Code walk-through to Clients before migration to Production.
Confidential
ETL Developer
Responsibilities:
- Responsible for analyzing, designing and developing ETL strategies and processes.
- Work in coordination with other application system (developers), Team Lead to understand the requirements and fetch data correctly.
- Involved in developing Mappings and reusable transformations using Informatica Designer.
- Source system data from the distributed environment was extracted, transformed and loaded into the Data warehouse database using Informatica.
- Extracted Data from flat files and various relational databases to Oracle Data Warehouse Database.
- Created the unit test cases for mappings developed and verified the data.
- Conducted the peer review for mappings developed and tested.
- Most of the transformations were used like the Source qualifier, expressions, and joiner, Aggregators, Lookups, Filter, Rank and Update Strategy.
- Developed and modified changes in mappings according to business Logic.
- Used Informatica Workflow Manager to create, schedule, monitor sessions and send pre and post session emails to communicate success or failure of session execution.
