Sr Big Data Developer Resume
0/5 (Submit Your Rating)
NJ
SUMMARY
- 8 years of experience as an IT professional, specialized in Data Integration, Data Ingestion, Data Processing and Big Data ecosystem.
- Experience in development and design of ETL pipelines for supporting data ingestion and processing.
- 3 years of experience using Spark and AWS cloud integration services.
- 3+ years of experience using Talend Data Integration/Big Data Integration (7.2) and SAP Data Services.
- Acquired profound knowledge on spark ecosystem and Architecture in developing production - ready Spark applications utilizing Spark Core, Spark SQL, Data Frames, Datasets.
- Experience in job workflow scheduling and monitoring tools like Glue and Cloud Watch.
- Experience in AWS cloud environment and widely used S3, Athena, Glue, Cloud watch and Step functions for the data integration.
- Experience in using boto3 and pandas’ libraries for handling the large excel data sets in S3.
- Experience in using various IDE, SFTP, Version Control tools.
- Expertise in creating mappings in Talend using Database, File, Error, Orchestration, Processing, Storage, System and other components.
- Expertise in Data modelling techniques like Dimensional/ Star Schema and Snowflake modelling along with Slowly Changing Dimensions (SCD Type 1, Type 2, and Type 3).
- Strong working experience with databases like SAP HANA, SSMS, MySQL, DB2, Oracle, and Snowflake
- Excellent knowledge of Hadoop cluster architecture and its key concepts - Distributed storage and distributed processing, High availability, Fault tolerance and Scalability.
- Excellent working experience in Waterfall, Agile methodologies
- Worked independently and collaboratively. Involved in complex troubleshooting, root-cause analysis, and solution development.
- Self-Starter and Team Player with excellent communication, organizational and interpersonal skills with the ability to grasp things quickly.
- Closely collaborated with business products, production support, Effective decision making in critical situations
- Experience in implementing the end-to-end solution (SDLC) full life cycle.
TECHNICAL SKILLS
Big Data Technologies: HDFS, Apache Spark, Hive, Hue
Operating Systems: Windows, Linux, Unix.
ETL Tools: Talend, Data Services
Languages: Python, SQL, CSS 3, HTML 5.
Databases: Snowflake, MySQL, DB2, SSMS, Oracle, and SAP HANA.
Applications\Tools: Eclipse, Gitlab, WinSCP, Putty, Jenkins, Jira, Confluence, Insomnia, Service Now
AWS Services: AWS S3, EC2, Athena, Glue, Cloud Watch, Step Functions AWS Developer Associate
PROFESSIONAL EXPERIENCE
Confidential, NJ
Sr Big Data Developer
Responsibilities:
- Worked with business to build extremely efficient and reliable data pipelines to move data across several platforms, including cloud storage and Data Lake.
- Created the requirement document by working closely with the business and revised it as per the business suggestion and needs.
- Collaborated with cross-functional teams like infrastructure and admin team for handling requests like SFTP -S3 sync and snowflake user access.
- Used WinSCP and Putty for the secure FTP between cloud and on-premises systems.
- Involved in the design and development of spark data pipelines using AWS services to support the ETL process.
- Experience in using boto3 and pandas’ libraries for handling the large excel data sets in S3.
- This improved the efficiency of the operations by 60 % for the key functional processes.
- Experienced in handling large data sets using Spark in memory capabilities, effective Joins and Transformations.
- Used AWS S3 as a staging layer for raw data and transformed data.
- Worked with non-native binary file formats like Parquet and Snappy to leverage the storage in S3.
- Used Athena on top of Amazon S3 data lake for the data validation.
- Defined complex functions and numerous joins to handle data transformations as per business rules, which ran through AWS Glue to load data into Snowflake.
- Used AWS Glue and Step Functions to orchestrate and schedule spark jobs on AWS Cloud.
- Created tables, procedures, sequences, views for the storage and handling of data in Snowflake.
- Involved in extensive data validation by performing back-end testing to eliminate data quality issues.
- Identified and resolved issues involving data errors during data processing and calculations, and coordinated with other teams such as IT, support, and development to resolve any data processing or client issues.
Confidential, MO
ETL Developer
Responsibilities:
- Working with the business for the requirements gathering and the implementation of the requirements.
- Designed and developed data pipelines to provide an enhanced view of data to the customers through Talend.
- Configured a Kafka consumer to listen to the KAFKA topic and write the delta changes to SAP HANA.
- Design and Implemented ETL for data load from heterogeneous Sources to SAP HANA as target database by using Slowly Changing Dimensions of SCD-Type1 and SCD-Type2.
- Extensively used tMap component which does lookup & Joiner Functions, tjava, txml, tdelimitedfiles, tlogrow components.
- Used Talend most used components (tMap, tDie, tConvertType, tLogCatcher, tRowGenerator, tSetGlobalVar, tHashInput & tHashOutput and many more).
- Created many complex ETL jobs for data exchange from and to Database Server and various other systems including RDBMS, XML, CSV, and Flat file structures.
- Performance tuning - Using the tmap cache properties, Multi-threading and Parallelize components for better performance in case of huge source data. Tuning the SQL source queries to restrict unwanted data in an ETL process.
- Involved in Preparing Detailed design and technical documents from the functional specifications.
- Used Talend Admin Console Job conductor to schedule ETL Jobs on daily, weekly based Cron Triggers.
- Developed Talend jobs to load the health care, crop science sources from Snowflake and Azure API respectively, and to send back the matched data to the corresponding destinations.
- Worked with Aggregate and Analytical functions like Rank and Dense rank to perform data integration.
- Created Joblets and user-defined routines to implement and inherit them into multiple jobs.
- Created local, and global Context variables in the jobs.
- Worked on various Talend components such as tMap, tFilterRow, tAggregateRow, tFileExist, tFileCopy, tFileList, tDie etc.
- Used Gitlab for Configuring the builds and managing the application releases by using version control.
- Collaborated with the nearshore for the application support and handling service requests.
Confidential, NJ
ETL developer
Responsibilities:
- Developed data pipelines using Talend to load the data from SSMS tables into the data lake.
- Created parent child hierarchical jobs and passed context parameters dynamically.
- Implemented delta capture technology using SQL procedures to load the delta changes into data warehouse dimensions.
- Created Tables, Indexes, Views, Stored Procedures and Packages in SSMS database.
- Created mappings in Talend using Database, File, Error, Orchestration, Processing, Storage, System and other components.
- Responsible for improving the jobs performance by modifying the mappings and parallelized the sessions using performance-tuning techniques.
- Developed email notification functionality, which consolidates all the errors and sends them after every run, which helps for easy debugging.
- Deployed the jobs using Nexus and created job execution plans to schedule and execute multiple tasks in TAC.
- Designed jobs to cleanse the server logs after every 15 days to reduce the load on the application server.
- Coordinated with multiple partner teams for the support handover and the MDM knowledge transfer.
- Migration Request using HP Quality Centre and Incidents regarding project were maintained.
- Created Test cases for the mappings modified and then created Unit Test Case Document.
- Did the defect analysis and provided quick solutions to maintain uninterrupted testing of various mappings/workflows/scripts/Jobs.
Confidential, Detroit
SQL developer
Responsibilities:
- Designed, developed, and maintained MySQL relational database (RDBMS) for the project.
- Involved in gathering software requirements from clients and writing complex queries based on the requirement.
- Developed JAVA APIs and used My SQL-JDBC connector to set up data pipelines connecting Google Drive.
- Designed a database to store the patient data with over 20 million records feeding from the smartwatch.
- Created stored procedures to reduce the overall load and improve the performance of the database.
- Unit testing maintenance after completing the development and documented it.
- Prepared audit architecture for reporting team to monitor job analysis on maintaining data lake.
Confidential, Detroit
SQL Developer
Responsibilities:
- Migration of current webpage based out of Oracle to Create/Process and show data related to the customer in the frontend application
- Migrated all completed orders from an old database to the new interface
- Prepared queries to show data in the front end like old pages for users and customers.
- Take care of migration activities like Dev to PROD Deployment by maintaining the code in GitLab through putty.
- Migration Request using HP Quality Centre and Incidents regarding the project was maintained.
- Created Test cases for the modified requests and then created Unit Test Case Document.
- Did the defect analysis and provided quick solutions to maintain uninterrupted testing of various mappings /workflows/scripts/Jobs.
