We provide IT Staff Augmentation Services!

Lead Data Engineer Resume

0/5 (Submit Your Rating)

SUMMARY

  • Have more than 13 years of experience in implementing and supporting data warehouses of multiple terabytes to peta bytes using tools like Microsoft Azure, Google Cloud, Snowflake, Apache Spark, Informatics and languages like SQL, Python.
  • Have more than 4 years of experience in developing data processing applications using Python
  • Expertise in Microsoft Azure cloud to perform ETL & ELT operations on data and optimize performance by building robust, extensible and reusable framework using technologies like Azure Data Factory, Apache Spark, Azure Data bricks, Azure Synapse and Azure data lake
  • Experience in Developing Spark applications using Spark - SQL, PySpark in Data bricks for data extraction, transformation and aggregation from multiple file formats for analysing & transforming the data to uncover insights into the customer usage patterns.
  • Strong knowledge in developing, debugging, fine tuning data processing jobs using ETL tools like Informatics and DM Express
  • Fine-tuned several complex ETL Reporting applications with a goal of providing faster and more efficient BI platform for business users.
  • In-depth experience in translating key strategic objectives into actionable and governable roadmaps and designs using best practices and guidelines. Worked on all facets of software development life cycle.
  • Good understanding of spark architecture with Azure Data bricks, Hands on experience with Data bricks, Managing Data bricks notebooks
  • Tuning the performance of existing Azure data bricks pipelines to improve the load times
  • Successfully led executed several Simplification Process Optimization initiatives to bring efficiency into business processes.
  • Experienced in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
  • Experience in enforcing Data Quality, Data Validation standards during the data load process
  • Experience in data modelling and designing OLTP and OLAP systems
  • An effective communicator with excellent relationship building interpersonal skills. Follow effective time management and goal driven approaches. Strong analytical, problem solving organizational abilities
  • Have strong relational Database concepts and worked on several databases such as SQL and DB2.
  • Hands-On experience in designing and developing scalable applications to process data in multiple terabytes to petabytes using tools like Microsoft Azure, Azure Data Factory, Azure Data bricks, Apache Spark, Apache Hadoop, Hive, Pig
  • Experience in processing data from large files and preparing dashboards using python libraries like Pandas, Numpy, Matplotlib
  • Microsoft certified Data Engineer Associate, Informatics and Oracle SQL certified developer

TECHNICAL SKILLS

Programming Languages: SQL, PL/SQL, Python, UNIX Shell Scripting

DBMS/Databases: Oracle 8i/9i/10g, Microsoft SQL Server, DB2, My SQL 4.x/5.x,Snowflake

Big Data Ecosystem: Apache Spark, Spark SQL, Hive, Pig, Hadoop, Airflow

Methodologies: Agile, Water Fall.

ETL Tools: Informatica, DMExpress and DBT

Cloud Technologies: Microsoft Azure, Google Cloud, Snowflake

PROFESSIONAL EXPERIENCE

Confidential

Lead Data Engineer

Responsibilities:

  • Analysing different database options available and identifying the right fit for the needs
  • Designing and developing ETL jobs to ingest inbound files using Azure Data Factory, Azure Databricks
  • Developing spark application using pyspark, spark-sql in Databricks for data extraction, transformation and aggregation
  • Developing scalable code in python to ingest multiple files and prepare them for loading into database
  • Designing the database for best performance
  • Maintaining the data pipelines in production, to make sure the end users always have accurate data
  • Reaching out to source teams to share requirements of dashboard and establishing pipeline to receive the data on monthly basis

Confidential

Sr. Data Engineer

Responsibilities:

  • Designing cloud migration strategies, data delivery architecture
  • Determining appropriate tools to use for migration purpose
  • Performs long-term evaluations of systems, databases, and solutions for security risks. Considers innovative methods and designs to continuously improve our technical capabilities / solutions
  • Model processes to clarify technical designs, and to enhance or re-engineer business processes, prior to, or in parallel with, solution design and implementation, as necessary
  • Building data pipelines in cloud to meet business requirements using pyspark and Azure Databricks
  • Developing python scripts to copy data between on prem and cloud databases
  • Built data pipelines using Spark framework that are extensible and reusable in cloud to ingest and transform structured/semi-structured data.
  • Post migration data testing using tools like python

Confidential

Sr. Data Engineer

Responsibilities:

  • Designing and developing ETL jobs to ingest inbound files using Azure Data bricks, spark, spark-sql
  • Supporting and improving the existing data pipelines in production
  • Developing scalable code in python to ingest multiple files and prepare them for loading into database
  • Maintaining the data pipelines in production, to make sure the end users always have accurate data
  • Performs long-term evaluations of systems, databases, and solutions for security risks. Considers innovative methods and designs to continuously improve our technical capabilities / solutions
  • Model processes to clarify technical designs, and to enhance or re-engineer business processes, prior to, or in parallel with, solution design and implementation, as necessary
  • Working on pilot programmes with vendors, to share the data as per the requirements
  • Developing pipelines to share data with external vendors as per business needs

Confidential

Data Engineer

Responsibilities:

  • Designing the back end database system for the application front end
  • Designing and developing ETL jobs to support the data load process
  • Using Azure synapse, Azure Data Factory and Apache Spark developed data pipelines to Extract & Transform from various source systems (SQL server and flat files) by incorporating business rules using different objects and functions that the tool supports.
  • Implemented slowly changing dimensions (SCD) for some of the Tables as per user requirement.
  • Developed Stored Procedures and used them in Stored Procedure transformation for data processing and have used data migration tools
  • Performance tuning the jobs for better performance
  • Maintaining the data pipelines in production, to make sure the end users always have accurate data within the timeline as per the SLA

Confidential

Sr. ETL Developer

Responsibilities:

  • Understanding existing informatica jobs that loads data into DB2
  • Acquiring knowledge about the existing application in production and handling the production issues and closing them within SLA
  • Working on Proof of Concept, with Hadoop technology to improve the data processing speeds
  • Actively involved in gathering requirements for the enhancements and upcoming projects
  • Developing scripts to create and load data into Hive tables as per the business rules
  • Acquiring knowledge about the existing application in production and handling the production issues and closing them within SLA
  • Providing impact analysis of the requirements on current production system
  • Analyzing the data from different sources based on End user requirements and extensively involved in discussions with end users in decision-making.
  • Represented/lead discussions related to product/application/modules/team (for example, leads technical design reviews). Builds relationships with internal customers/stakeholders. Represent the team in front of the customer
  • Enforcing quality processes such as extensive testing, root cause analysis, code review to deliver the quality product on time as per deadline
  • Providing estimates, resource needs, milestones, risks as per the requirements
  • Developed ETL routines using Informatica Power Center and created mappings involving transformations like Lookup, Aggregator, Ranking, Expressions, Mapplets, SQL overrides usage in Lookups and source filter usage in Source qualifiers and data flow management into multiple targets using Routers.
  • Extensively used Mapping Variables, Mapping Parameters to execute complex business logic
  • Design and development of complex ETL mappings making use of Connected/Unconnected Lookups, Normalizer.
  • Monitoring the jobs on daily basis and undressing the failures immediately.
  • Involving in discussions with reporting team, to identify the data fields and loading the data accordingly

Confidential

ETL Developer

Responsibilities:

  • Analysing the data from different sources based on End user requirements and extensively involved in discussions with end users in decision-making.
  • Represented/lead discussions related to product/application/modules/team (for example, leads technical design reviews). Builds relationships with internal customers/stakeholders. Represent the team in front of the customer
  • Enforcing quality processes such as extensive testing, root cause analysis, code review to deliver the quality product on time as per deadline
  • Providing estimates, resource needs, milestones, risks as per the requirements
  • Developed ETL routines using Informatica Power Center and created mappings involving transformations like Lookup, Aggregator, Ranking, Expressions, Mapplets, SQL overrides usage in Lookups and source filter usage in Source qualifiers and data flow management into multiple targets using Routers to read data from OLTP source systems.
  • Extensively used Mapping Variables, Mapping Parameters to execute complex business logic
  • Design and development of complex ETL mappings making use of Connected/Unconnected Lookups, Normalizer.
  • Monitoring the jobs on daily basis and undressing the failures immediately.
  • Involving in discussions with reporting team, to identify the data fields and loading the data accordingly
  • Actively involved in gathering requirements and acquiring application knowledge from the Business.
  • Analyzed sources and End user requirements and extensively involved in discussions with end users in decision-making.
  • Created ETL detail design document and ETL standards document.
  • Developed ETL routines using Informatica Power Center and created mappings involving transformations like Lookup, Aggregator, Ranking, Expressions, Mapplets, SQL overrides usage in Lookups and source filter usage in Source qualifiers and data flow management into multiple targets using Routers.
  • Extensively used Mapping Variables, Mapping Parameters to execute complex business logic
  • Design and development of complex ETL mappings making use of Connected/Unconnected Lookups, Normalizer.
  • Proficient in using Source Analyzer, Warehouse Designer, Transformation Designer, Mapping Designer and mapplet Designer.
  • Used debugger in debugging some critical mapping by setting breakpoints and trouble shot the issues by checking sessions and workflow logs.
  • Involved in identifying bottlenecks in source, target, mappings and sessions and resolved the bottlenecks by doing Performance tuning techniques like increasing block size, data cache size, sequence buffer length.
  • Developed UNIX shell scripts to create parameter files, rename files, compress files and for prebalancing the flat file extracts.
  • Performed active interaction with the client and effort estimation for the new requirements in the project.
  • Responsible for the documentations of the different processes carried out like design documents and mapping documents, share point for the version control of the documents.
  • Worked closely in setting up the environment for various file transfer activities between the systems using SFTP as the file transfer protocol.

Confidential

Associate

Responsibilities:

  • Responsible for the design, development, coding, testing, debugging and documentation of applications to satisfy the requirements of one or more user areas.
  • Using Informatica PowerCenter Designer, 9.0.1, 8.6, DMExpress 7.1 (Task editor, Job editor, Server) analyzed the source data to Extract & Transform from various source systems (oracle 10g and flat files) by incorporating business rules using different objects and functions that the tool supports.
  • Using Informatica PowerCenter created mappings and mapplets to transform the data according to the business rules to read data from OLTP source systems and write into OLAP systems.
  • Using DMExpress created tasks and jobs to transform the data from flat files according to the business rules.
  • Used various transformations like Source Qualifier, Joiner, Lookup, Sql, Router, Filter, Expression and Update Strategy in Informatica
  • Used various tasks like Join, Sort and Merge in DMExpress.
  • Implemented slowly changing dimensions (SCD) for some of the Tables as per user requirement.
  • Developed Stored Procedures and used them in Stored Procedure transformation for data processing and have used data migration tools
  • Documented Informatica mappings, DMExpress tasks in Excel spread sheet.
  • Tuned the Informatica mappings for optimal load performance.
  • Created and Configured Workflows and Sessions to transport the data to target warehouse Oracle tables using Informatica Workflow Manager.
  • This role carries primary responsibility for problem determination and resolution for each application system
  • Worked along with UNIX team for writing UNIX shell scripts to customize the server scheduling jobs.
  • Constantly interacted with business users to discuss requirements.

We'd love your feedback!