We provide IT Staff Augmentation Services!

Sr. Data Engineer Resume

0/5 (Submit Your Rating)

Branchburg, NJ

SUMMARY

  • 9+ years of experience as a Hadoop/Spark Developer, Data Engineer, and Programmer Analyst in design, development, and deploying large - scale distributed systems.
  • Hands-on experience in installing, configuring, and using Hadoop components like Hadoop MapReduce, HDFS, Yarn, Pig, Hive, HBase, Spark, Kafka, Flume, Sqoop, Oozie, Avro, Spark integration with Cassandra, Solr, and Zookeeper.
  • Proficient in Data Modelling / DW/ Bigdata/ Hadoop/ Data Integration/ Master Data Management, Data Migration and Operational Data Store, BI Reporting projects with a deep focus in design, development and deployment of BI and data solutions using custom, open source and off the shelf BI tools.
  • Logical and physical database designing like Tables, Constraints, Index, etc. using Erwin, ER Studio, TOAD Modeler and SQL Modeler.
  • Good understanding and hands-on experience with AWS S3 and EC2.
  • Experience with event-driven and scheduled AWS Custom Lambda (Python) functions to trigger various AWS resources.
  • Experience in creating Pyspark scripts and Spark Scala jars using IntelliJ IDE and executing them.
  • Good experience on programming languages Python, Scala.
  • Experience in troubleshooting Spark/Map Reduce jobs.
  • Experience in developing Python ETL jobs run on AWS services and integrating with enterprise systems like Enterprise logging and alerting, enterprise configuration management and enterprise build and versioning infrastructure.
  • Experience in using Terraform for building AWS infrastructure services like EC2, Lambda and S3.
  • Have experience in Apache Spark, Spark Streaming, Spark SQL, and NoSQL databases like HBase, Cassandra, and MongoDB.
  • Expertise in configuring the monitoring and alerting tools according to the requirement like AWS CloudWatch.
  • Experience in designing both time driven and data driven automated workflows using Oozie and developing high-performance batch processing applications on Apache Hive, Spark, Impala, Sqoop HDFS.
  • Experienced in using Integrated Development environments like Eclipse, NetBeans, Kate and Gedit. Migration from different databases (i.e., Oracle, DB2, MYSQL, MongoDB) to Hadoop interactive Dashboards and Creative Visualizations using Visualization tools like Tableau, Power BI.
  • Working experience on NoSQL databases like HBase, Azure, SSIS, MongoDB and Cassandra with functionality and implementation.
  • Proficient in programming with SQL, PL/SQL, and Stored procedures.
  • Experience in Database Design, Database Management and Data Migration using Oracle, MS SQL, SQL, and Technical Support.
  • Experience in Developing ETL workflows using Informatica PowerCenter 9.X/8.X and IDQ. Worked extensively with Informatica client tools- Designer, Repository Manager, Workflow Manager, and Workflow Monitor.
  • Experience in working with business intelligence and data warehouse software, including SSAS, Pentaho, Cognos Database, Amazon Redshift, or Azure Data Warehouse.
  • Worked on Informatica Power Centre Tools-Designer, Repository Manager, Workflow Manager.
  • Expertise in broad range of technologies, including business process tools such as Microsoft Project, MS Excel, MS Access, MS Visio.

PROFESSIONAL EXPERIENCE

Sr. Data Engineer

Confidential, Branchburg, NJ

Responsibilities:

  • Designed, developed, and maintained data integration programs in Hadoop and RDBMS environment with both RDBMS and NoSQL data stores for data access and analysis.
  • Used all major ETL transformations to load the tables through Informatica mappings.
  • Created Hive queries and tables that helped line of business identify trends by applying strategies on historical data before promoting them to production.
  • Installed Hadoop, Map Reduce, HDFS, AWS and developed multiple Map Reduce jobs in PIG and Hive for data cleaning and pre-processing.
  • Implemented solutions for ingesting data from various sources and processing the Data-at-Rest utilizing Big Data technologies such as Hadoop, Map Reduce Frameworks, HBase, Hive.
  • Implemented Spark GraphX application to analyse guest behaviour for data science segments.
  • Exploring with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Pair RDD's, Spark YARN.
  • Worked on batch processing of data sources using Apache Spark, Elastic search.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Scala.
  • Worked on migrating PIG scripts and Map Reduce programs to Spark Data frames API and Spark SQL to improve performance.
  • Developed the Talend mappings using various transformations, Sessions and Workflows. Teradata was the target database, Source database is a combination of Flat files, Oracle tables, Excel files and Teradata database.
  • Created Hive External tables to stage data and then move the data from Staging to main tables.
  • Implemented Installation and configuration of multi-node cluster on Cloud using Amazon Web Services (AWS) on EC2.
  • Created Data Pipelines as per the business requirements and scheduled it using Oozie Coordinators.
  • Worked with NoSQL database HBase in getting real time data analytics.
  • Able to assess business rules, collaborate with stakeholders and perform source-to-target data mapping, design, and review.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as MapReduce Hive, Pig, and Sqoop.
  • Created scripts for importing data into HDFS/Hive using Sqoop from DB2.
  • Loading data from different source (database & files) into Hive using Talend tool.
  • Conducted POC's for ingesting data using Flume.
  • Objective of this project is to build a data lake as a cloud-based solution in AWS using Apache Spark.
  • Prepared the complete data mapping for all the migrated jobs using SSIS.
  • Created SSIS Packages using SSIS Designer for export heterogeneous data from OLE DB Source (Oracle)
  • Created SSIS packages for File Transfer from one location to the other using FTP task.
  • Worked on Data modelling, Advanced SQL with Columnar Databases using AWS.
  • Worked on Sequence files, RC files, Map side joins, bucketing, Partitioning for Hive performance enhancement and storage improvement.

Environment: Hadoop, Cloudera, SSIS, Talend, Scala, Spark, HDFS, Hive, Pig, Sqoop, DB2, SQL, Linux, Yarn, NDM, Informatica, AWS, Windows & Microsoft Office, MS-Visio, MS-Excel.

Sr. Data Engineer

Confidential, Weehawken, NJ

Responsibilities:

  • Evaluate, extract/transform data for analytical purpose in a Big data environment. Involved in designing ETL processes and developing source to target mappings.
  • In-depth understanding of Spark Architecture including Spark Core, Spark SQL, Data Frames.
  • Developed spark application by using python (pyspark) to transform data according to business rules.
  • We have worked on many ETL's creating aggregate tables by doing many transformations and actions like JOINS, sum of the amounts etc.
  • Worked with building data warehouse structures, and creating facts, dimensions, aggregate tables, by dimensional modeling, Star and Snowflake schemas.
  • Used Spark-SQL to Load data into Hive tables and Written queries to fetch data from these tables. Implemented partitioning and bucketing in hive
  • Experienced in performance tuning of Spark Applications for setting correct level of Parallelism and memory tuning.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins.
  • Sourced Data from various sources (Teradata, Oracle) into Hadoop Eco system using big data tools like sqoop.
  • Imported data from Snowflake query and into Spark Dataframes and performed transformations and actions on Dataframes.
  • Developed shell script to install snowflake jars, Python packages and spark executors from artifacts.
  • Worked on Exporting and analyzing data to the Snowflake using for visualization and to generate reports for the BI team. Use Amazon Elastic Cloud Compute (EC2) infrastructure for computational tasks and Simple Storage Service (S3) as storage mechanism.
  • Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in In Azure Databricks.
  • Design and implement database solutions in Data Warehouse, snowflake.
  • Good understanding of Spark Architecture with Databricks, Structured Streaming. Setting Up AWS and Microsoft Azure with Databricks, Databricks Workspace for Business Analytics, Manage Clusters In Databricks, Managing the Machine Learning Lifecycle.
  • Responsible for loading Data pipelines from web servers using Kafka and Spark Streaming API.
  • Responsible for estimating the cluster size, monitoring, and troubleshooting of the Spark DataBricks cluster.
  • Capable of using AWS utilities such as EMR, S3 and Cloud Watch to run and monitor Hadoop and Spark jobs on AWS.
  • Creating S3 buckets also managing policies for S3 buckets and utilized S3 bucket and Glacier for storage and backup on AWS. Spin up an EMR Cluster to run all Spark jobs.
  • Deployed Bigdata Hadoop application using Talend on cloud AWS.

Environment: Apache Hadoop, Hive, Map Reduce, SQOOP, Spark, Python, AWS, Databricks, Azure Data Bricks, Azure Data Factory, Delta Lake, HDFS, Oozie, Putty.

Data Engineer

Confidential, Vernon Hills, IL

Responsibilities:

  • Worked on implementation of a log producer in Scala that watches for application logs, transform incremental log, and sends them to a Kafka and Zookeeper based log collection platform.
  • Used Talend for Big Data Integration using Spark and Hadoop.
  • Worked in developing Pig Scripts for data capture change and delta record processing between newly arrived data and already existing data in HDFS.
  • Optimized Hive queries to extract the customer information from HDFS.
  • Used Polybase for ETL/ELT process with Azure Data Warehouse to keep data in Blob Storage with almost no limitation on data volume.
  • Analysed large and critical datasets using HDFS, HBase, MapReduce, Hive, Hive UDF, Pig, Sqoop, Zookeeper and Spark.
  • Loaded and transformed large sets of structured, semi structured, and unstructured data using Hadoop/Big Data concepts.
  • Performed Data transformations in HIVE and used partitions, buckets for performance improvements.
  • Developing Spark scripts, UDF's using both Spark DSL and Spark SQL query for data aggregation, querying, and writing data back into RDBMS through Sqoop.
  • Designed and developed a Data Lake using Hadoop for processing raw and processed claims via Hive and Informatica.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster processing of data.
  • Ingested data into HDFS using SQOOP and scheduled an incremental load to HDFS.
  • Using Hive to analyse data ingested into HBase by using Hive-HBase integration and compute various metrics for reporting on the dashboard.
  • Created ETL/Talend jobs both design and code to process data to target databases.
  • Wrote Pig Scripts to generate Map Reduce jobs and performed ETL procedures on the data in HDFS.
  • Responsible for the Extraction, Transformation and Loading of data from Multiple Sources to Data Warehouse using SSIS.
  • Experienced in loading the real-time data to NoSQL database like Cassandra.
  • Performed Logging and Deployment of various packages in SSIS.
  • Implemented various SSIS packages having different tasks and transformations and scheduled SSIS packages.
  • Developing scripts in Pig for transforming data and extensively used event joins, filtered, and done pre- aggregations.

Environment: Spark, YARN, HIVE, Pig, Scala, Mahout, NiFi, Python, Hadoop, Azure, Dynamo DB, Kibana, NOSQL, Sqoop, SSIS, MYSQL.

Big Data developer

Confidential, Columbia, SC

Responsibilities:

  • Involved in designing data warehouses and data lakes on regular (Oracle, SQL Server) high performance on big data (Hadoop - Hive and HBase) databases. Data modelling, Design, implement, and deploy high-performance, custom applications at scale on Hadoop /Spark.
  • Generated ad-hoc SQL queries using joins, database connections and transformation rules to fetch data from legacy DB2 and SQL Server database systems.
  • Participated in supporting Data Governance, and Risk Compliance platform utilizing Mark Logic.
  • Participated in the Data Governance working group sessions to create Data Governance Policies.
  • Loaded data into MDM landing table for MDM base loads and Match and Merge.
  • Designed ETL process using Talend Tool to load from Sources to Targets through data Transformations.
  • Translated business requirements into working logical and physical data models for OLTP &OLAP systems.
  • Data Modeler/Analyst in Data Architecture Team and responsible for Conceptual, Logical and Physical model for Supply Chain Project.
  • Created and maintained Logical &Physical Data Models for the project. Included documentation of all entities, attributes, data relationships, primary and foreign key structures, allowed values, codes, business rules, glossary terms, etc.
  • Owned and managed all changes to the data models. Created data models, solution designs and data architecture documentation for complex information systems.
  • Developed Advance PL/SQL packages, procedures, triggers, functions, Indexes and Collections to implement business logic using SQL Navigator.
  • Worked with Finance, Risk, and Investment Accounting teams to create Data Governance glossary, Data Governance framework and Process flow diagrams.
  • Involved in Extract, Transform and Load (ETL) data from spreadsheets, flat files, database tables and other sources using SQL Server Integration Services (SSIS) and SQL Server Reporting Service (SSRS) for managers and executives.
  • Designed Star Schema Data Models for Enterprise Data Warehouse using Power Designer.
  • Experienced in Mark Logic Infrastructure sizing assessment, and Hardware evaluation.
  • Developed the Talend jobs and make sure to load the data into HIVE tables & HDFS files and develop the Talend jobs to integrate with Teradata system from HIVE tables
  • Created the best fit Physical Data Model based on discussions with DBAs and ETL developers.
  • Created conceptual, logical, and physical data models, data dictionaries, DDL and DML to deploy and load database table structures in support of system requirements.
  • Designed ER diagrams (Physical and Logical using Erwin) and mapping the data into database objects.

Environment: OLTP, DBAs, DDL, DML, UML, diagrams, Snow-flake schema, SQL, Data Mapping, Metadata, OLTP, SAS, Informatica 9.5, MS-Office

Data Analyst

Confidential

Responsibilities:

  • Work with functional colleagues to develop interactive dashboards and management reports using business intelligence software.
  • This includes data quality, data management, data policies, business process management, and risk management related to the processing of business data.
  • Performed exploratory data analysis for health extractions in Python (Pandas, NumPy & Plotting libraries).
  • Objects involved in creating a database, such as tables, views, procedures, triggers, functions, etc. that use T-SQL to provide definitions and structures to efficiently manage data.
  • Create and publish custom interactive reports and dashboards and schedule reports in Tableau Server.
  • Create action filters, parameters, and calculation sets to prepare dashboards and worksheets in Tableau.
  • Developed Tableau visualizations and dashboards using Tableau Desktop.
  • SQL queries fine-tuned for maximum efficiency and performance.
  • Design database schemas, create tables to store data, and perform data operations.
  • Develop, maintain, and execute SQL content / questions and approve the simplicity, reliability, and accuracy of information.

Environment: SQL, Tableau, Python (Pandas, NumPy & Plotting libraries)

We'd love your feedback!