We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

Pleasonton, CA

SUMMARY

  • Over 8 years of working experience in IT Industry as Data Engineer with the help of big data, Cloud technology with proficiency in Data Modeling/Data Analysis.
  • Extensive use of cloud computing infrastructure such asAmazon Web Services (AWS), GCP and Azure.
  • Adept Software Development Life Cycle (SDLC) and various methodologies such as Agile and Waterfall
  • In - depth knowledge of. Snowflake Database and Table structures.
  • Good exposure to Python programming.
  • Strong experience with Big Data processing using Hadoop technologies Map Reduce, Apache Spark, Hive, Pig.
  • Good understanding of cloud configuration in Amazon web services (AWS).
  • Good experience in writing Spark applications using Python and Scala.
  • Experience working with SQL, PL/SQL and RDBMS/NoSQL databases like Microsoft SQL Server, Oracle, HBase and MongoDB.
  • Experience in using the spark application master to monitor the sparkjobs and capture the logs for the spark jobs.
  • Extensive experience in using ER modeling tools such as Erwin and ER/Studio, Teradata and Netezza.
  • Good knowledge in streaming applications using Apache Kafka.
  • Experience in designing both time driven and data driven automated workflows using Oozie.
  • Hands on experience with data acquisition into Hadoop cluster usingSqoop.
  • Experience with ETL working with Hive and Map-Reduce.
  • Experience in Working Dimensional Data modeling, Star Schema/Snow flake schema, Fact & Dimensions Tables.
  • Excellent technical and analytical skills with clear understanding of design goals of ER modeling for OLTP and dimension modeling for OLAP.
  • Having experience in developing a data pipeline using Kafka to store data into HDFS.
  • Experience inoptimizing Hive queriesby tuning configuration parameters.
  • Experienced in performing real time analytics on HDFS using HBase.
  • Knowledge of Data Cleaning, Exploratory Data Analysis, Data Visualization, and Data Mining using SQL/Python.
  • Capable in DAX Expressions, Power BI Power Pivot and Power integrated with Share Point, and in creating dashboards in Power BI and Tableau.
  • Good experience in design the jobs and transformations and load the data sequentially & parallel for initial and incremental loads in Talend.
  • Experience inData transformation, Data mappingfrom source to targetdatabase schemas, Data Cleansing procedures.
  • Experience in developing and scheduling ETL workflows in Hadoop using Oozie.
  • Proficient in Data extraction, Data cleaning, Data Loading, Statistical Data Analysis, Data Wrangling, Predictive Modeling using Python.
  • Performing extensivedata profilingandanalysisfor detecting and correcting inaccurate data from the databases and to trackdata quality.
  • Experience in designing a component usingUML Design-Use Case, Class, Sequence,andDevelopment, Component diagrams for the requirements.
  • Strong experience in using Excel and MS Access to dump the data and analyze based on business needs.
  • Good expertise knowledge with theUNIX commandslike changing the permissions of the file to file and group permissions.
  • Good communication skills, work ethics and the ability to work in a team efficiently with good leadership skills.

TECHNICAL SKILLS

Big Data tools: Hadoop 3.3, HDFS, Hive 3.1.2, Kafka 3.0, Scala, Oozie, HBase2.3, Sqoop 1.4

Data Modeling Tools: Erwin 9.8/9.7, ER/Studio V17

Cloud Services: Azure DevOps, Azure Synapse, Azure Data Lake, Azure Data Factory and, AWS, Amazon Redshift, Kinesis, GCP and BigQuery

NoSQL Databases: HBase and MongoDB

Scripting Languages: Python, Java, Scala, R, PowerShell Scripting, HiveQL.

Project Execution Methodologies: JAD, Agile, SDLC, Waterfall, and RAD

Database Tools: Oracle 12c/11g, Teradata15/14, Netezza, SQL Server, MySQL, Snowflake DB.

Reporting tools: SQL Server Reporting Services (SSRS), Tableau, Crystal Reports, Strategy, Business Objects

ETL/BI Tools: SSIS, Informaticav10. Snowflake, Informatica, Talend, SSRS, SSAS, ER Studio, Tableau, Power BI.

Programming Languages: SQL, T-SQL, UNIX shells scripting, PL/SQL.

Operating Systems: Microsoft Windows 10/8/7, UNIX

Version Control: Git, SVN, Bitbucket.

PROFESSIONAL EXPERIENCE

Confidential, Pleasonton, CA

Data Engineer

Responsibilities:

  • Responsible for building scalable and distributed data solutions using Cloudera CDH.
  • Worked with Impala for massive parallel processing of queries for ad-hoc analysis. Designed and developed complex queries using Hive and Impala for a logistics application.
  • Created Bash scripts to add dynamic partitions to Hive staging tables. Responsible for loading bulk amount of data into HBase using MapReduce jobs.
  • Loaded data from web servers using Flume and Spark Streaming API. Used flume sink to write directly to indexers deployed on cluster, allowing indexing during ingestion.
  • Involved in migrating large amounts of data from on-prem Cloudera cluster to EC2 instances deployed on Elastic MapReduce (EMR) cluster.
  • Gathered data and performed analytics using Azure Data Factory, Azure Data bricks, Azure Data Lake.
  • Developed an ETL pipeline to extract archived logs from disparate sources and stored in S3 data lake. Used AutoSys schedulers for weekly automation.
  • Implemented Spark Java UDF's to handle data quality, filter, and data validation checks.
  • Analyzed and optimized pertinent data stored in Snowflake using PySpark and SparkSQL.
  • Co-ordinated with Kafka team and built an on-premises data pipeline. Supported Kafka Integrations, performance tuning and identified bottlenecks to improve performance and throughput.
  • Used Stream Sets for analytics and involved in debugging and optimizing data pipelines collecting logs and metrics from various application APIs.
  • Involve in creating database schema and objects like tables, views, stored procedures, triggers, packages, and functions to provide structure and maintain dataefficiently.
  • Managed and deployed configurations for the entire datacenter infrastructure using Terraform.
  • Used Cloudera Hue and Zeppelin notebooks to interact with HDFS cluster. Used Cloudera Manager, Search and Navigator to configure and monitor resource utilization across the cluster.
  • Used Arcadia to connect with Impala, designed interactive dashboards and reports for the BI team.
  • Presented creative business insights with KPI reports and delivered actionable insights by identifying significant and correlated variables.
  • Involved in setting up CI/CD pipelines using Jenkins. Worked along with DevOps team and managed the Jenkins integration service with Puppet
  • Also worked on resolving several tickets generated when issues arise in production pipelines.
  • Used Kerberos for authentication and Apache Sentry for authorization.
  • Used Git for version control, Interacted with Onsite team for deliverables.

Environment: Azure Data Factory (ADF v2), Azure Databricks (PySpark), Azure Data Lake, Spark (Python/Scala), Hive, Apache Nifi 1.8.0, Jenkins, Kafka, Spark Streaming, Docker Containers, PostgreSQL, RabbitMQ, Celery, Flask, ELK Stack, MS-Azure, Azure SQL Database, Azure functions Apps, Azure Data Lake, BLOB Storage, SQL server, Windows remote desktop, UNIX Shell Scripting, AZURE PowerShell, ADLS Gen 2, Azure Cosmos DB, Azure Event Hub, Sqoop, Flume, Impala, Kafka, AWS S3, Spark SQL, SQL, Agile Methodology.

Confidential, Neenah, WI

Big Data Engineer

Responsibilities:

  • Using Sqoop to import and export datafrom Oracle and PostgreSQL into HDFS to use for the analysis.
  • Migrated Existing MapReduce programs to Spark Models using Python.
  • Migrating the data from DataLake (Hive) into S3 Bucket.
  • Done data validation between data present in Data Lake and S3 bucket.
  • Used Spark Data Frame API over Cloudera platform to perform analytics on hivedata.
  • Used Kafka for real-time dataingestion.
  • Created different topics for reading the data in Kafka
  • Read data from different topics in Kafka.
  • Moved data from S3 bucket to Snowflake DataWarehouse for generating the reports.
  • Written Hive queries for data analysis to meet the business requirements.
  • Developed Latin scripts to extract thedatafrom the webserver output files and to load into HDFS.
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs and Scala.
  • Implementing different performance optimization techniques such as using distributed cache for small datasets, partitioning, and bucketing in the Hive, doing map side joins, etc.
  • Good knowledge of Spark platform parameters like memory, cores, and executors
  • By using Zookeeper implementation in the cluster, provided concurrent access for Hive Tables with shared and exclusive locking.

Environment: Linux, Apache Hadoop Framework, HDFS, YARN, HIVE, HBASE, AWS (S3, EMR), Scala, Spark, SQOOP.

Confidential, Columbus OH

Data Engineer

Responsibilities:

  • As a Data Engineer reviewed the entire requirement and collect the all necessary documents with team and stockholders.
  • Responsible for data governance rules and standards to maintain the consistency of the business element names in the different data layers.
  • Involved in installation of HDP Hadoop, configuration of the cluster and the eco system components like Sqoop, Hive, HBase and Oozie.
  • Responsible for creating on-demand tables on S3 files using Lambda Functions and AWS Glue using Python and PySpark.
  • Coordinated with team and Developed framework to generate Daily ad-hoc, Report’s and Extracts from enterprise data and automated using Oozie.
  • Implemented Kafka producers create custom partitions, configured brokers and implemented High level consumers to implement data platform.
  • Used Jira as an agile tool to keep track of the stories that were worked on using the agile methodology.
  • Worked closely with business, transforming business requirements to technical requirements
  • Developed complete end to end Big-data processing in Hadoop eco system.
  • Used AWS Cloud with Infrastructure Provisioning / Configuration.
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting on the dashboard.
  • Used AWS glue catalog with crawler to get the data from S3 and perform Sql query operations.
  • Used AWS Glue for the data transformation, validate and data cleansing.
  • Worked on AWS Data Pipeline to configure data loads from S3 to into Redshift.
  • Used JSON schema to define table and column mapping from S3 data to Redshift.
  • Wrote various data normalization jobs for new data ingested into Redshift.
  • Defined and deployed monitoring, metrics, and logging systems on AWS.
  • Designed and developed an entire module called CDC (change data capture) in python and deployed in AWS GLUE using PySpark library and python.
  • Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
  • Worked extensively with Sqoop for importing and exporting the data from HDFS to Relational Database system and vice-versa.
  • Created reports for the BI team using Sqoop to export data into HDFS and Hive.
  • Optimized MapReduce Jobs to use HDFS efficiently by using various compression mechanisms.
  • Written Shell scripts to monitor the health check of Hadoop daemon services and respond accordingly to any warning or failure conditions.
  • Involved in Hadoop cluster task like Adding and Removing Nodes without any effect to running jobs and data.
  • Followed agile methodology for the entire project.
  • Tested raw data and executed performance scripts and Assisted with data capacity planning and node forecasting.
  • Worked with JIRA to follow agile approach.
  • Document process flow diagrams & data strategies for informative presentations with cross functional teams.

Environment: Hadoop 3.0, Python 3.8, AWS, Sqoop 1.4, Glue, HBase 2.3, Redshift, Hive, JSON, HDFS, PySpark 3.0 & Agile/Scrum.

Confidential

Data Analyst/Data Modeler

Responsibilities:

  • Worked as a Data Analyst / Data Modeler I was responsible for all data related aspects of a project.
  • Created Source to Target Mappings (STM) for the required tables by understanding the business requirements for the reports
  • Developed a Conceptual Model and Logical Model using ER/Studio based onrequirements analysis.
  • Generated parameterized queries for generating tabular reports usingglobal variables, expressions, functions,andstored procedures using SSRS.
  • Created mappings using pushdown optimization to achieve good performance in loading data into Netezza.
  • Developed Data mapping, Transformation and Cleansingrules for the Data Management involving OLTP and OLAP.
  • Designed both 3NF data models for OLTP systems and dimensional data models usings tarandsnowflake Schemas
  • Worked on Informatica Utilities Source Analyzer, warehouse Designer, Mapping Designer, Mapplet Designer and Transformation Developer.
  • Created a Data Mapping document after each assignment and wrote the transformation rules for each field as applicable
  • Involved in developing Unix Shell Scripts for automation of ETL process.
  • Worked on data transformations anddata qualityrules.
  • Worked on Data Miningand data validationto ensure the accuracy of the data between the warehouse and source systems.
  • Worked on Normalization and De-Normalization techniques for OLAP systems.
  • Developed and presented Business Intelligence reportsand product demos to the team using SSRS (SQL Server Reporting Services).
  • Designed and Developed PL/SQL procedures, functions and packages to create Summary tables.
  • Worked on visualizing the reports usingTableau.
  • Worked on Performance Tuning of the database which includes indexes, optimizingSQL Statements.
  • Executed change management processes surrounding new releases of SAS functionality
  • Prepared complex T-SQL queries, viewsandstored proceduresto load data into staging area.
  • Participated indata collection, data cleaning, data mining,developing models and visualizations.
  • Wrote complex SQL queries for validatingthe data against different kinds of reports generated by Business Objects.
  • Used Excel with VBA scripting to maintain existing and develop new reports as required by the business.

Environment: ER/Studio, Netezza, SAS, SSRS, SQL, Tableau, UNIX, OLAP, OLTP, VBA, MS Excel.

We'd love your feedback!