Sr. Data Engineer Resume
Charlotte North, CarolinA
SUMMARY
- Around 8+ years of IT experience as a Data Engineer, Hadoop and application developer including profound expertise and experience on statistical data analysis such as transforming business requirements into analytical models, designing algorithms, and strategic solutions that scales across massive volumes of data.
- Worked in various projects like Data Lake, Application migrations, Cloud migrations, and Automation projects for various clients.
- Extensive experience in Database Development in Spark, Hadoop, Python, Oracle, DB2, SQL Server, Teradata, Data bricks, MySQL.
- Experienced working on cloud - based technologies like AWS which includes EMR, EC2, S3, Cloud monitoring and different data bases provided by AWS to manage application end to end.
- Experience in launching services like EC2, EMR and triggering auto scaling to the application.
- Experience in AWS IAM roles to perform actions and give permission access to multiple roles and s3 buckets.
- Strong working experience in scalability, high availability, Disaster Recovery, Cloud Infrastructure design, Providing solutions on Cloud technologies and services.
- Experienced working on Pyspark data frames to perform data validations, data analytics on the cloud.
- Experienced in writing code in Python to manipulate data for data loads, extracts, statistical analysis, modeling, and data munging.
- Experienced on python third party libraries like java pandas, NumPy and Dask data frames.
- Hands on experience in developing Stored Procedures, Functions, Views and Triggers, Complex SQL queries using SQL and TSQL.
- Experience to handle enormous amounts of data for Data Cleansing, Data Profiling Data Scrubbing, Data lineage and Data Quality.
- Extensive knowledge on Version control tools (Git, GitHub, Gitlab) and Incident/Defect tracking tools (Service Now).
- Proficient in developing Web Services, RESTful in Python using XML, JSON.
- Experienced in various types of testing such as Unit testing, Integration testing, User acceptance testing, Functional testing.
- Experienced working on Kafka services to manage and set up Kafka clusters to build a pipeline.
- Expert in Python scripting, worked in stats function with NumPy, visualization using Matplotlib and Pandas for organizing data.
- Strong understanding of Data Modeling and ETL process in data warehouse environment such as star schema, snowflake schema.
- Extensive experience in Tableau Desktop reporting features like Measures, Dimensions, Folder, Hierarchies, Extract, Filters, Table Calculations, calculated fields, Sets, Groups, Parameters, Forecasting, Blending and Trend Lines.
TECHNICAL SKILLS
Big Data Ecosystem: HDFS, MapReduce, HBase, Pig, Hive, Sqoop, KafkaFlume, Cassandra, Impala, Oozie, Zookeeper, MapR, Amazon Web Services (AWS), EMR
Cloud Technologies: AWS, Azure, Google cloud platform (GCP)
IDE’s: IntelliJ, Eclipse, Spyder, Jupyter.
OperatingSystems: Linux, Unix, Windows 8, Windows 7, Windows Server 2008/2003
Programming languages: Python, Scala, Linux shell scripts, Java Scripting, PL/SQL, Java, Pig Latin, HiveQL
Databases: Oracle, MySQL, DB2, MS-SQL Server, MongoDB, HBASE
Web Dev. Technologies: HTML, XML, JSON, CSS, JQUERY, JavaScript
Java Technologies: Core Java, Servlets, JSP, JDBC, Java Beans, J2EE
Business Tools: Tableau, Power BI
PROFESSIONAL EXPERIENCE
Sr. Data Engineer
Confidential - Charlotte, North Carolina
Responsibilities:
- The project is mainly focused on Retail data, Migrating to AWS cloud from Teradata, snowflake, and internal databases.
- We built a pipeline (True Source) to handle huge amount of data which comes through our internal and external vendors were pushed mainly Retail bank data to load directly to the cloud.
- Built AWS infrastructure, in such a way that NIFI, KYLO, EMR which work seamlessly in contrast with each other.
- Hands on writing code in Python and Pyspark to manipulate data for data loads, extracts, statistical analysis, modeling, and data validation.
- Developed streaming & batch applications using Apache spark and Python.
- I was mainly part of doing true sourcing the data through our pipeline includes various techniques to handle the data which are mainly Data quality checks, schema validation, row count validation, data standardization and data lineage.
- To manage and store the data we are working on cloud-based technology like AWS which includes all the services from AWS like Ec2 instances, EMR and S3 services.
- Developed Python and Java scripts to manage AWS resources from API calls using BOTO3 SDK and worked with AWS CLI.
- Created custom CFT (cloud Formation Template) for EMR, Lambda, EC2 in AWS as per requirement.
- Responsible for launching Amazon EC2 cloud instances using AWS services and configuring launched instances with respect to specific application and regions.
- Developed Python scripts to transfer files between cross-region s3 buckets.
- Written python code to Encrypt and Decrypt the sensitive data files using aws services and store it to S3 buckets.
- Written python custom Regex functions to mask the non-public information and plastic card info data.
- Worked on Java code to convert ebcdic format files to delimited formats.
- To read the big files and different format files like parquet, ebcdic and text format files we use python, pandas, and spark to validate the millions records of data.
- In Python I worked on Pandas, NumPy and Dask data frames to analyze millions of transactions and perform validation reports on the data.
- Mostly for the large files I used Pyspark data frames to analyze and save time to perform validations on the data.
- Converting SQL codes to Spark codes using Scala, PySpark, and Spark -SQL for faster testing and processing of data.
- Written custom API calls to register metadata and created tables for the end users to consume date from cloud.
- In pandas used loc and iloc functions to filter the data based on conditions and performed data slicing and data indexing on the big files.
- Written custom scripts in Python and Pyspark to make refined and transformed files as per the consumers requests.
- Designed and developed end to end Infrastructure on aws and data transformation scripts in pyspark to integrate our platform with Enterprise Streaming Data Platform.
- Worked on Jenkins continuous integration tool for deployment of project.
- Perform data analysis and data modeling to support the business user's needs.
- Backup databases and test the integrity of the backups
Environment: Agile, Python, pandas, Jupyter Notebook, Java, Pyspark, Snowflake, EMR, EC2, AWS S3, Kylo, Nebula, NIFI, MS Visio, Word, Excel, PowerPoint, SQL, SAS, SharePoint 2010, MS Project.
Data Engineer
Confidential, St. Louis, Missouri
Responsibilities:
- Design conceptual data models by analyzing the data requirements needed to support the business processes.
- Developed python scripts for data cleaning, transforming, and extracting important metrics from streaming / log data to ensure the successful daily load of raw data to the server.
- ETL Data Cleansing, Integration & Transformation using python scripts for managing data from disparate sources.
- Worked on python code to extract data from ETL and load it to cloud.
- Worked on launching an application on cloud which includes AWS resources like EMR, EC2 and S3 services.
- Transformation of the logical data model to a physical data model and working with DBAs to create the data warehouse(redshift) on cloud.
- Developed Map reduce jobs, Hive queries and Python scripts for analyzing large datasets, data cleaning and transformations.
- Composed python scripts to parse XML and JSON reports and load the information in database.
- Worked on Big Data Integration &Analytics based on Hadoop, Spark, Kafka clusters and web Methods.
- Implemented Apache-spark code to read multiple tables from the real-time records and filter the data based on the requirement.
- Implemented Talend for automated loading and processing of data into HDFS which reduced manual loading and processing of data per app basis.
- Had knowledge on Kibana and Elastic search to identify the Kafka message failure scenario
- Worked on the ETL Informatica mappings and other ETL Processes (Data Warehouse).
- Perform data analysis on large volumes of data to identify duplications, data anomalies, missing data, etc.
- Experience in using DML statements to perform different operations on Hive Tables.
- Handled different file formats like EBCDIC, Text files, Sequence files, Avro data files using different Serde properties in Hive.
- Provided 24x7 on-call database support and Backup databases to test the integrity of the backups.
- Work with clients to establish eligibility data exchange and write data mapping specifications for custom file layouts.
- Developed Python modules to automate processes in AWS (AWS cloud formations/ EC2).
- Closely worked with Kafka Admin team to set up Kafka cluster setup on the QA and Production environments.
- Set-up reporting environment, develop strategies for data collection and automation, model the data environment and implement the solution.
Environment: Waterfall, Python, Data Mining, pyspark, Java, MS Visio, Word, Excel, PowerPoint, Kafka, UNIX shell scripting SQL, SharePoint 2010, OLAP, OLTP, MS Project, Mainframe-SAS, TABLEAU.
Data Engineer
Confidential, Sunny vale, CA
Responsibilities:
- Worked effectively on SQL Profiler, Index Tuning Wizard, Estimated Query Plan to optimize the performance tuning of SQL Queries and Stored Procedures.
- Developed tools using python and Shell Scripting to automate some of menial tasks.
- Experienced in developing Web Services with Python programming language.
- Used Spark and Scala for developing machine learning algorithms that analyze clickstream data.
- Implemented a CI/CD pipeline using Jenkins, Airflow for Containers from Docker and Kubernetes.
- Working directly with Development and Data team to validate fields from Source DB to Target DB using the Data Mapping document.
- Involved in development of Web Services using SOAP for sending and getting data from the external interface in the XML format.
- Support current and new services that leverage AWS cloud computing architecture including EC2, S3, and other managed service offerings.
- Created a task scheduling application to run in an EC2 environment on multiple servers.
- Designed built and deployed a set of python modeling APIs for customer analytics, which integrate multiple machine learning techniques for various user behavior prediction and support multiple marketing segmentation programs.
- Used Spark for data analysis and store final computation result to HBase tables.
- Developed and tested many features for dashboard using Python, Java, Bootstrap, CSS, Java script and jQuery.
- Developed a fully automated continuous integration system using Git, Jenkins, MySQL, and custom tools developed in Python and Bash.
- Managed large datasets using Pyspark, Pandas and Dask Data frames.
- Designed front end using HTML, AngularJS, CSS, and JavaScript.
- Written python custom scripts to transform data in to ETL logic and perform the Data driven analysis, Data quality checks and Data profiling.
- Developed a fully automated continuous integration system using Git, Jenkins, MySQL, and custom tools developed in Python and Bash.
- Created interactive data charts on web application using High charts Java script library with data coming from Apache Cassandra.
- Performed Data profiling, preliminary data analysis and handle anomalies such as missing, duplicates, outliers, and imputed irrelevant data.
- Worked effectively on SQL Profiler, Index Tuning Wizard, Estimated Query Plan to optimize the performance tuning of SQL Queries and Stored Procedures.
- Wrote Python scripts to parse XML and JSON documents and load the data in database.
- Troubleshooted data related issues in data warehouse to correct reporting discrepancies.
- Write SQL queries effectively and efficiently - including inner/outer joins, inserts, and table creation.
Environment: python, SQL, Data Warehouse, Rational Tools, Java, Apache MS Visio, MS SharePoint, ETL, Pyspark, Jupyter Notebook, Pandas, Oracle, and XML.
Data Engineer
Confidential
Responsibilities:
- Implemented and followed a Scrum Agile development methodology within the cross functional team and acted as a liaison between the business user group and the technical team.
- Involved in defining the source to target data mappings, business rules and data definitions.
- Created data trace map and data quality mapping documents.
- Performed data analysis and data profiling using complex SQL queries on various sources systems
- Worked on various UNIX commands and shell scripting using VI editor and ultra-edit.
- Performed Exploratory Data Analysis using R Also involved in generating various graphs and charts for analyzing the data using Python Libraries.
- Worked with internal architects and assisting in the development of current and target state enterprise data architectures.
- Designed and developed the UI for the website with HTML, XHTML, CSS, Java Script and AJAX.
- Created SQL scripts using OLAP functions to improve the query performance while pulling the data from large tables.
- Developed SQL queries to perform data extraction from existing sources to check format accuracy.
- Co-developed the SQL server database system to maximize performance benefits for clients.
- Creating Unix scripts for automation.
- Worked with project team representatives to ensure that logical and physical data models were developed in line with corporate standards and guidelines.
- Wrote and executed various MySQL database queries from Python-MySQL connector and MySQL dB package.
- Performed data validation with Redshift and constructed pipelines designed over 100TB per day.
- Working with business users on the new Tableau versions features and explaining self-service capabilities.
- Designed & Developed logical & physical data model using data warehouse methodologies
- Created Summary and detail dashboards for identifying mismatch of the data in Source and reporting systems using Tableau Desktop.
- Performed Data profiling, preliminary data analysis and handle anomalies such as missing, duplicates, outliers, and imputed irrelevant data.
Environment: Rapid SQL, Data Organization, python, Spark, Data Profiling, MVS Assembler, MS Visio, MS Project, MS Office, Windows.
