We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

PROFESSIONAL EXPERIENCE:

Confidential

Data Engineer

Responsibilities:

  • Involved in the entire project life cycle, from design discussions to production deployment. Expertise in Cluster Analysis using various big data analytic tools such as MapReduce and Hive Supported the Hadoop Architect team in developing a Database Design in HDFS using HBase Architecture Design. Used Spark applications to ease the transition to Hadoop. Developed scripts and batch jobs for scheduling various Hadoop programs and was involved in the maintenance and review of Hadoop log files. Used Sqoop for data import and export from RDBMS to HDFS Extracted the required data from the server into
  • HDFS and Bulk Loaded cleaned data into HBase. Built a data pipeline for a compliance report proof - of-principle consisting of Sqoop, Hadoop (HDFS), SparkSQL, Elasticsearch, and Kibana. Analyzed SQL scripts and designed them by using PySpark SQL for faster performance. Involved in developing shell script, where the logs generated by the users are collected and stored in AWS S3 (Simple storage service) buckets. Worked on JSON schema to define tables and column mapping from AWS S3 data to AWS Redshift and used AWS Data Pipeline for configuring data loads from AWS S3 to AWS Redshift. The open-source Python web scraping framework was used to crawl and extract data from web pages, and the conversion was performed using Hadoop, Hive, and MapReduce. Created Airflow Scheduling scripts in Python. Expert in data pipeline operational services to coordinate clusters and plan workflows.
  • Used Cloudera Manager to continuously monitor and manage the Hadoop cluster. Mappings done with reusable components such as worklets and mapplets, as well as other transformations. Expert in Developing SSIS packages to ETL data into heterogeneous data warehouses. Used python subprocess module to perform UNIX shell commands and extracted data from Agent. Used Kibana visualizations to highlight compliance metrics in the translated report. Wrote Scala user-defined functions for SQL functionality that lacked a Spark-SQL counterpart.

Environment: Hadoop, Cloudera Manager, HDFS, Map Reduce, Hive, Spark SQL, Scala, Python, Oozie, Sqoop, ETL, SSIS, Kibana, RDBMS, UNIX.

Confidential

Data Engineer

Responsibilities:

  • Worked in customer data and Notification management team. Designed and Implemented Extraction, Processing, and storing the data from various sources. Designed and Implemented Data Engineering Solutions on batch data pipeline on - premises. Worked on designing conceptual, logical, and physical data models for the analytical solutions. Worked on translating the customer needs and processing into logical and physical data models, data diagrams, and automated workflows. Used python libraries for Data transformations and handling the voids in the data. ETL and load data to Azure services (Azure
  • Data Lake, Azure Storage, Azure SQL, Azure DW) and process data in Azure Databricks. Worked on developing ETL architecture diagrams. Extracted, Transformed, and loaded data using Cloudera Morphlines. Developed Spark applications using Pyspark and Spark-SQL for data extraction, transformation, and aggregation from multiple file formats for analyzing and transforming data to uncover insights into the customer usage patterns. Developed python scripts to build custom ETL Data pipelines and store the data in Data warehouses. Collected data from websites and mobiles services and integrated it into the
  • Hybrid Connection. Developed SQL queries to store the data into Microsoft SQL Server using Hybrid Connection manager. Implemented Data management frameworks across the client organization. Migrated data from MySQL to Hive using Apache Sqoop in CDP with the Teradata connector. Performed ACID v2 transactions at the row-level. Experience working on Apache HBase shell. Experience using HBase REST API to interact with HBase services, tables, and regions using HTTP endpoints. Designed and configured SSIS packages to migrate the data from Oracle, legacy system using various transformations. Collaborated with the Data warehouse team and Data Analytics team for performing the data models, analysis, and dashboards. Experience working on various Entity relationships and dimension models.

Environment: Cloudera Morphlines, Sqoop, MySQL, Hive, HBase, Python, Entity Relationships, Dimension Models Pyspark, Visio, Oracle, Spark SQL, JSON Format

Confidential

AWS Data Engineer

Responsibilities:

  • Designing and deploying multi - tier applications using all AWS services (EC2, Route53, S3, RDS, Dynamo DB, SNS, SQS, IAM) with an emphasis on high-availability, fault tolerance, and auto-scaling in AWS Cloud Formation. Developing and maintaining an appropriate Data Pipeline design. Responsible for loading data from the internal server and the Snowflake data warehouse into S3 buckets. Created Airflow DAGs for Batch Processing to orchestrate Python data pipelines for csv files preparation pre-ingestion, using conf to parameterize for a multitude of input files from different Hospitals launching separate TaskInstances Wrote Python scripts utilizing the Boto3 library to automatically launch instances on AWS EC2 and OpsWorks stacks, integrating Auto scaling with preset AMIs. Constructed the framework for effective data extraction, transformation, and loading (ETL) from several data sources. Launch Amazon
  • EC2 Cloud Instances and configure launched instances for specific applications using Amazon Web Services (Linux/Ubuntu). Extensive work was performed to migrate data from Snowflake to S3 for the TMCOMP/ESD feeds. Extensively utilized AWS Athena to import structured data from S3 into several systems, including RedShift, and to provide reports. For developing the common learner data model, which obtains data from Kinesis in near real time, we utilized Spark-Streaming APIs to execute on-the-fly conversions and operations. Created Snowflake views for loading and unloading data from and to an AWS S3 bucket, as well as deploying the code to production. Data sources are extracted, processed, and fed into CSV data files using Python programming and SQL queries. Utilized Informatica Power Center Workflow manager to generate sessions, workflows, and batches to execute with the logic encoded within the mappings. Created DAGs in Airflow to automate the process using Python schedule jobs. Analyzed Hive data with the Spark API and the EMR Cluster Hadoop YARN. Enhancements to existing Hadoop algorithms utilizing Spark Context, Spark-SQL, Data Frames, and Pair RDDs. Assisted with the development of Hive tables and the loading and analysis of data using Hive queries. Python was used for exploratory data analysis and data visualization (matplotlib, numpy, pandas, seaborn).

Environment: AWS services, AWS EC2, AWS S3, Dynamo Database, SNS, AWS Athena, Amazon EC2 Cloud Instances, Boto3, AMI, ETL, Hive, Spark API, Spark SQL, Hive Table, Python, Spark, HDFS, Sqoop, MySQL, Linux, Snowflake.

Confidential

Data Analyst

Responsibilities:

  • Identified and documented accurate business rules and use cases based on requirements analysis.* Investigated and resolved issues about data flow integrity into databases.* Use historical data to put data prediction systems to the test.* Analyzed transactions to provide a consistent business intelligence model for real - time reporting needs.* Draft and optimize SQL scripts to evaluate the flow of online quotes into the database and validate the data.* Collecting and iteratively refining management specifications to complement current pivot table reporting with high-quality Excel dashboard graphics.* Application

    SQL, PL/SQL stored procedures, triggers, partitions, Primary Keys, various indexes, constraints, and views are developed and maintained.* Create and run new or existing SQL queries to connect to business intelligence (BI) tools including Jupyter Notebooks (Python), Tableau Dashboard, and Excel reports.* Extracted and analyzed data patterns to translate insights into practical outcomes.* Created multiple Excel documents to help collect metrics data and present information to stakeholders to provide straightforward explanations of optimal resource deployment.* Analyzed datasets provided by more than 20 APIs and developed reports from underlying tables and fields to offer the leadership team a better understanding of crucial data trends.* Designing and delivering interactive dashboards with a range of charts for simple comprehension.* In Tableau, I created joins, relationships, data blending (when merging several sources), calculated fields, Level-of-Detail (LOD) expressions, parameters, hierarchies, sorting, groups, and applied actions (filter, highlight, and URL) to dashboards.

Environment: Jupitar Notebook, SQL, Python, Tableau, Microsoft Excel.

Confidential

Python Developer

Responsibilities:

  • Worked on an online logistics platform to increase data storage efficiency.* Used the Django framework.* Hands - on experience with database issues and connections to SQL and NoSQL databases such as MongoDB via the installation and configuration of various Python packages.* Git, a software version control system, was utilized by a team of programmers to organize and monitor development changes.* Used Python core packages and modules such as NumPy to boost data processing and analysis efficiency.* Developed REST APIs for web applications to improve web system interoperability.* Wrote

    Python code in an integrated development environment, such as Visual Studio Code, to perform Git and Terminal activities more simply.* Worked in a Linux environment and ran Unix-based commands* Developed complex SQL queries and PL/SQL procedures.* Used Python's XML parser framework (SAX) and DOM API to track small amounts of data without the requirement for a database.* Used the Python package Beautiful Soup for web scraping.* Contributed to the automation of VLAN, Trunk port, and Routing configurations.* Designed the Linux Services to run REST web services using a Shell script.* Create an RPM package for the product that allows feature upgrades..* Designed and developed REST API test cases; participated in developing REST API test framework.* Design and build test cases for CLI automation in Python.* Used the PyUnit framework to participate in unit testing and construct unit test cases.

Environment: Django, SQL, NoSQL, MongoDB, Python, REST, XML parser framework, DOM, RPM, PyUnit framework.

We'd love your feedback!