Data Engineer Resume
Dallas, TexaS
SUMMARY
- Exceptionally effective data engineer around 7 years of expertise in machine learning, data mining with big datasets of structured and unstructured data, data acquisition, data validation, predictive modeling, data visualization, and web crawling and web scraping. skilled in Big Data technologies like Hadoop and Hive as well as statistical programming languages like R and Python.
- Experienced on Hadoop Ecosystem and Big Data components including Apache Spark, Scala, Python, HDFS, Map Reduce, KAFKA.
- Knowledge of the Software Development Life Cycle, SCRUM, Agile, and Waterfall techniques as participation in the development lifecycle of an Agile project using Git and Jenkins for CI/CD.
- Good Exposure on the Map in Apache Hadoop Reduce programming, distribute applications, HDFS, and have solid understanding of Cloudera Hadoop cluster design and cluster monitoring.
- Knowledge of implementing various standards and procedures for the design and deployment of Hadoop based applications.
- Hands on experience in installing, configuring, and using Hadoop ecosystem components like Hadoop Map Reduce, HDFS, HBase, Hive, Spark, Sqoop, Zookeeper and Flume.
- Developed Spark Applications for data extraction, transformation, and aggregation from multiple file formats for analysing and transforming the data to uncover insights into customer usage patterns.
- Experienced in performance tuning of Spark Applications for setting right batch interval time, correct level of parallelism and memory tuning.
- Optimization of existing jobs and improving the performance in Hadoop using Spark context, Spark - SQL and Spark YARN using Scala.
- Involved in migrating MapReduce programs into Spark transformations using Apache Spark and Scala.
- Strong understanding in conversion of SQL queries into Spark Transformations using Spark RDDs, Data Frames and Scala, and performed map-side joins on RDD's.
- Experience in managing Hadoop clusters using Cloudera Manager Tool.
- Involved in importing the real time data to Hadoop using Kafka and implemented the Oozie job for daily imports.
- Experience in development of Hive-Map Reduce-streaming Python modules.
- Creating AWS Lambda functions to run python scripts.
- Ability to work independently as well as in a team and able to effectively communicate with customers, peers and management at all levels in and outside the organization.
- Experience in all the phases of Data warehouse life cycle involving Requirement Analysis, Design, Coding, Testing, and Deployment.
- Extensive knowledge of utilizing cloud-based technologies using Amazon Web Services (AWS), VPC, EC2, Route S3, Dynamo DB, Elastic Cache Glacier, RRS, Cloud Watch, Cloud Front, Kinesis, Redshift, SQS, SNS, RDS.
- Strong Hand’s on in Azure cloud components (HDInsight, Databricks, DataLake, Blob Storage, Data Factory, Storage Explorer, SQL DB, SQL DWH, CosmosDB).
- Well-equipped in Statistical thinking which include Graphical and Quantitative EDA, Correlation, Hypotheses modelling, Collaborative Filtering, Recommender systems, Time-series, Inferential Statistics and data warehousing in scripting language python.
- Expertise in programming languages like scripting language python and R. Skilful in using Python libraries Pandas, NumPy, Mat Plot Lib, Seaborn, Scikit Learn, NLTK, Keras, Tensor Flow.
- Work on data integration and ingestion from SAP, Oracle, SQL Server source systems and EDW into Hadoop.
- Worked on setting up key components for the project like Kerberos authentication renewals, password encryption mechanism in Hadoop and creation of environment profiles for ease of code deployments to higher environments.
- Installed and configured Apache airflow for workflow management and created workflows in python. Involved in converting Hive/SQL queries into Spark transformations using Spark RDD and Pyspark concepts.
- Extensive experience in Text Analytics, generating data visualizations using R, Python and creating dashboards using tools like Tableau.
- Hands on experience with big data tools like Hadoop, Spark, Hive, Pig, Impala, Pyspark, SparkSql.
- Extensive experience in Data Visualization including producing tables, graphs, listings using various procedures and tools such as Tableau.
TECHNICAL SKILLS
Big Data Ecosystem: Hadoop Map Reduce, Impala, HDFS, Hive, Pig, HBase, Flume, Storm, Sqoop, Oozie
Hadoop Distributions: Airflow, Kafka, Spark and Zookeeper Apache Hadoop 2.x/1.x, Cloudera CDP, Hortonworks HDP, Amazon EMR (EMR, EC2, EBS, RDS, S3, Athena, Glue, Elasticsearch, Lambda, DynamoDB, Redshift, ECS, Quick sight)
Programming Languages: Python, R, Scala, SAS, Java, SQL, HiveQL, PL/SQL, UNIX shell Scripting, Pig Latin
Cloud: AWS EC2, VPC, EBS, SNS, RDS, EBS, S3, Autoscaling, Lambda, Redshift, Cloud Watch, Azure Cloud, Azure Data Factory (ADF v2), Azure functions Apps, Azure Data Lake, BLOB Storage, Azure Cosmos DB, Data bricks
Databases: Snowflake, MySQL, Teradata, Oracle, MS SQL SERVER, PostgreSQL, DB2 .
NoSQL Databases: HBase, Cassandra, Mongo DB, DynamoDB and Cosmos DB
Version Control: Git, SVN, Bitbucket
ETL/BI: Informatica, SSIS, SSRS, SSAS, Tableau, Power BI, QlikView, Arcadia.
Devops Tools: Jenkins, Docker, Maven
Operating System: Mac OS, Windows 7/8/10, Unix, Linux, Ubuntu
PROFESSIONAL EXPERIENCE
Confidential, Dallas, Texas
Data Engineer
Responsibilities:
- Performed advanced data processing procedures leveraging Spark's in-memory computing capabilities.
- Developed Spark code using Scala and Spark-SQL/Streaming for faster processing of data.
- Build Data Pipeline Structure to transform Raw data to final output and store the data into Hive tables.
- Developed Spark-Streaming applications to consume the data from Kafka topics and to insert the processed streams to Hive/BQ.
- Used AWS Data Pipeline to configure data loads from S3 to Redshift, as well as JSON schema to define table and column mapping from S3 data to Redshift.
- Takes care of the day to day running of Spark, Kafka and Hive cluster Jobs
- Developed Apache spark jobs using Scala in test environment for faster data processing and used spark SQL for querying.
- Working on AWS Glue cloud, Snowflake, python and pyspark programming language.
- Worked on Airflow 1.8(Python2) and Airflow 1.9(Python3) for orchestration and familiar with building custom Airflow operators and orchestration of workflows with dependencies involving multi-clouds and Strong experience in migrating other databases to Snowflake.
- In-depth knowledge of the Snowflake database, schema, and table structure will be defined, as will virtual warehouse sizing for Snowflake for various types of workloads.
- Experience with Apache spark streaming and Batch framework. Create Spark jobs for data transformation and aggregation.
- Hands-on experience with Snowflake utilities, Snow-SQL, Snow-Pipe, Big Data model techniques using Python
- Performed some operations, visualization using libraries like Matplotlib, Pandas.
- Tableau development skills, including visualization development, the use of parameters, and report optimization.
- Monitored and tracked issues within the team using JIRA and working application in Agile methodology with SCRUM meetings.
- Responsible for using GIT for version control to commit the code developed, which was then used for deployment using the build and release tool Jenkins.
Environment: Hive, HBase, Flume, Spark, Oozie, Oracle, GitHub, Tableau, Unix, Flume, Sqoop, HDFS, Python.
Confidential, New Jersey
Data Engineer
Responsibilities:
- Strong Knowledge on Software Development Lifecycle (SDLC), Application Maintenance Change Process (AMCP) and Agile.
- Monitored Spark cluster using Log Analytics and Ambari Web UI. Transitioned log storage from MS SQL to CosmosDB and improved the query performance.
- Extensively involved in creating database objects like tables, views, stored procedures, triggers, packages, and functions using T-SQL to provide structure and maintain data efficiently.
- Highly involved in analytical platform, handled data quality, and improved the performance using Scala’s higher order functions, lambda expressions, pattern matching and collections.
- Designed and developed Automated ETL jobs in Talend and pushed the data to Azure SQL data warehouse.
- Heavily used Azure Synapse to manage processing workloads and served data for BI and prediction needs.
- Strong Involvement in developing Spark Scala scripts for mining data and performed transformations on large datasets to provide real time insights and reports.
- Experience in working with Azure cloud platform (HDInsight, Databricks, Data Lake, Blob, Data Factory, Synapse, SQL DB, SQL DWH).
- Performed data cleansing and applied transformations using Databricks and Spark data analysis.
- Designed and automated Custom-built input adapters using Spark, Sqoop and Oozie to ingest and analyse data from RDBMS to Azure Data Lake.
- Hands on building an Enterprise Data Lake using Data Factory and Blob storage, enabling other teams to work with more complex scenarios and ML solutions.
- Strong involvement in developing automated workflows for daily incremental loads, moved data from traditional RDBMS to Data Lake.
- Used Azure Data Factory, SQL API and Mongo API and integrated data from MongoDB, MS SQL, and cloud (Blob, Azure SQL DB).
- Extensively used Databricks notebooks for interactive analytics using Spark APIs.
- Managed resources and scheduling across the cluster using Azure Kubernetes Service.
- Extensive knowledge in Data transformations, Mapping, Cleansing, Monitoring, Debugging, performance tuning and troubleshooting Hadoop clusters.
Environment: Azure (HDInsight, Databricks, Data Lake, Blob Storage, Data Factory, SQL DB, SQL DWH, AKS), Scala, Python, Hadoop 2.x, Spark v2.0.2, NLP, Airflow v1.8.2, Hive v2.0.1, Sqoop v1.4.6, HBase, Oozie, Talend, CosmosDB, MS SQL, MongoDB, Ambari, PowerBI, Azure DevOps, Ranger, Git.
Confidential, Irvine, CA
Python Developer
Responsibilities:
- Communicated effectively with stakeholders to gather requirements for various projects.
- Used MySQL db package and Python-MySQL connector to write and execute several MYSQL database queries from Python.
- Increased the speed and confidence of the learning algorithm by combining cutting-edge technology, statistical methods, and Parsed data, producing concise conclusions from raw data in a clean, well structured, and easily maintainable format.
- Python was used to create clustering for customer segmentation and Created functions, triggers, views and stored procedures using My SQL.
- Worked closely with back-end developer to find ways to push the limits of existing Web technology and Involved in the JIRA review meetings.
- Implemented a job which leads an electronic medical record, extract data into Oracle Database and generate an output.
- Analyse the data and provide the insights about the customers using Tableau. • Developed entire frontend and backend modules using Python on Django Web Framework.
- Developed the presentation layer using HTML, CSS, and JavaScript.
- Scheduled Time-based Oozie workflow by developing Python scripts.
- Contributed to the development of stored procedures in Oracle and Designed and built a data management system with Oracle, optimizing database queries to improve performance. Designed, implemented and automated modelling and analysis procedures on existing and experimentally created data.
- Created dynamic linear models to perform trend analysis on customer transactional data in Python.
Environment: Python, MySQL, Oracle, Tableau, Linux & Windows, Django.
Confidential
SQL Developer
Responsibilities:
- Developed SSIS packages to Extract, Transform and Load data into the data warehouse from SQL Server in MS Excel and Flat files and Used various transforms in SSIS to load data from flat files to the SQL databases.
- Developed in SSIS to extract data from relational databases, transform it, and then load it into the data mart.
- Creating Visio diagram documents to track and implement the blueprint of the ETL dataflow process implementation from the beginning to the end.
- Created SSIS packages to mitigate into the data warehouse database from heterogeneous databases and data sources.
- Developed SSIS packages to extract, transform, and load data from disparate databases and data sources into the data warehouse.
- Created packages with different control flow options and data flow transformations such as Conditional Split, Multicast, Union all and Derived Column.
- Worked on Design and implement row level data warehouse security.
- Developed ETL packages with various data sources such as SQL Server, Flat Files, Excel source files, and XML files, and then loaded the data into destination tables using various transformations using SSIS/DTS packages.
- Worked on report design, development, debugging, and testing in SQL Server Reporting Services (SSRS).
- Developed reports in SSRS with a variety of properties such as chart controls, filters, Interactive Sorting, SQL parameters.
Environment: SQL, Data Collection, Statistical Analysis, Excel, Data Cleansing, MySQL
