We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

NC

SUMMARY

  • Data Engineer having 8+ years of experience with strong background in end - to-end enterprise data warehousing and big data projects.
  • Proficient in big data tools like Hive and Spark and relational data warehouse tool Teradata etc.
  • Excellent hands-on business requirement analysis, designing, developing, testing and maintaining the complete data management & processing systems, process documentation and ETL technical and design documents.
  • Responsible for data engineering functions including, but not limited to data extract, transformation, loading, integration in support of enterprise data infrastructures - data warehouse, operational data stores and master data management.
  • Expertise in resolving production issues, hands-on experience in handling all phases of the software development Life cycle.
  • Adept at multitasking, working independently and as part of a team as required. Very flexible at adapting to changing client needs and deadlines. Possessing strong problem solving and communication skill.
  • A solid experience and understanding of designing and operationalization of large-scale data and analytics solutions on Snowflake Data Warehouse.
  • Developing ETL pipelines in and out of data warehouse using combination of Python and Snowsql.
  • Experience in Google Cloud components, Google container builders and GCP client libraries and cloud SDK's
  • Substantial experience in Spark 3.0 integration with Kafka 2.4
  • Experience in setting up monitoring infrastructure for Hadoop cluster using Nagios and Ganglia.
  • Sustaining the BigQuery, PySpark and Hive code by fixing the bugs and providing the enhancements required by the Business User.
  • Worked with GCP suite: Google Cloud storage, Data-Proc, Data Flow, Big- Query as well as AWS core components EMR, S3, Glacier and EC2 Instance with EMR cluster.
  • Proficient in Statistical Methodologies including Hypothetical Testing, ANOVA, Time Series, Principal Component Analysis, Factor Analysis, Cluster Analysis, Discriminant Analysis.
  • Expertise in transforming business resources and requirements into manageable data formats and analytical models, designing algorithms, building models, developing data mining and reporting solutions that scale across a massive volume of structured and unstructured data.
  • Worked with various text analytics libraries like Word2Vec, GloVe, LDA and experienced with Hyper Parameter Tuning techniques like Grid Search, Random Search, model performance tuning using Ensembles and Deep Learning.
  • Skilled in System Analysis, E-R/Dimensional Data Modeling, Database Design and implementing RDBMS specific features.
  • Knowledge of working with Proof of Concepts (PoC's) and gap analysis and gathered necessary data for analysis from different sources, prepared data for data exploration using data munging and Teradata.
  • Experience in developing customized UDF's in Python to extend Hive and Pig Latin functionality.
  • Experienced in building Automation Regressing Scripts for validation of ETL process between multiple databases like Oracle, SQL Server, Hive, and Mongo DB using Python.
  • Proficiency in SQL across several dialects (we commonly write MySQL, PostgreSQL, Redshift, SQL Server, and Oracle)
  • Hands-on use of Spark and Scala API's to compare the performance of Spark with Hive and SQL, and Spark SQL to manipulate Data Frames in Scala.
  • Expertise in Python and Scala, user-defined functions (UDF) for Hive and Pig using Python.
  • Experience in developing Map Reduce Programs using Apache Hadoop for analyzing the big data as per the requirement.
  • Hands on Spark MLlib utilities such as including classification, regression, clustering, collaborative filtering, dimensionality reduction.
  • Experience in working with Flume and NiFi for loading log files into Hadoop.
  • Experience in working with NoSQL databases like HBase and Cassandra.
  • Experienced in creating shell scripts to push data loads from various sources from the edge nodes onto the HDFS.
  • Good Experience in implementing and orchestrating data pipelines using Oozie and Airflow.
  • Worked with Cloudera and Hortonworks distributions.
  • Expert in developing SSIS/DTS Packages to extract, transform and load (ETL) data into data warehouse/ data marts from heterogeneous sources.
  • Experience in Data Analysis, Data Profiling, Data Integration, Migration, Data governance and Metadata Management, Master Data Management and Configuration Management.
  • Experience in developing customized UDF's in Python to extend Hive and Pig Latin functionality.
  • Expertise in designing complex Mappings and have expertise in performance tuning and slowly changing Dimension Tables and Fact tables
  • Extensively worked with Teradata utilities Fast export, and Multi Load to export and load data to/from different source systems including flat files.
  • Good knowledge of Data Marts, OLAP, Dimensional Data Modeling with Ralph Kimball Methodology (Star Schema Modeling, Snow-Flake Modeling for FACT and Dimensions Tables) using Analysis Services.
  • Strong analytical and problem-solving skills and the ability to follow through with projects from inception to completion.

TECHNICAL SKILLS

Big Data Ecosystem: HDFS, MapReduce, HBase, Pig, Hive, Sqoop, Kafka Flume, Cassandra, Impala, Oozie, Zookeeper, MapR, Amazon Web Services (AWS), EMR

Machine Learning: Classification Algorithms Logistic Regression, Decision Tree, Random Forest, K-Nearest Neighbor (KNN), Gradient Boosting Classifier, Extreme Gradient Boosting Classifier, Support Vector Machine (SVM), Artificial Neural Networks (ANN), Naïve Bayes Classifier, Extra Trees Classifier, Stochastic Gradient Descent, etc.

Cloud Technologies: AWS, Azure, Google cloud platform (GCP)

IDE’sIntelliJ: Eclipse, Spyder, Jupyter

Ensemble and Stacking: Averaged Ensembles Weighted Averaging, Base Learning, Meta Learning, Majority Voting, Stacked Ensemble, Auto ML - Scikit-Learn, ML jar, etc.

Databases: Oracle 11g/10g/9i, MySQL, DB2, MS SQL Server, HBASE

Programming: Query Languages Java, SQL, Python Programming (Pandas, NumPy, SciPy, Scikit-Learn, Seaborn, Matplotlib, NLTK), NoSQL, PySpark, PySpark SQL, SAS, R Programming (Caret, Glmnet, XGBoost, rpart, ggplot2, sqldf), RStudio, PL/SQL, Linux shell scripts, Scala.

PROFESSIONAL EXPERIENCE

Confidential, NC

Data Engineer

Responsibilities:

  • Responsible for architecting Hadoop clusters Translation of functional and technical requirements into detailed architecture and design.
  • Comparing the results of traditional system to Hadoop environment to identify any differences and fix them by finding the route cause.
  • Storing Data Files in Google Cloud S3 Buckets daily basis. Using DataProc, Big Query to develop and maintain GCP cloud base solution.
  • Hands of experience in GCP, Big Query, GCS bucket, G - cloud function, cloud data flow, Pub/sub cloud shell, GSUTIL, BQ command line utilities, Data Proc, Stack driver
  • Create a complete processing engine, based on Hortonworks distribution, enhanced to performance.
  • Design, Develop and test ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.
  • Deployed the HBase cluster in cloud (AWS) environment with scalable nodes as per the incremental business requirement.
  • Developed multiple ETL Hive scripts for data cleansing and transformations for data.
  • Developed spark applications in python (PySpark) on distributed environment to load huge number of CSV files with different schema in to Hive ORC tables.
  • Involved in egressing the Hive Data into GCP Cloud storage in the form of DAT files using Unix Shell script.
  • Used distinctive data formats while egressing the Hive data into GCP Cloud Storage.
  • Create Views in Big Query Dataset using Cloud-Shell.
  • Worked in GCP environment for development and deployment of custom Hadoop applications
  • Used Apache Nifi for loading PDF Documents from Microsoft SharePoint to HDFS. Worked on the Publish component to read the source data, extract metadata and apply transformations to build Solr Documents, index them using SolrJ.
  • Exported data from Hive to AWS s3 bucket for further near real time analytics.
  • Ingested data in real time from Apache Kafka to Hive and HDFS.
  • Developed the Apache Storm, Kafka, and HDFS integration project to do a real-time data analysis.
  • Use of Sqoop to import and export data from RDBMS to HDFS and vice-versa.
  • Responsible for migrating the code base to Amazon EMR and evaluated Amazon eco systems components like Redshift.
  • Analysed the sql scripts and designed it by using PySpark SQL for faster performance.
  • Implemented Kerberos Security Authentication protocol for existing cluster.
  • Good experience in troubleshooting production level issues in the cluster and its functionality.
  • Complete end-to-end design and development of Apache Nifi flow, which acts as the agent between middleware team and EBI team and executes all the actions mentioned above.
  • Utilized Ansible and Chef as configuration management tools to deploy consistent infrastructure across multiple environment
  • Involved in advanced procedures like text analytics and processing using the in-memory computing capabilities like Apache Spark written in Scala.
  • Developed Ansible Plays to configure, deploy and maintain software components of the existing infrastructure
  • Regular Commissioning and Decommissioning of nodes depending upon the amount of data.

Environment: GCP,S3, Hive, Spark, Java, SQL Server, PySpark, Ansible, Hortonworks, Apache Solr, Apache Tika, Linux, Azure, Redshift, Maven, GIT, JIRA, ETL, Toad 9.6, UNIX Shell Scripting, Scala, Apache Nifi.

Confidential, Ohio, Columbus

Data Engineer

Responsibilities:

  • Involved in complete Implementation lifecycle, specialized in writing custom MapReduce, and Hive
  • Extensively used Hive/HQL or Hive queries to query or search for a string in Hive tables in HDFS
  • Continuous monitoring and managing teh Hadoop cluster using Cloudera Manager
  • Implemented Spark using Python and Spark SQL for faster processing of data
  • Used Spark for interactive queries, processing of streaming data and integration wif popular NoSQL database
  • Used teh Spark -Cassandra Connector to load data to and from Cassandra
  • Implemented test scripts to support test driven development and continuous integration.
  • Dumped teh data from HDFS to Oracle database and vice-versa using Sqoop
  • Extensively involved in Installation and configuration of Cloudera Hadoop Distribution.
  • Provided support for EBS, Trusted Advisor, Cloud Watch, Cloud Front, IAM, Security Groups, Auto-Scaling, AWS CLI and Cloud Watch Monitoring creation and update.
  • Worked wif Amazon Web Services (AWS) using EC2 for computing and S3 as storage mechanism
  • Deployed Lambda and other dependencies into AWS to automate EMR Spin for Data Lake jobs
  • Scheduled spark applications/Steps in AWS EMR cluster.
  • Extensively used event-driven and scheduled AWS Lambda functions to trigger various AWS resources.
  • Implemented advanced procedures like text analytics and processing using teh in-memory computing capabilities like Apache Spark written in Scala.
  • Developed spark applications for performing large scale transformations and denormalization of relational datasets.
  • Developed and executed a migration strategy to move Data Warehouse from SAP to AWS Redshift.
  • Loaded data into teh cluster from dynamically generated files using Flume and from relational database management systems using Sqoop.
  • Used Spark Streaming to divide streaming data into batches as an input to spark engine for batch processing.
  • Worked on analyzing Hadoop cluster and different Big Data analytic tools including Pig, hive, Environment: Hadoop, HDFS, Hive, MapReduce, Impala, Sqoop, SQL, Informatica, Python, Flume, PySpark, Yarn, Pig, Oozie, Linux, AWS, Tableau, Maven, Jenkins, Cloudera, SAS (BI & DI),PL/SQL, Autosys, Oracle, Sql Server,No Sql, TeraData.

Environment: Hadoop, Hive, MapReduce, Sqoop, Kafka, Spark, Yarn, Pig, PySpark, Cassandra, Oozie, Nifi, Solr, Shell Scripting, Hbase, Scala, Maven, Java, JUnit, agile methodologies, Horton works, Soap, Python, Teradata, MySQL, Dataproc.

Confidential, Pataskala, OH

Big Data Engineer

Responsibilities:

  • Involved in analysing business requirements and prepared detailed specifications that follow project guidelines required for project development.
  • Involved in Data Ingestion using Sqoop/Flume
  • Worked with Flume in bringing click stream data from front facing application logs.
  • Used Kafka functionalities like distribution, partition, replicated commit log service for messaging systems by maintaining feeds.
  • Involved in writing Spark applications using Scala/Java.
  • Experience in using Kafka as a messaging system to implement real-time Streaming solutions using Spark Streaming
  • Worked on Spark Data sources, Spark Data Frames, Spark SQL and Streaming using Scala.
  • Good knowledge in creating data frames using Spark SQL.
  • Migrated Hive QL queries on structured into Spark SQL to improve performance.
  • Developed pyspark/Spark SQL scripts to analyze various customer behaviors.
  • Extensively worked with Pyspark/Spark SQL for data cleansing and generating Data Frames and RDDs.
  • Developed Spark code using python for pyspark, scala and Spark-SQL for faster testing and processing of data.
  • Used Spark API over Hortonworks Hadoop YARN to perform analytics on data in Hive.
  • Implemented existing Hive script in Spark Scala for better performance.
  • Explored with Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark context, Spark-SQL, Spark Streaming, Data Frame, pair RDD’s, Spark YARN.
  • Used Python for data validation and analysis purposes.
  • Involved in creating Hive tables with Dynamic Partitions and Buckets, loading various formats of data like Avro, Parquet into the tables and analyzed data using HiveQL.
  • Created Hive, HBase tables and HBase integrated Hive tables as per the design using ORC file format and Snappy compression.
  • Performed Hive Query Optimization for better performance.
  • Built reusable Hive UDF libraries for business requirements which enabled users to use these UDFs in Hive Querying.
  • Involved in loading data into Cassandra NoSQL Database.
  • Worked hands on with ETL processes. Handled imported data from various data sources, performed transformations.

Environment: Hortonworks, Hadoop, Big Data, HDFS, MapReduce, Sqoop, Oozie, NiFi

Confidential

Hadoop Developer

Responsibilities:

  • Developed different MapReduce applications on Hadoop.
  • Mining the location of users on social media sites in semi supervised environment on Hadoop cluster using Map Reduce.
  • Implementing single source shortest path on Hadoop cluster.
  • Involved in loading and transforming large sets of Structured, Semi-Structured and Unstructured data and analyzed them by running Hive queries and Pig scripts.
  • Evaluated suitability of Hadoop and its ecosystem to the above project and implemented various proof of concept (POC) applications to eventually adopt them to benefit from the Big Data Hadoop initiative.
  • Estimated Software & Hardware requirements for the Name Node and Data Node & planning the cluster.
  • Participated in requirement gathering from the Experts and Business Partners and converting the requirements into technical specifications.
  • Extracted the needed data from the server into HDFS and Bulk Loaded the cleaned data into HBase.
  • Written the Map Reduce programs, Hive UDFs in Java where the functionality is too complex.
  • Involved in running Hadoop jobs for processing millions of records of text data.
  • Involved in loading data from LINUX file system to HDFS.
  • Prepared design documents and functional documents.
  • Based on the requirements, addition of extra nodes to the cluster to make it scalable.
  • Developed HIVE queries for the analysis, to categorize different items.
  • Assisted application teams in installing Hadoop updates, operating system, patches and version upgrades when required.
  • Designing and creating Hive external tables using shared meta-store instead of derby with partitioning, dynamic partitioning and buckets.
  • Given POC of FLUME to handle the real time log processing for attribution reports.
  • Maintained System integrity of all sub-components (primarily HDFS, MR, HBase, and Hive).

Environment: Hadoop, HDFS, MapReduce, Yarn, Hive, PIG, Oozie, Sqoop, HBase, Flume, Linux, Shell scripting, Java, Eclipse, SQL.

Confidential

Hadoop Developer

Responsibilities:

  • Requirement discussions, design the solution.
  • Estimated the Hadoop cluster requirements
  • Responsible for choosing the Hadoop components (hive, pig, map-reduce, Sqoop, flume etc)
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Hadoop cluster building and ingestion of data using Sqoop
  • Imported streaming logs to HDFS through Flume
  • Used Flume to collect, aggregate, and store the web log data from different sources like web servers, mobile and network devices and pushed to HDFS
  • Developed Use cases and Technical prototyping for implementing Hive,and Pig.
  • Worked in analyzing data using Hive, Pig and custom MapReduce programs in Java.
  • Implemented partitioning, dynamic partitions and buckets in HIVE
  • Installed and configured Hive, Sqoop, Flume, Oozie on the Hadoop cluster.
  • Involved in scheduling Oozie workflow engine to run multiple Hive and Pig jobs.
  • Tuned the Hadoop Clusters and Monitored for the memory management and for the Map Reduce jobs.
  • Responsible for Cluster maintenance, Adding and removing cluster nodes, Cluster Monitoring and Troubleshooting.
  • Developed a custom Framework capable of solving small files problem in Hadoop.
  • Deployed and administered 70 node Hadoop clusters. Administered two smaller clusters.

Environment: Map Reduce, HBase, HDFS, Hive, Pig, Java, SQL, Cloudera Manager, Sqoop, Flume, Oozie, Java (JDK 1.6), Eclipse.

We'd love your feedback!