We provide IT Staff Augmentation Services!

Data Engineer/azure Resume

3.00/5 (Submit Your Rating)

Charlotte, NC

SUMMARY

  • Expertise in all components of the Hadoop Ecosystem - Spark, Hive, Pig, Sqoop, HBase, Oozie, Impala, Hue
  • Almost 10 years of professional IT experience and over 8 Years of Big Data Ecosystem experience in ingestion, storage, querying, processing and analysis of big data in diverse industries.
  • Experience in working with ETL pipeline jobs with a large set of data using hive, spark with Scala and PySpark.
  • Acute knowledge on Spark architecture and real-time streaming using Spark
  • Worked on real-time data integration using Kafka, Spark streaming and HBase.
  • Hands-on experience with Spark Core, Spark SQL, and Data Frames/Data Sets/RDD API.
  • Good knowledge on Amazon Web Services (AWS) cloud services like EC2, S3, and EMR.
  • Executing Apache Hadoop on CDH and Map-R distros dubbed Elastic MapReduce(EMR) on (EC2)
  • Exporting and importing data into S3
  • Experience in implementing end-to-end full life cycle (SDLC).
  • Hands-on experience with NoSQL Databases like HBase for performing analytical operations.
  • Good knowledge on DataStage and GCP ecosystems big query, big table.
  • Worked with the SCRUM team in delivering agreed user stories on time for every sprint.
  • Experience with configuration of Hadoop Ecosystem components: Hadoop HDFS, Mapreduce, Hive, Impala, Kafka, Storm, Pig, Scoop, Oozie, HBase, Zookeeper and Flume.
  • Experience in building and managing scalable Hadoop clusters including Cluster designing, provisioning, custom configurations, monitoring and maintaining using different Hadoop distributions: Hortonworks, Cloudera CDH, and Apache Hadoop.
  • Experienced in using Apache Storm (Topology, Data model, Spouts and Bolts, Grouping and parallelism).
  • Experienced in handling databases: Netezza, Oracle and Teradata, Postgres DB.
  • Good Knowledge with NoSQL Databases like HBase, Cassandra, CouchDB, and MongoDB.

TECHNICAL SKILLS

Big Data Ecosystem: HDFS, MapReduce, Pig 0.12, Hive, Oozie 1.4.7, Kafka 2.10, Flume, Sqoop, Impala, Spark 2.x/1.x, Talend and HBase

Cloud Ecosystem: Amazon Web services (EC2, EMR, and S3). Hadoop Distributions Cloudera CDH 6.1/5.12/5., Hortonworks Languages Java, Scala, SQL, Shell scripting, and Python

Databases: MSSQL, Teradata, Netezza, Postgres SQL, Maria DB and Oracle

Operating systems: UNIX, Linux, and Windows Variants.

Tools: Maven, SBT, Jenkins, IntelliJ, Eclipse, GIT and SVN, Laravel 5.1.

System Software: WordPress, Local List, Google analytics, Tableau

PROFESSIONAL EXPERIENCE

Confidential, Charlotte, NC

Data Engineer/Azure

Responsibilities:

  • Data profiling and generating reports for the missing data and inconsistent data.
  • Building data pipelines to build game analytics on top of the raw data which will help game developers and marketing team
  • Developing data models for the data marts in Erwin tool
  • Develop new and existing modules in Scala while working with developers across the globe
  • Involved in converting HiveQL/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
  • Develop Scala and Python software in an agile environment using continuous integration
  • Running Apache Hadoop and dubbed Elastic MapReduce (EMR) on (EC2).
  • Understanding and exposure to troubleshooting, configuring Azure VM’s and other services like Azure Data factory, Azure Data Lake
  • Expertise in building Azure native enterprise applications and migrating applications from on-premises to Azure environments
  • Worked on migrating MapReduce programs into Spark transformations using Spark and Scala.
  • Analyzed the SQL scripts and designed the solution to implement using Spark
  • Involved in creating Hive tables, and loading and analyzing data using hive queries
  • Implemented schema extraction for Parquet and Avro file Formats in Hive.
  • Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
  • Building spark pipeline for real time analysis of active users across different country
  • Working with Data science team to build data sets and pipe line for their models

Confidential, Buffalo, NY

Data Engineer

Responsibilities:

  • Developed a Java Spring Boot application to ingest AWS S3 objects to Kafka and implemented a Spark Scala consumer to read data from Kafka.
  • Implemented a Spark ETL pipeline to prepare Analytical datasets from the Ingested raw data.
  • Wrote Hive, Spark-Scala scripts that were used as a part of DI/DQ checks and other post validation processes.
  • Worked extensively on performance tuning of Spark jobs and Hive queries.
  • Created an SFTP ingestion framework using Python Paramiko library to ingest data from multiple vendors into HDFS
  • Worked on Pyspark - DataFrame, Datasets and RDD’s
  • Responsible for developing data pipelines with Snowflake to extract the data from weblogs and store in HDFS.
  • Building Data pipeline using Airflow and SnowSQL to compute business KPI’s on daily basis
  • Formulated procedures to integration of R programming plans with data sources and delivery systems
  • Performed statistical analyses in R programming/R analysis.
  • Developed scripts to automate routine pipeline running tasks using Shell Scripts/Python.
  • Developed Coordinator and Oozie workflows to automate the jobs

Confidential, Raleigh, NC

Sr. Data Engineer / Spark Developer

Responsibilities:

  • Ingested variety of source systems data into Hadoop DataLake from RDBMS and SFTP servers.
  • Created Hive tables on top of the ingested data sets and performed standard cleansing operations using pig Latin
  • Teamed up with Architects to design Spark model for the existing Map Reduce model
  • Written all ETL transformation pipelines using PySpark and loaded the exposed the final transformed data sets from Hive Foundation layer tables.
  • Helped the Business Insights team with statistical predictions, business intelligence and data science efforts
  • Developed scripts, UDFs using both Spark SQL and Spark-Scala for aggregative operations Experienced in implementing Spark RDD/Data Frame transformations, actions to implement business analysis and Worked with Spark accumulators and broadcast variables
  • Cloudera Manager was used to monitor the health of Jobs which are running on the cluster
  • Worked closely with the App Support team in production deployment and to schedule jobs in TIDAL/CTL-M.

Confidential, Commerce, TX

Associate Data Engineer

Responsibilities:

  • Maintained and added new functionalities to the department website built on HTML, CSS and JavaScript.
  • Supported department faculty with administrative assistance as well as curriculum and research support.
  • Involved in installing Hadoop Ecosystem Components. Performed performance tuning of ETL jobs.
  • Developed data ingest workflows with stream processing systems like spark-streaming, Kafka streams.
  • Checked incoming data for accuracy and integrity to identify and resolve problems.
  • Handled microservices architecture components, including docker and Kubernetes

Confidential

Hadoop Developer

Responsibilities:

  • Worked with Business Analyst and helped representing the business domain details, prepared documentation
  • Optimized long running Hive ETL jobs by rewriting Hive SQLs, restructuring the Hive tables and changed the underlying file format structure.
  • Migrated existing analytical business models written in Hive to PySpark and developed new models directly in PySpark.
  • Output produced by these models is used to generate reports in Tableau.
  • Migrated existing Data Pipelines to cloud environment and developed new applications by choosing the right AWS components.
  • Sourced thousands of S3 JSON Objects into Spark SQL and created Hive external tables.
  • Developed Airflow Hook and Operator from the scratch to connect with Google Ad Manager API that pulls Advertising data from the API
  • Used Bitbucket and Jenkins for CICD process and exported data to Snowflake for Tableau Dashboards.

We'd love your feedback!