We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

Albany, NY

SUMMARY:

  • Overall, 8+ working experience as a Big Data/Python Developer in designed and developed various applications like big data, Hadoop, AWS, GCP, Python, PySpark and open - source technologies.
  • AWS Certified Solution Architect.
  • Hands-on experience wif AWS GLUE, Lambda, S3, VPC, IAM Role, Redshift, SNS, CloudWatch, EMR, Athena, RDS, Ec2, Step Function, etc.,
  • Hands-on experience wif Google Cloud Platform services like Cloud Function, Data Proc, Kubernetes, Cloud Storage, Big Query, Cloud Run, Cloud Registry, Cloud Composer, Cloud Monitor, etc.,
  • Experience in leveraging big data tools such as Spark, Hadoop, Hive, HBase, Kafka, Zookeeper, Flume, MapReduce, Oozie, Yarn and Pig.
  • Experience wif Redshift, BigQuery and Snowflake
  • Solved performance issues in Hive and Pig scripts wif understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs.
  • Experienced wif continuous integration and continuous deployment tools (CI/CD)
  • Hands on experience in Test-driven development, Software Development Life Cycle (SDLC) methodologies like Agile and Scrum.
  • Developed transformation logic using snow pipeline
  • Worked wif large scale data processing collected data from a variety of structured and unstructured data sources, stores data in a scale-out data lake and prepared teh data using ELT techniques in preparation for teh data science, data exploration and analytic modeling..
  • Procedural knowledge in data Cleaning and Analyzing using HiveQL and custom MapReduce programs.
  • Participates in teh development improvement and maintenance of snowflake database applications
  • Hands-on experience in GIT, AWS Code Build, Docker (Container Based Tool).
  • Good understanding of designing attractive data visualization dashboards using Tableau.
  • Experience in using various IDEs/Text Editors such as PyCharm, Jupyter, Nano, Emacs and repositories such as Git, SVN.
  • Developed Scala scripts, UDFs using both Data frames and RDDs in Spark for Data Aggregation, queries and writing data back into OLTP Systems.
  • Created batch data by using spark wif teh halp of Spark in developing Data Ingestion pipelines using Kafka.
  • Hands on experience in designing and developing POCs in Spark to compare teh performance of Spark wif Hive and SQL/Oracle using PySpark.
  • Strong Experience in implementing Data warehouse solutions in Confidential Redshift; Worked on various projects to migrate data from on premise databases to Confidential Redshift, RDS and S3.
  • Used Flume and Kafka to direct data from different sources to/from HDFS.
  • Worked wif AWS cloud and created EMR clusters wif spark for analyzing raw data processing and access data from S3 buckets.
  • Scripted an ETL Pipeline on Python that ingests files from AWS S3 to Redshift Table.
  • Hands on experience wif various file formats such as ORC, Avro, Parquet, and JSON.
  • Ability to develop MapReduce program using Python.
  • Good understanding and exposure to Python programming.
  • Exporting and importing data to and from Oracle using SQL developer for analysis.
  • Good experience in using Sqoop for traditional RDBMS data pulls and worked wif different distributions of Hadoop like Hortonworks and Cloudera.
  • PySpark Developer using AWS services like AWS Glue, Lambda services, Athena.
  • Involved in teh Design and building of project from teh scratch. Key contributor in building teh complete workflow, trigger, PySpark jobs, crawlers, Lambda functions.
  • Excellent analytical and problem-solving skills and ability to work on own besides being a valuable and contributing team player

TECHNICAL SKILLS:

Big Data Tools: Hadoop, Hive, Apache Spark, PySpark, HBase, Kafka, YARN, Sqoop, Impala, Oozie, Pig, Map Reduce, Zookeeper and Flume

Hadoop Distributions: EMR, Cloudera, Hortonworks.

Cloud Services: AWS, EC2, S3, EMR, RDS, Glue, Presto, Lambda, RedShift and Azure - Data Lakes, BLOB,GCP, Cloud Storage, BigQuery, Compute Engine, Cloud Composer, Data Proc, Data Flow, Pub-sub

BI and Data Visualizations: ETL -Informatica, SSIS, Talend, Tableau and Power BI

Relational Databases: Oracle, SQL Server, Teradata, MySQL, PostgreSQL and Netezza

No SQL Databases: Cassandra, MongoDB and HBase

Programming Languages: Scala, Python and R

Scripting: Python and Shell scripting

Build Tools: Apache Maven and SBT, Jenkins, Bitbucket

Version Control: GIT and SVN

Operating Systems: Unix, Linux, Mac OS, CentOS, Ubuntu and Windows

Tools: PUTTY, Putty-Gen, Eclipse, IntelliJ and Toad

PROFESSIONAL EXPERIENCE:

Confidential, Albany, NY

Data Engineer

Responsibilities:

  • Worked wif AWS cloud and created EMR clusters wif spark for analyzing raw data processing and access data from S3 buckets.
  • Involved in creating IAM Users, Groups, Roles, Identify Providers and defining Policies and applying to IAM Users and Groups. Experience creating EC2 AMI’s and migrating it to different regions. Also automate EC2 instance backup.
  • Designed and developed ETL pipelines using AWS S3, Redshift, Step Function, Lambda, Glue and EMR.
  • Worked on Docker based containers for using Airflow
  • Created Python scripts using Boto3 to automate AWS infrastructure provisioning tasks such as creating S3 buckets in AWS, create EMR cluster, add steps to running EMR cluster to exe Spark jobs
  • Integrated teh end to end data pipeline to take data from source systems to target data repositories ensuring teh quality and consistency of data is maintained at all times.
  • Designed and implemented highly performant data ingestion pipelines from multiple sources using Apache Spark
  • Configured S3 versioning and lifecycle policies to and backup files and archive files in Amazon Glacier
  • Working wifin an Agile delivery to deliver proof of concept and production implementation in iterative sprints.
  • Designed and implemented Incremental framework to handle incremental load process TEMPeffectively in Airflow.
  • Worked on various file formats such as ORC, Avro, Parquet, and JSON.
  • Developed PySpark script to setup teh data pipeline.
  • Worked on developing ETL streams using Databricks
  • Involved in teh Design and building of project right from teh scratch.
  • Designed and created automation workflows and execution
  • Developed python code for different tasks, dependencies, SLA watcher and time sensor for each job for workflow management and automation using Airflow tool
  • Worked on Developing DAG, Performance tuning of teh DAGs and task implementation.

Environment: PySpark, Python, AWS Glue, Lambda, Athena, Teradata, Snowflake, Airflow, S3, Boto3, Databricks, IAM, SNS, DAGS.

Confidential, Atlanta, GA

Data Engineer

Responsibilities:

  • Worked on multi cloud environment both AWS and GCP services.
  • Designed and developed batch and streaming pipelines using AWS and GCP services for different clients.
  • Designed and developed ETL Jobs using AWS services like S3, EMR, Glue, Lambda, Athena, Step Function and Redshift.
  • Involved in teh design and building of project right from teh scratch.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDD, Scala, and Python.
  • Worked in Agile development environment in sprint cycles of two weeks by dividing and organizing tasks. Participated in daily scrum and other design related meetings.
  • Collaborated wif product teams, data analysts and data scientists to design and built data-forward solutions.
  • Key contributor in building teh complete workflows, triggers, PySpark jobs, crawlers, Lambda functions.
  • Using AWS Glue to schedule and track teh job performance by enabling job.
  • Using AWS Cloud watch to track all log activities.
  • Hands-on experience wif Snowflake utilities, Snow SQL, Snow Pipe, Big Data model techniques using Python
  • ETL pipelines in and out of data warehouse using combination of Python and Snowflakes Snow SQL Writing SQL queries against Snowflake.
  • Using AWS Redshift as a database/data warehouse to perform all analytics.
  • Using AWS Lambda to perform rest API activities and dump data into S3.
  • Wrote various data normalization jobs for new data ingested into Redshift.
  • Designed and developed batch and streaming pipelines using GCP services like Cloud Storage, BigQuery, Compute Engine, Cloud Composer, Cloud Monitoring and cloud Function.
  • Migrated teh on-perm solutions to GCP cloud, created Data lake in cloud storage and data warehouse in BigQuery.
  • Automated and Orchestrated pipelines using Google composer, Created Airflow dags.
  • Teh Process involved Design, Development, Build, Testing, Implementation, and support till their were no further issues.

Environment: PySpark, Python, Scala, AWS Glue, Lambda, Athena, Aurora DB, Dynamo DB CloudWatch, S3, MySql, NoSql, Snowflake, RedShift, IAM, SNS.

Confidential, Charlotte, NC

Data Engineer

Responsibilities:

  • Worked as a Big Data/Hadoop Developer wif Hadoop Ecosystems components like HBase, Sqoop, Zookeeper, Oozie, Hive and Pig wif Cloudera Hadoop distribution.
  • Involved in Agile development methodology active member in Scrum meetings.
  • Worked in Azure Environment for development and deployment of Custom Hadoop Applications.
  • Created Hive schemas using performance techniques like partitioning and bucketing.
  • Developed analytical components using Kafka and Spark Stream.
  • Developed POC using Scala and deployed on teh Yarn cluster, compared teh performance of Spark, wif Hive and SQL.
  • Involved in converting Hive queries into Spark transformations using Spark RDDs, Python and Scala.
  • Performed transformations like event joins, filter boot traffic and some pre-aggregations using Pig.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
  • Used windows Azure SQL reporting services to create reports wif tables, charts, and maps.
  • Configured Oozie workflow to run multiple Hive and Pig jobs which run independently wif time and data availability.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
  • Imported and exported teh analyzed data to teh relational databases using Sqoop for visualization and to generate reports for teh BI team.
  • Worked as a Big Data/Hadoop Developer for providing solutions for big data problem.
  • Worked in Agile development environment in sprint cycles of two weeks by dividing and organizing tasks. Participated in daily scrum and other design related meetings.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDD, PySpark, and Python.
  • Developed data flow to move data from different sources to HDFS and from HDFS to S3buckets
  • Worked on Spark SQL, created Data frames by loading data from Hive tables and created prep data and stored in AWS S3.
  • Responsible for loading teh customer's data and event logs from Kafka into HBase using RESTAPI.
  • Created Partitions, Buckets based on State to further process using Bucket based Hive joins.
  • Installed and Configured Apache Hadoop clusters for application development and Hadoop tools like Hive, Pig, HBase, Zookeeper and Sqoop.
  • Wrote complex Hive queries and UDFs in Python.
  • Analyzed teh data by performing Hive queries (HiveQL), ran Pig scripts, Spark SQL and Spark streaming.
  • Developed tools using Python, Shell scripting, XML to automate some of teh menial tasks.

Environment: Hadoop, Agile, Spark, Python, HDFS, Hive, AWS, NoSQL, HBase, Kafka, EMR, MapReduce, MySQL, Hadoop, HBase, Zookeeper, Hive.

Confidential

Hadoop Developer

Responsibilities:

  • Worked as Hadoop Developer and responsible for taking care of everything related to teh clusters.
  • Responsible for building scalable distributed data solutions using Hadoop cluster environment wif Horton works distribution.
  • Developed Spark scripts by writing custom RDDs in PySpark for data transformations and actions on RDDs.
  • Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
  • Involved in performance tuning of Spark jobs using Cache and using complete advantage of cluster environment.
  • Developed Spark scripts by using Scala Shell commands as per teh requirement.
  • Developed in scheduling Oozie workflow engine to run multiple Hive and Pig jobs.
  • Worked wif different file formats such as Text, Sequence files, Avro, ORC and Parquet.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Scala.
  • Used Spark API over Hadoop Yarn as execution engine for data analytics using Hive.

Environment: Hadoop, Pig, Hive, HBase, Oozie, Sqoop, Kafka, Spark, Scala, Zookeeper, HDFS, Oozie, JSON, XML, Oracle, MySQL, Cassandra, Jenkins.

We'd love your feedback!