We provide IT Staff Augmentation Services!

Data Engineer/hadoop Developer Resume

5.00/5 (Submit Your Rating)

SUMMARY

  • Data Engineer with 7+ years of IT experience in Big data, Cloud Computing, SQL/NoSQL Databases, Software Development, DevOps, Application Support.
  • Good experience of development of cloud native application.
  • Extensive experience on Bigdata Analytics with hands on experience in writing MapReduce jobs on HadoopEcosystem including Hive and Pig.
  • Experience in implementing automated pipeline using AWS Resources like AWS Lambda, API Gateway, S3, Step Function, EMR, EC2, Glue, Athena, Redshift, Cloud Formation etc.
  • Strong experience in implementing Data warehousing applications using ETL tool, Oracle, and Unix.
  • Experience in implementing robust automated data pipeline to perform ETL on data that can produce valuable business insights and driving innovation.
  • Experience in Designing, creating, testing and maintaining the complete data management & processing systems.
  • Experienced to generate SQL and PL/SQL scripts to install building database objects including tables, views, primary keys, indexes, constraints, packages.
  • Good experience in creating procedures, packages, functions, triggers views, tables, indexes, cursors, SQL collections and optimizing query performance and other database objects using SQL and have good knowledge in writing SQL queries.
  • Enabled speedy reviews and first mover advantages by using EC2 server to automate data loading from AWS S3 bucket into the Hadoop Distributed File System and Hive, Drill and Spark to pre - process the data, Python and Bash scripting to batch process the data.
  • Strong work experience in Spark SQL, Spark Structured Streaming with PySpark and consuming payloads from Kafka.
  • Worked on various file formats like Parquet, Avro, Delta, CSV & JSON formats.
  • Worked on performance tuning of Spark Jobs.
  • A quick learner and motivated team player with excellent analytical, inter-personal and communication skills.
  • Proficient in understanding Client requirements, mapping them into the System, quickly develop the poc (proof of concept).
  • Excellent Organization, Analytical and Problem-Solving skills, and ability to quickly learn new technologies.

TECHNICAL SKILLS

Big Data Technologies: Spark, Spark Structured Streaming, Hive, Kafka, HDFS, Airflow etc.

Programming Language: Python, SQL, R

Database/DW: Hive, Snowflake, DynamoDB, Oracle, MySQL

Cloud Technology: AWS, Databricks

AWS Services: Lambda, EMR, Glue, EC2, Athena, S3, Redshift, DynamoDB, SQS, SNS, RDS, IAM, Cloud formation etc.

Utilities: MS Word, Excel, Macros, Access, Power Point

DevOps: GitHub, Terraform, Jenkins, CloudFormation etc.

PROFESSIONAL EXPERIENCE

Confidential

Data Engineer/Hadoop Developer

Responsibilities:

  • Design and development of framework to load data from various source tables and transform the data as per business requirement.
  • Implemented Spark Structured Streaming to receive data in near real time and process the data to store in Delta tables.
  • Gained good knowledge in troubleshooting and performance tuning Spark applications and Hive scripts to achieve optimal performance.
  • Optimize and fine tuning of databricks jobs and delta tables.
  • Developed the event pipeline using AWS Services like S3, Step function, Lambda, AWS Glue, Athena, Redshift etc.
  • Used Oozie and Oozie Coordinators for automating and scheduling our data pipelines.
  • Loaded the processed data into Redshift tables for allowing downstream ETL and Reporting teams to consume the processed data.
  • Worked extensively in automating creation/termination of EMR clusters as part of starting the data pipelines.
  • Orchestrated data pipelines using Airflow and to interact with Databricks jobs.
  • Moving existing ETL jobs from traditional SQL database processing to the cloud based big data processing platform, ensure that the jobs are designed to scale.
  • Provide technical support for optimize and performance tuning of ETL jobs in distributed cloud environment that result in cost saving and robust data management & processing platform.
  • Good experience working on reporting tool Tableau for creating datasources, building reports, dashboards and performing data analysis
  • Involved in continuous Integration of application using Jenkins.
  • Provided Knowledge transfer session offshore team to support project.

Confidential

Data Engineer /Hadoop Developer

Responsibilities:

  • Monitoring Sqoop jobs which is schedule on daily basis.
  • Verifying incremental loaded data.
  • Writing PySpark script to import data into landing zone from third party system.
  • Developed data pipelines using Spark, Hive and Sqoop to ingest, transform and analyze operational data.
  • Involved in Enrich Layer for data cleaning process using PySpark.
  • Redefined/Consumption layer interaction through Oracle ODI tool.
  • Worked on different file formats like Text, Avro, Parquet, JSON, XML files using Map Reduce Programs.
  • Analyzed the SQL scripts and designed the solution to implement using Spark.
  • Tuning existing Spark codes time to time based on the business requirement.
  • Configured Nifi to check for new data time to time from Oracle DB server.
  • Writing SQL queries for business for getting insight of data.
  • Exposure to Kafka cluster, involved in connecting various Kafka Topics.
  • Imported data into data lake using various file format using Sqoop jobs.
  • Spark jobs deployed to AWS EMR cluster and stored the result to Amazon S3 storage.
  • Extensively worked with Partitions, Dynamic Partitioning, bucketing tables in Hive, designed both Managed and External tables, also worked on optimization of Hive queries.
  • Involved in collecting and aggregating large amounts of log data using Flume and staging data in HDFS for further analysis.

Confidential

System Analyst

Responsibilities:

  • Analyzed data and remove unwanted data to remove errors within the applications
  • Check data entry fields to make sure valid values are entered
  • Primarily customer support role interacting with outside clients in resolving their issue
  • Worked on preparing the STM document
  • Worked on validation of data in QA and UAT environment.
  • Tested and documented the code.
  • Provided feedback to the development group to improve Agile Project management methods.

We'd love your feedback!