We provide IT Staff Augmentation Services!

Sr. Data Engineer Resume

4.00/5 (Submit Your Rating)

Englewood Cliffs, NJ

SUMMARY

  • Above 8+ Years of working experience in IT field as Data Engineer with Big data/Hadoop Applications and product development.
  • Extensive working experience in Hadoop, Big Data ecosystem related experience in developing Spark applications.
  • Expertise in complete Software Development Life Cycle (SDLC) in Waterfall and Agile, Scrum models.
  • Extensive working experience in agile environment using a CI/CD model.
  • Solid understanding in Design Patterns, MVC, Python Algorithms, Python Data Structures.
  • Have also done some database work regarding AWS platform and hands on EC2, S3, Redshift, Snowflake, Lambda and Databricks.
  • Experience in using SQOOP for importing and exporting data from RDBMS to HDFS and Hive.
  • Good experience in Technical consulting and end - to-end delivery with data governance.
  • Strong experience in database migration to Snowflake.
  • Hands on experience in SQL and NOSQL database such as Snowflake, HBase, Cassandra and MongoDB.
  • Experience using design patterns including MVC, Singleton, Frontend Controller, Service Locator and Decorator.
  • Experience in implementing the various services using Micro services architecture in which the services working dependently, implemented Spring Boot Microservices to divide the application into various sub modules.
  • Developed and deployed J2EE applications on both Web and Application Servers including Apache Tomcat, Web Logic, JBoss and IBM Web Sphere.
  • Good experience with Amazon Web Services (Amazon EC2, Amazon S3, Amazon RDS, Amazon Elastic Load Balancing) using the Elastic Search APIs.
  • Experience in AWS, implementing solutions using services like (EC2, S3, RDS, Redshift, VPC)
  • Extensive experience working with structured data using Spark SQL, Data frames, Hive QL, optimizing queries, and in corporate complex UDF's in business logic.
  • Worked on SparkSQL, Spark Streaming and using Core Spark API to explore Spark features to build data pipelines.
  • Experience in installation, configuration, supporting and managing Hadoop clusters.
  • Experience in working with MapReduce programs using Apache Hadoop for working with Big Data.
  • Experience in installation, configuration, supporting and monitoring Hadoop clusters using Apache, Cloudera distributions and AWS.
  • Experience in Big Data Hadoop Ecosystem in ingestion, storage, querying, processing and analysis of Big data.
  • Experience in creating tables, dropping and altered at run time without blocking updates and queries using Spark and Hive.
  • Experienced in converting Hive/SQL queries into Spark transformations using Spark RDD and Python.
  • Good exposure to Development, Testing, Implementation, Documentation and Production support.

TECHNICAL SKILLS

Big Data Technologies: Hadoop 3.3, HDFS, MapReduce, Hive 2.3, Sqoop 1.4, Apache Impala 2.1, Oozie 4.3, Yarn, Apache Flume 1.8, Kafka 1.1, Zookeeper

Cloud Platform: Amazon AWS, EC2, EC3, MS Azure, Azure Analysis Services, HDInsight, Azure Data Lake, Data Factory.

NoSQL DB: HBase 2.4/2.3, Cassandra 3.11, Mongo DB, Couch DB, Snowflake DB

Programming Language: Scala, Python 3.6, SQL, PL/SQL, Shell Scripting

Hadoop Distributions: Cloudera, Hortonworks, MapR

SDLC Methodologies: Agile, Waterfall

Operating Systems: Windows 10, Linux and Unix

PROFESSIONAL EXPERIENCE

Confidential

Sr. Data Engineer

Responsibilities:

  • As Senior Data Engineer worked on Architecture Design for Multistate implementation or deployment.
  • Developed understanding of key business, product and user questions.
  • Followed agile methodology for the entire project.
  • Defined the business objectives comprehensively through discussions with business stakeholders, functional analysts and participating in requirement collection sessions.
  • Provided a summary of the Project's goals, and the specific expectation of business users from BI and how it aligns with the project goals.
  • Lead the estimation, review the estimates, identify the complexities and communicate to all the stakeholders.
  • Installed and configured Hadoop Ecosystem components.
  • Responsible for data governance rules and standards to maintain the consistency of the business element names in the different data layers.
  • Collaborated with Business users for requirement gathering for building Tableau reports per business needs.
  • Involved in Migration of data from Amazon Redshift data warehouse to Snowflake.
  • Involved in code migration of quality monitoring tool from AWS EC2 to AWS Lambda and built logical datasets to administer quality monitoring on snowflake warehouses.
  • Implemented One time Data Migration of Multistate level data from SQL server to Snowflake by using Python and SnowSQL.
  • Created Snowpipe for continuous data load from staged data residing on cloud gateway servers.
  • Used AWS glue catalog wif crawler to get the data from S3 and perform SQL query.
  • Worked on analyzing data using hive.
  • Developed ETL Pipelines in and out of data warehouse, develop major regulatory and reports using advanced SQL queries in snowflake.
  • Pulling the data from data lake (HDFS) and massaging the data with various RDD transformations.
  • Involved in Kafka and building use case relevant to our environment.
  • Stage the API or Kafka Data (in JSON file format) into Snowflake DB by flattening the same for different functional services.
  • Designed efficient and robust Hadoop solutions for performance improvement and end-user experiences.
  • Worked in a Hadoop ecosystem implementation/administration, installing software patches along with system upgrades and configuration.
  • Created the XML control files to upload the data into Data warehousing system.
  • Wrote Python scripts to parse XML documents and load the data in database.
  • Used Python to extract weekly information from XML files.
  • Developed Python scripts to clean the raw data.
  • Operations and JSON schema to define table and column mapping from S3 data to Redshift.
  • Used Flatten table function to produce lateral view of VARIENT, OBJECT and ARRAY column.
  • Worked with both Maximized and Auto-scale functionality while running the multi-cluster warehouses.
  • Build Docker Images to run airflow on local environment to test the Ingestion as well as ETL pipelines.
  • Maintaining Docker container clusters managed by Kubernetes.
  • Utilization of Kubernetes and Docker for the runtime environment of the CI/CD system to build, test and deploy.
  • Created Airflow DAGs to schedule the Ingestions, ETL jobs and various business reports.
  • Support Production Environment and debug issues using Splunk logs.
  • On call support for production job failures and lead the effort on working with various teams to resolve the issues.

Environment: Amazon Web Services, Elastic Map Reduce cluster, EC2s, Cloud Formation, Amazon S3, Amazon Redshift, Hive, Snowflake, Shell Scripting, Tableau and Kafka.

Confidential, Englewood Cliffs, NJ

Data Engineer

Responsibilities:

  • As a Data Engineer involved in Agile Scrum meetings to help, manage and organize a team of developers with regular code review sessions.
  • Participated in Code Reviews, Enhancement discussion, maintenance of existing pipelines & systems, testing and bug-fix activities on-going basis.
  • Worked closely with the business analysts to convert the Business Requirements into Technical Requirements and prepared low and high level documentation.
  • Worked on Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Pair RDD's
  • Developed ETL Processes in AWS Glue to migrate data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.
  • Worked on Ingesting data by going through cleansing and transformations and leveraging AWS Lambda, AWS Glue and Step Functions.
  • Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data.
  • Seamlessly worked on Python to build data pipelines after the data got loaded from Kafka.
  • Used Kafka Streams to Configure Spark Streaming to get information and then store it in HDFS.
  • Worked on loading data into Spark RDD's, perform advanced procedures like text analytics using in-memory data computation capabilities of Spark to generate the Output response.
  • Implemented usage of Amazon EMR for processing Big Data across a Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service (S3)
  • Created AWS Lambda functions and assigned IAM roles to schedule python scripts using Cloud Watch Triggers to support the infrastructure needs (SQS, Event Bridge, SNS)
  • Involved in converting MapReduce programs into Spark transformations using Spark RDD's using Scala and Python.
  • Integrated Kafka-Spark streaming for high efficiency throughput and reliability.
  • Developed a python script to hit REST API’s and extract data to AWS S3
  • Conducted ETL Data Integration, Cleansing, and Transformations using AWS glue Spark script
  • Worked on functions inLambdathat aggregates the data from incoming events, and then stored result data in AmazonDynamo DB.
  • Deployed the project on Amazon EMR with S3 connectivity for setting a backup storage.
  • Designed and Developed ETL jobs to extract data from oracle and load it in data mart in Redshift
  • Worked on AWS Data Pipeline to configure data loads from S3 to into Redshift
  • Used JSON schema to define table and column mapping from S3 data to Redshift
  • Connected Redshift to Tableau for creating dynamic dashboard for analytics team
  • Used JIRA to track issues and Change Management
  • Involved in creating Jenkins jobs for CI/CD using GIT, Maven and Bash scripting.
  • Involved in daily Scrum meetings to discuss the development/progress and was active in making scrum meetings more productive.

Environment: Spark 3.3, AWS S3, Redshift, Glue, EMR, IAM, EC2, Tableau, Jenkins, Jira, Python, Kafka, Agile

Confidential - Houston, TX

Hadoop Engineer

Responsibilities:

  • Installed and configured various components of Hadoop ecosystem and maintained their integrity on Cloudera.
  • Extensively involved in Cluster Capacity planning, Hardware planning, Installation, Performance Tuning of the Hadoop Cluster.
  • Followed agile methodology and Scrum process.
  • Actively involved on proof of concept (POC) for Hadoop cluster in AWS.
  • Used EC2 instances, EBS volumes and S3 for configuring the cluster.
  • Involved in migrating the ON Premise data to AWS.
  • Worked with big data developers, designers and scientists in troubleshooting job failures and issues.
  • Exported the analyzed data to the relational databases using Sqoop.
  • Monitored systems and services through Cloudera manager dashboard to make the clusters available for the business.
  • Involved in upgrading clusters to Cloudera Distributed versions and deployed into CDH5.
  • Extensively working on Spark using Python for testing and development environments.
  • Executed tasks for upgrading cluster on the staging platform before doing it on production cluster.
  • Changed the configurations based on the requirements of the users for the better performance of the jobs.
  • Designed, configured and managed the backup and disaster recovery for HDFS data.
  • Extensively involved in configuring the Storm in loading the data from MySQL to HBase.
  • Used Impala to write sample queries to test connectivity or work flow.
  • Monitored multiple clusters using Cloudera Manager.
  • Setup Flume for different sources to bring the log messages from outside to Hadoop HDFS.

Environment: Hadoop, AWS, EC2, S3, CDH 5, MySQL, HBase, HDFS, Impala, Flume, Sqoop, Python, Spark & Agile.

Confidential - Piscataway, NJ

Software Engineer

Responsibilities:

  • Gathered requirements and translating the business details into Technical design.
  • Collaborated within a team using an agile development workflow and widely-accepted collaboration practices using Git.
  • Implemented responsive user interface and standards throughout the development and maintenance of the website using the HTML, CSS, JavaScript, and JQuery.
  • Developed Business Logic using Python on Django Web Framework.
  • Implemented AWS high-availability using AWS Elastic Load Balancing (ELB), which performed balance across instances.
  • Developed views and templates with Python and Django to create a user-friendly website interface.
  • Development of Python APIs to dump the array structures in the Processor at the failure point for debugging.
  • Enhanced by adding Python XML SOAP request/response handlers to add accounts, modify trades and security updates.
  • Logged user stories and acceptance criteria in JIRA for features by evaluating output requirements and format
  • Written test cases using PyUnit and Selenium Automation testing for better manipulation of test scripts.
  • Utilized Kubernetes and Docker for the runtime environment for the CI/CD system to build and test and deploy.
  • Wrote modules in Python to connect to Mongo DB with PyMongo and doing CRUD operations with MongoDB.
  • Involved in development of Web Services using SOAP for sending and getting data from the external interface in the XML format.
  • Used automation Jenkins for continuous integration on Amazon EC2.
  • Used JIRA for Bug tracking and issue tracking.
  • Used Django configuration to manage URLs and application parameters.
  • Worked as part of an Agile/Scrum based development team and exposed to TDD approach in developing applications.

Environment: Python 3.7.10, MongoDB, CI/CD, PyMongo, Django, JIRA, Amazon EC2, Git, HTML 5.5, CSS3, JavaScript, and JQuery.

We'd love your feedback!