We provide IT Staff Augmentation Services!

Sr. Hadoop Data Engineer Resume

0/5 (Submit Your Rating)

Rock Island, IL

SUMMARY

  • 8+ years of IT experience as a Data Engineer (MS Azure/ Amazon Web Services - AWS) with emphasis on Data Analysis and Data Engineering utilizing Hadoop Framework, Spark Framework & Big Data components such as Hadoop - HDFS, MapReduce, Yarn, Pig, HIVE, Spark - Spark Streaming, Spark SQL, Spark RDD, and Sqoop, Oozie, Python, Kafka, and Machine Learning Algorithms.
  • Experienced in the end-end process from requirements gathering to implementation using software development methodologies such as Agile Software Development, Scrum, Test Driven Development (TDD), Data Pipeline Design, development, and implementation using Continuous Integration and Continuous Deployment (CI/CD).
  • BIGDATA using HADOOP framework and Analysis, Design, Development, Documentation, Deployment, and Integration of Big Data technologies as well as Java / J2EE technologies with H, AZURE, specialized in Data Ecosystem including Data Aggregation, Querying, Storage, Analysis, Developing and Implementation of data models
  • Experience in designing and building Data Management Lifecycle covering Data Ingestion, Data integration, Data consumption, Data delivery, and integration Reporting, Analytics, and System-System integration.
  • Proficient in Big Data environment and Hands-on experience in utilizing Hadoop environment components for large-scale data processing including structured and semi-structured data.
  • Strong experience with all phases including Requirement Analysis, Design, Coding, Testing, Support, and Documentation.
  • Extensive experience with Azure cloud technologies like Azure Data Lake Storage, Azure Data Factory, Azure SQL, Azure Data Warehouse, Azure Synapse Analytical, Azure Analytical Services, Azure HDInsight, and Databricks.
  • Solid Knowledge of AWS services like AWS EMR, Redshift, S3, EC2, and concepts, configuring the servers for auto-scaling and elastic load balancing.
  • Experience with monitoring the web services using Hadoop and Spark for controlling the applications and analyzing their operation and performance.
  • Experienced in Python data manipulation for loading and extraction as well as with Python libraries such as NumPy, Pandas, and SciPy for data analysis and numerical computations.
  • Good knowledge and experience with NoSQL databases like HBase, Cassandra, and MongoDB and SQL databases like Teradata, Oracle, PostgreSQL, and SQL Server.
  • Hands-on experience in developing data integration and transformation code pipelines in object-oriented scripting languages like PySpark, and Python in Spark.
  • Experience in Developing Spark applications using Spark - SQL, Pyspark, and Data Lake in Databricks for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns.
  • Experience in configuring theZookeeperto coordinate the servers in clusters and to maintain data consistency which is important for decision-making in the process.
  • Capable of understanding and knowledge of jobworkflow schedulingand locking tools/ services likeOozie, Zookeeper, Airflow, and Apache NiFi.
  • Expertise in Creating, Debugging, Scheduling, and Monitoring jobs using Airflow and Oozie.
  • Experience in Migrating SQL database to Azure Data Lake, Azure Data Lake Analytics, Azure SQL Database, Data Bricks, and Azure SQL Data warehouse, Controlling and granting database access, and Migrating On-premise databases to Azure Data Lake store using Azure Data factory.
  • Worked on text analytics, data visualizations using R Python, dashboard creation using Tableau PowerBI.
  • Experience in writingREST APIsinPythonfor large-scale applications.
  • Experience with creating and using CI/CD (Continuous Integration and Continuous Delivery) tools such as Jenkins, and Azure DevOps pipelines.
  • Good understanding of Data Modeling (Dimensional and Relational) concepts like Star-Schema Modeling, Snowflake Schema Modeling, and Fact and Dimension Tables.
  • Experience in design, development, and implementation of database systems using MS-SQL Server for both OLTP & Data Warehousing Systems applications.
  • Strong Experience in working with Databases like Teradata and proficiency in writing complex SQL, PL/SQL for creating tables, views, indexes, stored procedures, and functions.
  • Knowledge and experience with Continuous Integration and Continuous Deployment using containerization technologies like Docker and Jenkins.

TECHNICAL SKILLS

SDLC: Agile, Scrum, Waterfall, Kanban

Big Data Ecosystem: Hadoop, MapReduce, Pig, Hive, HBase, YARN, Kafka, Flume, Sqoop, Impala, Oozie, Zookeeper, Spark, Ambari, Elastic Search, Parquet, Snappy, Airflow, NiFi

Hadoop Distributions: Cloudera (CDH3, CDH4, and CDH5), Hortonworks, MapReduce, Apache EMR

Cloud Platforms: Amazon Web Services (AWS), MS Azure

MS Azure Services: Azure SQL Database, Azure Data Lake (ADL), Azure Data Factory (ADF), Azure SQL Data Warehouse, Azure Service Bus, Azure Key Vault, Azure Analysis Service (AAS), Azure Blob Storage, Azure Search, Azure App Service, and Azure Data Platform Services

AWS Services: EMR, S3, EC2, VPC, Redshift, EMR, Lambda, Dynamo DB, RDS, SNS, SQS, Glue

ETL/ BI Tools: Informatica, SSIS, Tableau, PowerBI, SSRS

CI/ CD: Azure DevOps, Jenkins, Ant, Maven

Ticketing Tools: JIRA

Operating Systems: Linux, Windows, Ubuntu, Unix

Databases (RDBMS/ NoSQL): Oracle, SQL Server, Cassandra, Teradata, PostgreSQL, HBase, MongoDB

Programming Languages/Scripting: Scala, SQL, PL/SQL, R, Python (Pandas, NumPy, SciPy, Scikit-Learn, Seaborn, Matplotlib, NLTK), Shell Scripting.

DWH Schemas: Star Schema, Snowflake Schema

DataModeling Tools: Erwin, MS VISO

Web/ Application Server: Apache Tomcat, WebLogic, WebSphere

Version Control: Git, Subversion, GitHub

CI/CD Tools: Jenkins, Azure DevOps

PROFESSIONAL EXPERIENCE

Confidential, Rock Island, IL

Sr. Hadoop Data Engineer

Responsibilities:

  • Worked on both batch processing and streaming data Sources. Used Spark Streaming and Kafka for the streaming data processing.
  • Developed Spark Streaming script which consumes topics from distributed messaging source Kafka and periodically pushes a batch of data to spark for real-time processing.
  • Builtdatapipelines for reporting, alerting, anddatamining. Experienced with table design anddata management using HDFS, Hive, Impala, Sqoop, MySQL, and Kafka.
  • Worked on Apache Nifi to automate the data movement between RDBMS and HDFS.
  • Created shell scripts to handle various jobs like Map Reduce, Hive, Pig, Spark, etc, based on the requirement.
  • Used Hive techniques like Bucketing, Partitioning to create the tables.
  • Developing ETL pipelines in and out of data warehouses using a combination of Python and Snowflakes SnowSQL Writing SQL queries against Snowflake.
  • Worked on AWS to aggregate clean files in Amazon S3 and also on Amazon EC2 Clusters to deploy files into Buckets.
  • Performed end- to-end Architecture & implementation assessment of various AWS services like Amazon EMR, Redshift, S3.
  • Implemented Lambda to configure Dynamo DB Autoscaling feature and implemented Data Access Layer to access AWS DynamoDB data.
  • Designed and architected solutions to load multipart files which can't rely on a scheduled run and must be event-driven, leveraging AWS SNS.
  • Involvedin Data Modeling usingStar Schema, and Snowflake Schema.
  • Used Spark SQL for Scala, Python interface that automatically converts RDD case classes to schema RDD.
  • Import the data from different sources like HDFS/HBase into Spark RDD and perform computations using PySpark to generate the output response.
  • Creating Lambda functions with Boto3 to deregister unused AMIs in all application regions to reduce the cost for EC2 resources
  • Used AWS EMR clusters for creating Hadoop and spark clusters. These clusters are used for submitting and executing Scala and Python applications in production.
  • Responsible for developing a data pipeline with Amazon AWS to extract the data from weblogs and store it in HDFS.
  • Migrated the data from AWS S3 to HDFS using Kafka.
  • Integrating Kubernetes with the network, storage of security to provide comprehensive infrastructure, and orchestrating the Kubernetes containers across multiple hosts.
  • Implementing Jenkins and building pipelines to drive all microservices builds out to Docker registry and deploying to Kubernetes.
  • Worked with NoSQL databases like HBase, and Cassandra to retrieve and load the data for real-time processing using Rest API.
  • Worked on creating data models for Cassandra from the existing Oracle data model.
  • Responsible for transforming and loading large sets of structured, semi-structured, and unstructured data.

Environment: Hadoop 2.7.7, HDFS 2.7.7, Apache Hive 2.3, Apache Kafka 0.8.2.X, Apache Spark 2.3, Spark-SQL, Spark-Streaming, Zookeeper, Pig, Oozie, Java 8, Python3, S3, EMR, EC2, Redshift, Cassandra, Nifi, Talend, HBase, Cloudera (CHD 5 .X), snowflake, Power BI, Tableau

Confidential, Grants Pass, OR

Data Engineer

Responsibilities:

  • Provided 24X7 (including weekend) support to address critical failures.
  • Worked with Azure Cloud Services (PaaS & IaaS), Azure Synapse Analytics, SQL Azure, Data Factory, Azure Analysis Services, Application Insights, Azure Monitoring, Key Vault, and Azure Data Lake.
  • Worked on creating tabular models on Azure analysis services for meeting business reporting requirements.
  • Extract, transform, and load data from source systems to Azure Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL, and U-SQL of Azure Data Lake Analytics.
  • Data ingestion to one or more Azure services (Azure Data Lake, Azure Storage, Azure SQL DB, Azure SQL DW), and processing of the data in Azure Databricks.
  • Used Spark for Parallel data processing and better performances using Python.
  • Used Spark Streaming to divide streaming data into batches as an input to the Spark engine for batch processing.
  • Implement ad-hoc analysis solutions using Azure Data Lake Analytics/ Store, HDInsight
  • Completed online data transfer from AWS S3 to Azure Blob by using Azure Data Factory (ADF).
  • Used Azure Migrate to get started migrating AWS EC2 instances over to MS Azure.
  • Developed Spark Applications by using Scala, python, and Implemented Apache Spark data processing project to handle data from various RDBMS and Streaming sources.
  • Managed pipelines by using Airflow to schedule and orchestrate jobs.
  • Developed Spark Programs using Scala and performed transformations and actions on RDDs.
  • Involved in converting Hive/ SQL queries into Spark transformations using Spark RDD, Scala, and Python.
  • Developed analytical components using Scala, Spark, and Spark Stream.
  • Used Visualization tools such as Power View for Excel, and Power BI for visualizing and generating reports.
  • Perform validation and verify software at all testing phases which include Functional Testing, System Integration Testing, End to End Testing, Regression Testing, Sanity Testing, User Acceptance Testing, Smoke Testing, Disaster Recovery Testing, Production Acceptance Testing, and Pre-prod Testing phases.
  • Built CI/CD pipelines using Azure DevOps.
  • Generating various capacity planning reports (graphical) using Python packages like Numpy, Scipy, Pandas, and Matplotlib.
  • Worked with Snowflake utilities, SnowSQL, SnowPipe, and Big Data model techniques using Python.
  • Created ETL pipelines in and out of the data warehouse using a combination of Python and Snowflakes SnowSQL Writing SQL queries against Snowflake.

Environment: SDLC, R, Python, NumPy, Scipy, Matplotlib, Pandas, MS Azure, SQL Azure, ADF, Azure Databricks, Data Lake, Azure Blob, HDInsight, USQL, Big Data, Snowflake, SnowSQL, SnowPipe, Spark, Spark Streaming, Spark RDD, PySpark, Spark SQL, Scala, Hive, Oozie, HBase, Pig, Sqoop, Azure DevOps, Linux, Power BI, Informatica, SQL Server, MongoDB.

Confidential, Charlotte, NC

Data Engineer

Responsibilities:

  • Collaborate with cross-functional teams (business stakeholders, developers, product managers, product owners, and management) to identify requirements, data, and key insights needed for developing models to derive business value.
  • Worked on analyzing Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, and Sqoop.
  • Define and extract data from multiple sources, integrate disparate data into a common data model, and integrate data into a target database, application, or file using efficient programming processes
  • Worked on predictive and what-if analysis using Python from HDFS and successfully loaded files to HDFS from Teradata, and loaded from HDFS to HIVE and worked on NoSQL databases including HBase, Mongo DB, and Cassandra.
  • Developed MapReduce applications using Hadoop MapReduce programming framework for processing and used compression techniques to optimize MapReduce Jobs.
  • Developed work flow using Oozie for running MapReduce jobs and Hive Queries.
  • Worked on distributed computing architectures such as Hadoop, Kubernetes, and Docker containers.
  • Extensively used Pig for data cleansing and extracting the data from the web server output files to load into HDFS.
  • Developed a data pipeline using Kafka and Storm to store data into HDFS.
  • Worked on AWS Data Pipeline to configure data loads from S3 to into Redshift and have used AWS components (Amazon Web Services) - Downloading and uploading data files (withETL) to AWS system using S3 components and Used AWS Data Pipeline to schedule an Amazon EMR cluster to clean and process web server logs stored in Amazon S3 bucket.
  • Develop and implement scripts for database and data process maintenance, monitoring, and performance tuning.
  • Implemented Big Data Analytics and Advanced Data Science techniques to identify trends, patterns, and discrepancies on petabytes of data by using Hive, Hadoop, Python, HDFS, MapReduce, and Machine Learning.
  • Imported data from relational data sources to HDFS and imported bulk data into HBase using Map Reduce programs.
  • Design and develop scalable, efficient data pipeline processes to handle data ingestion, cleansing, transformation, and integration using Sqoop, Hive, Python, and Impala.
  • Developed Pig UDFs to know customer behavior and Pig Latin scripts for processing the data in Hadoop.
  • Developed simple to complex MapReduce streaming jobs using Python.
  • Scheduled automated tasks with Oozie for loading data into HDFS through Sqoop and pre-processing the data with Pig and Hive.
  • Developed various Python scripts to find vulnerabilities with SQL Queries by doing SQL injection, permission checks, and analysis.
  • Develop the Oozie actions like hive, shell, and java to submit and schedule applications to run in the Hadoop cluster.
  • Migrated ETL jobs to Pig Scripts and worked on different file formats like sequence files, XML, JSON, etc.
  • Created Reference Table using Informatica Analyst tool as well as Informatica Developer tool.
  • Capture data integration metadata, lineage, and catalog through configuration and parameterization.
  • Troubleshoot data processing issues and support the management of data migration.
  • Worked with building data warehouse structures, and creating facts, dimensions, and aggregate tables, by dimensional modeling Star and Snowflake schemas.
  • Worked on Oozie workflow engine for job scheduling,
  • Implement a Continuous Delivery Pipeline with Docker and Git Hub.
  • Involved in creating/modifying worksheets and data visualization dashboards in Tableau.

Environment: Agile, Scrum, Hadoop, MapReduce, HDFS, HBase, AWS, Redshift, S3, Kubernetes, Docker, UNIX, Hive, Sqoop, Oozie, BigData ECO systems, PIG, Cloudera, Python, Impala, Teradata, MongoDB, Cassandra, Informatica, Unix scripts, XML files, JSON, Rest API, Maven, GitHub, Tableau.

Sonata Software

Data Analytics Engineer

Responsibilities:

  • Analyzed the pertinent client data and worked with the Analytics team to inspect, cleanse, transform and model data and provide valuable feedback to the client.
  • Performed Data Analysis using visualization tools such as Tableau, Spotfire, and SharePoint to provide insights into the data.
  • Implemented Multidimensional and Tabular cubes to perform interactive analysis.
  • Configured Azure platform offerings for web applications, business intelligence using Power BI, Azure Data Factory etc.
  • Conceptualized and Designed Extraction, Transformation and Loading of client data from multiple sources into SQL Server Integration Services (SSIS).
  • Created DDL scripts for implementing Data Modeling changes. Created ERWIN crystal reports in HTML, RTF format depending upon the requirement,
  • Published Data model in model mart, created Skilled in System Analysis, E-R/DimensionalData Modeling, Database Design and implementing RDBMS specific features.
  • Created the conceptual model for the data warehouse using Erwin data modeling tool. Part of team conducting logical data analysis and data modeling JAD sessions, communicated data-related standards implementing Data modeling, Erwin, Dimensional Modeling, Ralph Kimball Approach, Star/Snowflake Modeling, Datamarts, OLAP, FACT & Dimensions tables, Physical & Logical data modeling and Oracle Design.
  • Responsible for development of workflow analysis, requirement gathering, data governance, data management and data loading.
  • Have setup data governance touch points with key teams to ensure data issues were addressed promptly.
  • Responsible for facilitating UAT (User Acceptance Testing), PPV (Post Production Validation) and maintaining Metadata and Data dictionary.
  • Responsible for source data cleansing, analysis and reporting using pivot tables, formulas (v-lookup and others), data validation, conditional formatting, and graph and chart manipulation in Excel.
  • Actively involved in data modeling for the QRM Mortgage Application migration to Teradata and developed the dimensional model.
  • Created views for reporting purpose which involves complex SQL queries with sub-queries, inline views, multi table joins, with clause and outer joins as per the functional needs in the Business Requirements Document (BRD).
Environment: Agile, Teradata, BTEQ, Fast Load, Fast Export, Multi load, Oracle, Unix Shell Scripts, SQL Server 2005/08,SAS, PROC SQL, MS Office Tools, MS Project, MS Access, Pivot Tables

We'd love your feedback!