Senior Big Data Engineer/scrum Master Resume
Dearborn, MI
SUMMARY
- Over 9+ years of IT experience in a variety of industries working on BigDatatechnology using technologies such as Cloudera and Hortonworks distributions. Hadoop working environment includes Hadoop, Spark, MapReduce, Kafka, Hive, Ambari, Sqoop, HBase, and Impala.
- Strong experience in Software Development Life Cycle (SDLC) including Requirements Analysis, Design Specification and Testing as per Cycle in both Waterfall and Agile methodologies.
- Enabled improvement in team delivery commitments and capacity planning for sprints by identifying and tracking hidden tasks that increased customer satisfaction.
- Lead daily stand - ups and scrum ceremonies for two scrum teams. Work with product owners to groom the backlog and plan sprints. Track, escalate and remove impediments. Report at daily Scrum of Scrum meetings. Track burn down, issues and progress in Version One. Work with component teams to resolve issues.
- Facilitated Agile Adoption Retrospective for the organization with the leadership and guided teams with outcome resulting in enhanced performance.
- Good working experience on Spark (spark streaming, spark SQL) with Scala and Kafka. Worked on reading multiple data formats on HDFS using Scala.
- Experience in Microsoft Azure/Cloud Services like SQL Data Warehouse, Azure SQL Server, Azure Databricks, Azure Data Lake, Azure Blob Storage, Azure Data Factory.
- Extensive knowledge in writing Hadoop jobs for data analysis as per the business requirements using Hive and worked on HiveQL queries for required data extraction, join operations, writing custom UDF's as required and having good experience in optimizing Hive Queries.
- Experience in importing and exporting the data using Sqoop from HDFS to Relational Database systems and vice-versa and load into Hive tables, which are partitioned.
- Experience in cloud data migration usingAWSandSnowflake.
- Developed custom Kafka producer and consumer for different publishing and subscribing to Kafka topics.
- Good understanding of distributed systems, HDFS architecture, Internal working details of MapReduce and Spark processing frameworks.
- Excellent programming skills with experience in Java, C, SQL and Python Programming.
- Experience in data modeling, data warehouse, and ETL design development using Ralph Kimball model with Star/Snowflake Model Designs with analysis - definition, database design, testing, and implementation process.
- Involved in converting Hive/SQL queries into Spark transformations using Spark Data frames and Scala.
- Experience in using Kafka and Kafka brokers to initiate spark context and processing livestreaming.
- Hands on experience in writingMap Reduceprograms using Java to handle different data sets usingMap and Reduce tasks.
- Working Experience with Amazon Web Services(AWS) Cloud Platform which includes services likeEC2,S3,VPC,ELB, IAM, DynamoDB, Cloud Front, Cloud Watch, Route 53, Elastic Beanstalk (EBS), Auto Scaling, Security Groups.
- Good understanding and knowledge of NoSQL databases like MongoDB, PostgreSQL, HBase and Cassandra.
- Strong experience in core Java,Scala, SQL, PL/SQL and Restful web services.
- Experience in developingcustomUDFsfor Pig and Hive to in corporate methods and functionality of Python/Java intoPig LatinandHQL(HiveQL) and Used UDFs from Piggybank UDF Repository.
- Worked on various programming languages using IDEs like Eclipse, NetBeans, and Intellij, Putty, GIT.
- Experience on ETL concepts using Informatica Power Center, AB Initio.
- Experience in extracting files from MongoDB through Sqoop and placed in HDFS and processed.
- Excellent understanding and knowledge of job workflow scheduling and locking tools/services like Oozie and Zookeeper.
- Good experience in database design, creating Tables, Views, Stored Procedures, Functions, Triggers and Indexes.
- Hands on experience with data ingestion tools Kafka, Flume and workflow management tools Oozie.
- Used Spark Data Frames API over Cloudera platform to perform analytics on Hive data and Used Spark Data Frame Operations to perform required Validations in the data.
- Having extensive knowledge on RDBMS such as Oracle, DevOps, Microsoft SQL Server, MYSQL
- Extensive experience working on various databases and database script development using SQL and PL/SQL.
- Expert in developing SSIS/DTS Packages to extract, transform and load (ETL) data into data warehouse/ data marts from heterogeneous sources.
- Working Knowledge of ETL methods for data extraction, transformation and loading in corporate-wide ETL Solutions and Data Warehouse tools for reporting and data analysis.
- Hands on learning with different ETL tools to get data in shape where it could be connected to Tableau through Tableau Data Extract.
TECHNICAL SKILLS
Big Data Tools: Kafka, Cassandra, Apache Spark, Spark Streaming, HBase, Impala, HDFS, MapReduce, Hive, Pig, Sqoop, Flume, Oozie, Zookeeper
Hadoop Distribution: Cloudera CDH, Apache, AWS, Horton Works HDP
Programming Languages: SQL, PL/SQL, Python, UNIX, Pyspark, Pig, HiveQL, Scala, Shell Scripting
Spark Components: RDD, Spark SQL, Spark Streaming
Data Modeling Tools: Erwin Data Modeler, ER Studio v17
Methodologies: RAD, JAD, System Development Life Cycle (SDLC), Agile
Cloud Management: MS Azure, Amazon Web Services (AWS)- EC2, EMR, S3, Redshift, EMR, Lambda, Athena
Databases: Oracle 12c/11g/ 10g, MySql, MS Sql, DB2, Snowflake
No Sql Databases: MongoDB, Hbase, Cassandra
OLAP Tools: Tableau, SSAS, Business Objects, and Crystal Reports 9
ETL/Data warehouse Tools: Informatica, and Tableau.
Version Control: CVS, SVN, Clear Case, Git
Operating System: Windows, Unix, Sun Solaris
PROFESSIONAL EXPERIENCE
Confidential, Dearborn, MI
Senior Big Data Engineer/Scrum Master
Responsibilities:
- Facilitated Scrum ceremonies such as daily stand-ups, backlog grooming, sprint planning, retrospectives, etc.
- Increased productivity by helping create business process management software function.
- Facilitated several Joint Application Development (JAD) workshops to drive out business requirements. Used rapid prototyping techniques to get design input from all parties, resulting in zero follow on work after the release to production.
- Lead Scrum/SAFe events, acting as Scrum Master and Agile Coach for multiple teams. Facilitating team level product increment planning, Sprint planning, goal setting and commitment to deliverables. Ensuring that all software development lifecycle (SDLC) practices are understood and practiced on my teams
- Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
- Extensive usage of Spark for data streaming and data transformation for real time analytics.
- Experienced in Importing and exporting data into HDFS and Hive usingSqoop.
- Provided daily support to scrum teams and Clearing roadblocks and insulating scrum teams from interruptions.
- Promoting positive team dynamics utilizing good communication skills. Tracking sprint goals and metrics. Facilitating and tracking continuous improvement goals focused on creating self-organized and high-performing teams
- Implemented large Lambda architectures using Azure Data platform capabilities like Azure Data Lake, Azure Data Factory, Azure Data Catalog, HDInsight, Azure SQL Server, Azure ML and Power BI.
- Extract Real time feed usingKafkaandSpark Streamingand convert it to RDD and process data in the form of Data Frame and save the data as Parquet format in HDFS.
- Responsible for analyzing large data sets and derive customer usage patterns by developing new MapReduce programs using Java
- Experienced in writing real-time processing and core jobs usingSpark StreamingwithKafkaas a data pipe-line system.
- Worked oncomplex SNOW SQL and Python Queries in Snowflake.
- Extensively using Web HDFS REST API commands in Perl scripting.
- Involved in converting Map Reduce programs into Spark transformations using Spark RDD's using Scala and Python
- Design and Develop Data Collectors and Parsers by using Perl or Python.
- Data Import and Export from various sources through Script and Sqoop.
- Experience in developing customized UDF’s in Python to extend Hive and Pig Latin functionality.
- Implemented Copy activity, Custom Azure Data Factory Pipeline Activities
- Primarily involved in Data Migration using SQL, SQL Azure, Azure Storage, and Azure Data Factory, SSIS, PowerShell.
- Recreated existing SQL Server objects in snowflake.
- Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in Azure Databricks.
- ConfiguredSpark streamingto get ongoing information from theKafkaand store the stream information to HDFS.
- Developed JSON Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the SQL Activity.
- Worked on Snowflake Schemas and Data Warehousing
- Enabled Python scripts to explode, parse and de-dupe JSON from Kafka and land in HDFS.
- Built Kafka monitoring scripts to monitor Kafka loads into Hadoop cluster.
- Experienced in running query-usingImpalaand used BI tools to run ad-hoc queries directly on Hadoop.
- IntegratedCassandraas a distributed persistent metadata store to provide metadata resolution for network entities on the network
- WroteMap Reducejobs using Java API and Pig Latin.
- Used Azure Databricks for fast, easy and collaborative spark-based platform on Azure.
- Used Databricks to integrate easily with the whole Microsoft stack.
- Used Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scala and NoSql databases such as HBase and Cassandra.
Environment: Hadoop, Cloudera, Map Reduce, Kafka, Impala, Spark, Snowflake, Azure, data bricks, data factory, data lake, Zeppelin, Hue, Impala, Pig, Hive, Sqoop, Java, Scala, Cassandra, SQL, Tableau, Pig, Zookeeper, Teradata, Zoom-Data, Linux Red-Hat and Oracle 12c.
Confidential, Columbia, SC
Big Data Engineer
Responsibilities:
- Used Agile methodology in developing the application, which included iterative application development, weekly Sprints, stand up meetings and customer reporting backlogs.
- Design the incremental, historical extract logic to load the data from flat files into Massive Event Logging Database (MELD) from various servers.
- Created and managed cloud VMs with AWS EC2 Command line clients and AWS management console.
- Migrated on premise database structure to Confidential Redshift data warehouse. Worked on AWS Data Pipeline to configure data loads from S3 into Redshift
- Involved in data migration to snowflake using AWS S3 buckets.
- Map Reduce jobs in Python for data cleaning and data processing.
- Extracting batch and Real time data from DB2, Oracle, Sql server, Teradata, Netezza to Hadoop (HDFS) using Teradata TPT, Sqoop, Apache Kafka, Apache Storm.
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, Spark and loaded data into HDFS.
- Monitoring resources and Applications using AWS Cloud Watch, including creating alarms to monitor metrics such as EBS, EC2, ELB, RDS, S3, SNS and configured notifications for the alarms generated based on events defined.
- Used PySpark to expose Spark API to Python.
- UsedAWS S3 Bucketsto store the file and injected the files into Snowflake tables usingSnow Pipeand run deltas usingData pipelines.
- Developed Talend jobs to populate the claims data to data warehouse - star schema, snowflake schema, Hybrid Schema
- Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs.
- Built different visualizations and reports in tableau using Snowflake data.
- Loading, analyzing and extracting data to and from Elastic Search with Python.
- Understanding of AWS Product and Service suite primarily EC2, S3, VPC, Lambda, Redshift, Spectrum, Athena, EMR(Hadoop) and other monitoring service of products and their applicable use cases, best practices and implementation, and support considerations
- Developed a Python Script to load the CSV files into the S3 buckets and created AWS S3buckets, performed folder management in each bucket, managed logs and objects within each bucket.
- Design and build ETL workflows, leading the efforts of programming data extraction from various sources into Hadoop file system, implement end to end ETL workflows using Teradata, SQL, TPT, SQOOP and load to HIVE data stores
- Assist with the analysis of data used for the tableau reports and creation of dashboards.
- Design and implement large scale distributed solutions in AWS.
- Analyze and develop programs by considering the extract logic and the data load type using Hadoop ingest processes using relevant tools such as Sqoop, Spark, Scala, Kafka, Unix shell scripts and others.
- Developing Apache Spark jobs for data cleansing and pre-processing.
- Developed UDFs in Java as and when necessary to use in PIG and HIVE queries.
- Automated the cloud deployments using chef, python and AWS Cloud Formation Templates.
- Optimized Map Reduce Jobs to use HDFS efficiently by using various compression mechanisms.
- Used ORC and Parquet file formats in Hive.
- Development of efficient pig and hive scripts with joins on datasets using various techniques.
- Write documentation of program development, subsequent revisions and coded instructions in the project related GitHub repository
- Writing spark programs to improve the performance and optimization of the existing algorithms in Hadoop using spark context, spark-sql, data frame, pair RDD's, spark yarn.
- Using Scala language to write programs for faster testing and processing of data.
- Writing code and creating hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Created functions and assigned roles in AWS Lambda to run python scripts, and AWS Lambda using java to perform event driven processing. Created Lambda jobs and configured Roles using AWS CLI.
- Experience in change implementation, monitoring and troubleshooting of AWS Snowflake databases and cluster related issues
Environment: RHEL, HDFS, Python, Django, Flask, Pyspark, Map-Reduce, Hive, Snowflake, AWS, EC2, S3, Lambda, Redshift, Pig, Sqoop, Oozie, Teradata, Oracle SQL, UC4, Kafka, GitHub, Hortonworks data platform distribution, Spark, Scala.
Confidential, New York, NY
Data Engineer
Responsibilities:
- Developed Map Reduce/Spark Python modules for machine learning & predictive analytics in Hadoop on AWS. Implemented a Python-based distributed random forest via Python streaming.
- Created a data-profiling dashboard by leveraging podium internal architecture, which drastically reduced the time to analyze data quality using Looker reporting.
- Imported data from RDBMS to HDFS and Hive using Sqoop on regular basis.
- Involved in Kafka and building use case relevant to our environment.
- Documented the requirements including the available code which should be implemented using Spark, Hive, HDFS, HBase and Elastic Search.
- Implemented SQL scripts to take over certain spark jobs that used mainly SparkSQL commands, as indexing is more efficient in SQL. Also documented run times for easier comparison.
- Developed Spark code using Scala for faster testing and processing of data.
- Handled optimization of spark scripts for better performance on a whole as well as validate and test changes.
- Helped Debug multiple bugs that were responsible for data not being shown on the client side.
- Developed Oozie workflow jobs to execute hive, Sqoop and MapReduce actions.
- Provided thought leadership for architecture and the design of Big Data Analytics solutions for customers, actively drive Proof of Concept (POC) and Proof of Technology (POT) evaluations and to implement a Big Data solution.
- Built tools using Tableau to allow internal and external teams to visualize and extract insights from big data platforms.
- Implemented Oozie Operational Services for batch processing and scheduling workflows dynamically.
- Involved in designing and developing tables in HBase and storing data
- By using Kafka HDFS connector load the data to the Hadoop clusters and integrate with Hive.
- Worked on Data modelling, Advanced SQL with Columnar Databases using AWS
Environment: Hadoop, AWS, HDFS, SQL, PL/SQL, Hive, Sqoop, T-SQL, Snowflake, OLTP, Pig, HBase, Teradata, ETL, Oracle, Tableau.
Confidential
Hadoop Developer
Responsibilities:
- Managed the imported data from different data sources, performed transformation using Hive and Map- Reduce and loaded data in HDFS.
- Involved in designing the row key in HBase to store Text and JSON as key values in HBase table and designed row key in such a way to get/scan it in a sorted order.
- Recommended improvements and modifications to existing data and ETL pipelines.
- Experience in writing stored procedures and complex SQL queries using relational databases like Oracle, SQL Server and MySQL.
- Performed Data Preparation by using Pig Latin to get the right data format needed.
- Experienced in working with spark ecosystem using Spark SQL and Scala queries on different formats like text file, CSV file.
- Created Hive schemas using performance techniques like partitioning and bucketing.
- Used Hadoop YARN to perform analytics on data in Hive.
- Developed a python script to transfer data from on-premises to AWS S3
- Used Hive to implement data warehouse and stored data into HDFS. Stored data into Hadoop clusters which are set up in AWS EMR.
- Developed and maintained batch data flow using HiveQL and Unix scripting
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDD, Scala and Python.
- Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself.
- Worked extensively with Sqoop for importing metadata from Oracle.
- Involved in creating Hive tables and loading and analyzing data using hive queries.
Environment: Hadoop, MapReduce, AWS, HBase, JSON, Spark, Kafka, Hive, Pig, Hadoop YARN, Spark Core, Spark SQL, Scala, Python, Java, Hive, Sqoop, Impala, Oracle, Yarn, Linux, Oozie.
Confidential
Hadoop Developer
Responsibilities:
- Installed and configuredHive, Pig, Sqoop, FlumeandOozieon the Hadoop cluster.
- Used Hive and created Hive tables and involved in data loading and writing Hive UDFs.
- Used Sqoop to import data into HDFS and Hive from other data systems.
- Continuous monitoring and managing theHadoop clusterthroughCloudera Manager.
- Developed Hive queries to process the data for visualizing.
- Setup and benchmarked Hadoop/Hbase clusters for internal use.
- Developed JavaMap Reduce programsfor the analysis of sample log file stored in cluster.
- Developed Map Reduce Programs for data analysis and data cleaning.
- Developed PIG Latin scripts for the analysis of semi structured data.
- Developed Hive queries to process the data and generate the data cubes for visualizing.
- Responsible for data extraction and data ingestion from different data sources into Hadoop Data Lake by creating ETL pipelines using Pig, and Hive
- Importing and exporting data intoHDFSandHiveusingSqoop.
- InstalledOozie workflowengine to run multiple Hive and Pig Jobs.
- Involved in loading data fromUNIXfile system toHDFS.
Environment: ApacheHadoop, HDFS, Cloudera Manager, Pig, Sqoop, Flume, HDFS, Hbase, Hive, MapReduce
