We provide IT Staff Augmentation Services!

Senior Big Data Engineer Resume

2.00/5 (Submit Your Rating)

Coppell, TX

SUMMARY

  • Over 8+ years of experience in Data Engineering, Data Pipeline Design, Development and Implementation as a Sr. Big Data Engineer/ Data Developer and Data Modeler.
  • Strong experience in Software Development Life Cycle (SDLC) including Requirements Analysis, Design Specification and Testing as per Cycle in both Waterfall and Agile methodologies.
  • Extensively worked with Teradata utilities Fast export, and Multi Load to export and load data to/from different source systems including flat files.
  • Experienced in building Automation Regressing Scripts for validation of ETL process between multiple databases like Oracle, SQL Server, Hive, and Mongo DB usingPython.
  • Experience in setting up monitoring infrastructure for Hadoop cluster using Nagios and Ganglia.
  • Sustaining teh BigQuery, PySpark and Hive code by fixing teh bugs and providing teh enhancements required by teh Business User.
  • Very good exposure in OLAP and OLTP.
  • Proficient inStatistical MethodologiesincludingHypothetical Testing,ANOVA,Time Series,Principal Component Analysis,Factor Analysis,Cluster Analysis,Discriminant Analysis.
  • Expertise in Python andScala, user - defined functions (UDF) for Hive and Pig using Python.
  • Experience in developing Map Reduce Programs using Apache Hadoop for analyzing teh big data as per teh requirement.
  • Hands on Spark MLlib utilities such as including classification, regression, clustering, collaborative filtering, dimensionality reduction.
  • Experience in working with Flume and NiFi for loading log files into Hadoop.
  • Experience in working with NoSQL databases like HBase and Cassandra.
  • Experienced in creating shell scripts to push data loads from various sources from teh edge nodes onto teh HDFS.
  • Good Experience in implementing and orchestrating data pipelines using Oozie and Airflow.
  • Worked with Cloudera and Hortonworks distributions.
  • Good working knowledge of Amazon Web Services(AWS) Cloud Platform which includes services likeEC2,S3,VPC,ELB, IAM, DynamoDB, Cloud Front, Cloud Watch, Route 53, Elastic Beanstalk (EBS), Auto Scaling, Security Groups, EC2 Container Service (ECS), Code Commit, Code Pipeline, Code Build, Code Deploy,DynamoDB, Auto Scaling, Security Groups, Red shift, CloudWatch, CloudFormation, CloudTrail, Ops Works, Kinesis, IAM, SQS, SNS, SES.
  • Knowledge of working with Proof of Concepts (PoC's) and gap analysis and gathered necessary data for analysis from different sources, prepared data for data exploration using data munging and Teradata.
  • Well experience in Normalization and De-Normalization techniques for optimum performance in relational and dimensional database environments.
  • Proficiency in SQL across several dialects (we commonly write MySQL, PostgreSQL, Redshift, SQL Server, and Oracle)
  • Strong experience in writing scripts usingPythonAPI, PySpark API and Spark API for analyzing teh data.
  • Extensively usedPythonLibraries PySpark, Pytest, Pymongo, cxOracle, PyExcel, Boto3, Psycopg, embedPy, NumPy and Beautiful Soup.
  • Experience in Microsoft Azure/Cloud Services like SQL Data Warehouse, Azure SQL Server, Azure Databricks, Azure Data Lake, Azure Blob Storage, Azure Data Factory
  • Worked with various text analytics libraries like Word2Vec, GloVe, LDA and experienced with Hyper Parameter Tuning techniques like Grid Search, Random Search, model performance tuning using Ensembles and Deep Learning.
  • Skilled in System Analysis, E-R/Dimensional Data Modeling, Database Design and implementing RDBMS specific features.
  • Experience in developing customizedUDF’sin Python to extend Hive and Pig Latin functionality.
  • Expertise in designing complex Mappings and have expertise in performance tuning and slowly changing Dimension Tables and Fact tables
  • Skilled in performing data parsing, data ingestion, data manipulation, data architecture, data modelling and data preparation with methods including describe data contents, compute descriptive statistics of data, regex, split and combine, Remap, merge, subset, reindex, melt and reshape.
  • Hands-on use of Spark andScalaAPI's to compare teh performance of Spark with Hive and SQL, and Spark SQL to manipulate Data Frames inScala.
  • Experience in developing customizedUDF’sin Python to extend Hive and Pig Latin functionality.
  • Expertise in designing complex Mappings and have expertise in performance tuning and slowly changing Dimension Tables and Fact tables
  • Experienced in building Automation Regressing Scripts for validation of ETL process between multiple databases like Oracle, SQL Server, Hive, and Mongo DB usingPython.
  • Proficiency in SQL across several dialects (we commonly write MySQL, PostgreSQL, Redshift, SQL Server, and Oracle)
  • Experience in designing star schema, Snowflake schema for Data Warehouse, ODS architecture.
  • Well experience in Normalization and De-Normalization techniques for optimum performance in relational and dimensional database environments.
  • Good knowledge of Data Marts, OLAP, Dimensional Data Modeling with Ralph Kimball Methodology (Star Schema Modeling, Snow-Flake Modeling for FACT and Dimensions Tables) using Analysis Services.
  • Excellent communication skills. Successfully working in fast-paced multitasking environment both independently and in collaborative team, a self-motivated enthusiastic learner.

TECHNICAL SKILLS

Programming languages: Python, PySpark, Shell Scripting, SQL, PL/SQL and UNIX Bash

Big Data: Hadoop, Sqoop, Apache Spark, NiFi, Kafka, Snowflake, Cloudera, Horton Works, PySpark, Spark, Spark SQL

Data Modeling Tools: Erwin Data Modeler, ER Studio v17

Operating Systems: UNIX, LINUX, Solaris, Mainframes

IDE Tools: Aginitiy for Hadoop, PyCharm, Toad, SQL Developer, SQL *Plus, Sublime Text, VI Editor

OLAP Tools: Tableau, SSAS, Business Objects, and Crystal Reports 9

ETL/Data warehouse Tools: Informatica 9.6/9.1, and Tableau.

Data bases: Oracle, SQL Server, My SQL, DB2, Sybase, Netezza, Hive, Impala

Cloud Technologies: AWS, Microsoft AZURE

Others: AutoSys, Crontab, ArcGIS, Clarity, Informatica, Business Objects, IBM MQ, Splunk

PROFESSIONAL EXPERIENCE

Confidential, Coppell, TX

Senior Big Data Engineer

Responsibilities:

  • Work in a fast-paced agile development environment to quickly analyze, develop, and test potential use cases for teh business.
  • Teh individual will be responsible for design and development of High-performance data architectures which support data warehousing, real-time ETL and batch big-data processing.
  • Worked with Hadoop infrastructure to storage data in HDFS storage and use Spark / HIVE SQL to migrate underlying SQL codebase in AWS.
  • Converting Hive/SQL queries into Spark transformations using Spark RDDs and Pyspark
  • Analyzing SQL scripts and designed teh solution to implement using PySpark
  • Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, Oracle, MongoDB, T-SQL, and SQL Server usingPython.
  • Involved as primary on-site ETL Developer during teh analysis, planning, design, development, and implementation stages of projects using IBM Web Sphere software (Quality Stage v9.1, Web Service, Information Analyzer, Profile Stage)
  • Prepared Data Mapping Documents and Design teh ETL jobs based on teh DMD with required Tables in teh Dev Environment.
  • Implemented Installation and configuration of multi-node cluster on Cloud using Amazon Web Services (AWS) onEC2.
  • Designed and developed architecture for data services ecosystem spanning Relational, NoSQL, and Big data technologies. Extracted Mega Data from Amazon Redshift, AWS, and Elastic Search engine using SQL Queries to create reports.
  • Used Talend for Big Data Integration using Spark and Hadoop.
  • Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregation on teh fly to build teh common learner data model and persists teh data in HDFS.
  • Generate metadata, create Talend etl jobs, mappings to load data warehouse, data lake.
  • Designed and Developed Real Time Stream Processing Application using Spark, Kafka, Scala and Hive to perform Streaming ETL and apply Machine Learning.
  • Involved in Relational and Dimensional Data modeling for creating Logical and Physical Design of Database and ER Diagrams with all related entities and relationship with each entity based on teh rules provided by teh business manager using ERWIN r9.6.
  • Export tables from Teradata to HDFS using Sqoop and build tables in Hive.
  • Loaded and transformed large sets of structured, semi structured and unstructured data usingHadoop/Big Data concepts.
  • Developing Spark programs with Python, and applied principles of functional programming to process teh complex structured data sets.
  • Use Spark SQL to load JSON data and create Schema RDD and loaded it into Hive Tables and handled structured data using Spark SQL.
  • Worked with Hadoop ecosystem and Implemented Spark using Scala and utilized Dataframes and Spark SQL API for faster processing of data.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster processing of data.
  • Develop RDD's/Data Frames in Spark using and apply several transformation logics to load data from Hadoop Data Lakes.
  • Filtering and cleaning data using Scala code and SQL Queries
  • Extensively worked with Join, Look up (Normal and Sparse) and Merge stages.
  • Applying teh Data modelling and Data Designing in-between staging and target for creating teh views.
  • Responsible for gathering requirements, system analysis, design, development, testing and deployment and
  • Worked on SQL Server concepts SSIS (SQL Server Integration Services), SSAS (Analysis Services) and SSRS (Reporting Services). Using Informatica & SSIS, SPSS, SAS to extract transform & load source data from transaction systems.
  • Developed reusable objects like PL/SQL program units and libraries, database procedures and functions, database triggers to be used by teh team and satisfying teh business rules.
  • Involved with writing scripts in Oracle, SQL Server and Netezza databases to extract data for reporting and analysis and Worked in importing and cleansing of data from various sources like DB2, Oracle, flat files onto SQL Server with high volume data
  • Evaluated big data technologies and prototype solutions to improve our data processing architecture. Data modeling, development and administration of relational and NoSQL databases (Big Query, Elastic Search)
  • Utilized Spark, Scala, Hadoop, HBase, Cassandra, MongoDB, Kafka, Spark Streaming, a broad variety of machine learning methods including classifications, regressions, dimensionally reduction etc.
  • Used Informatica power center for (ETL) extraction, transformation and loading data from heterogeneous source systems and studied and reviewed application of Kimball data warehouse methodology as well as SDLC across various industries to work successfully with data-handling scenarios, such as data
  • Experience with Data Analytics, Data Reporting, Ad-hoc Reporting, Graphs, Scales, PivotTables and OLTP reporting.

Environment: Hadoop, Spark, Scala, HBase, Hive, IBM Info Sphere Data Stage, PL/SQL, Oracle 12c, Flat files,Autosys, UNIX, Erwin, TOAD, MS SQL Server database, XML files, AWS, Cassandra, MongoDB, Kafka, MS Access database.

Confidential, Sunnyvale, CA

Senior Big Data Engineer

Responsibilities:

  • Designing teh business requirement collection approach based on teh project scope and SDLC methodology.
  • Creating Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform, and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.
  • Files extracted from Hadoop and dropped on daily hourly basis intoS3. Working with Data governance and Data quality to design various models and processes.
  • Involved in all teh steps and scope of teh project reference data approach to MDM, have created a Data Dictionary and Mapping from Sources to teh Target in MDM Data Model.
  • Experience managing Azure Data Lakes (ADLS) and Data Lake Analytics and an understanding of how to integrate with other Azure Services. Knowledge of USQL
  • Responsible for working with various teams on a project to develop analytics-based solution to target customer subscribers specifically.
  • Architect & implement medium to large scale BI solutions on Azure using Azure Data Platform services (Azure Data Lake, Data Factory, Data Lake Analytics, Stream Analytics, Azure SQL DW, HDInsight/Databricks, NoSQL DB).
  • Migration of on premise data (Oracle/ SQL Server/ DB2/ MongoDB) to Azure Data Lake and Stored (ADLS) using Azure Data Factory (ADF V1/V2).
  • Transforming business problems into Big Data solutions and define Big Data strategy and Roadmap. Installing, configuring, and maintaining Data Pipelines
  • Developed thefeatures,scenarios,step definitionsforBDD (Behavior Driven Development)andTDD (Test Driven Development)usingCucumber, Gherkinandruby.
  • Responsible for wide-ranging data ingestion using Sqoop and HDFS commands. Accumulate ‘partitioned’ data in various storage formats like text, Json, Parquet, etc. Involved in loading data from LINUX file system to HDFS
  • Monitored cluster health by Setting up alerts using Nagios and Ganglia
  • Data Integrationingests, transforms, and integrates structured data and delivers data to a scalable data warehouse platform using traditional ETL (Extract, Transform, Load) tools and methodologies to collect of data from various sources into a single data warehouse.
  • Applied variousmachine learning algorithmsand statistical modeling likedecision trees, text analytics, natural language processing (NLP),supervised and unsupervised, regression models, social network analysis, neural networks, deep learning, SVM, clusteringto identify Volume usingscikit-learn packageinpython, R, and Matlab. Collaborate withData Engineers and Software Developersto develop experiments and deploy solutions to production.
  • Designed and developed architecture for data services ecosystem spanning Relational, NoSQL, and Big Data technologies.
  • Used SQL Server Integrations Services (SSIS) for extraction, transformation, and loading data into target system from multiple sources
  • Built real time pipeline for streaming data usingKafkaandSparkStreaming.
  • Involved inUnit Testingthe code and provided teh feedback to teh developers. PerformedUnit Testingof teh application by usingNUnit.
  • Designed both 3NF data models for OLTP systems and dimensional data models using star and snowflake Schemas.
  • Created and maintained SQL Server scheduled jobs, executing stored procedures for teh purpose of extracting data from Oracle into SQL Server. Extensively used Tableau for customer marketing data visualization
  • Optimizealgorithmwithstochastic gradient descent algorithmFine-tuned thealgorithm parameterwith manual tuning and automated tuning such asBayesian Optimization.
  • Working on tickets opened by users regarding various incidents, requests
  • Writing UNIX shell scripts to automate teh jobs and scheduling cron jobs for job automation using commands with Crontab.
  • Developed various Mappings with teh collection of all Sources, Targets, and Transformations using Informatica Designer
  • Developed Mappings using Transformations like Expression, Filter, Joiner and Lookups for better data messaging and to migrate clean and consistent data
  • Used ApacheSpark Data frames, Spark-SQL, Spark MLLibextensively and developing and designing POC's using Scala, Spark SQL and MLlib libraries.
  • Write research reports describing teh experiment conducted, results, and findings and make strategic recommendations to technology, product, and senior management. Worked closely with regulatory delivery leads to ensure robustness in prop trading control frameworks using Hadoop, Python Jupyter Notebook, Hive and NoSql.
  • Wrote production level Machine Learning classification models and ensemble classification models from scratch using Python and PySpark to predict binary values for certain attributes in certain time frame.
  • Performed all necessary day-to-day GIT support for different projects, Responsible for design and maintenance of teh GIT Repositories, and teh access control strategies.

Environment: Hadoop, Kafka, Spark, Sqoop, Docker, Spark SQL, TDD, Spark-Streaming, Hive, Scala, pig, NoSQLImpala, Oozie, Hbase, Data Lake, Zookeeper, Azure, Unix/Linux Shell Scripting,Python, PyCharm, InformaticaLinux, Shell Scripting, Informatica PowerCenter

Confidential, Charlotte, NC

Big Data Developer

Responsibilities:

  • Involved in Analysis, Design and Implementation/translation of Business User requirements.
  • Generated ad-hoc SQL queries using joins, database connections and transformation rules to fetch data from legacy DB2 and SQL Server database systems.
  • Translated business requirements into working logical and physical data models for OLTP &OLAP systems.
  • Creation of BTEQ, Fast export, Multi Load, TPump, Fast load scripts for extracting data from various production systems.
  • Reviewed Stored Procedures for reports and wrote test queries against teh source system (SQL Server-SSRS) to match teh results with teh actual report against teh DataMart (Oracle)
  • Perform Data profiling, preliminary data analysis and handle anomalies such as missing, duplicates, outliers, and imputed irrelevant data. Remove outliers using Proximity Distance and Density based techniques.
  • Responsible for analyzing large data sets to develop multiple custom models and algorithms to drive innovative business solutions
  • Involved in designing data warehouses and data lakes on regular (Oracle, SQL Server) high performance on big data (Hadoop - Hive and HBase) databases. Data modeling, Design, implement, and deploy high-performance, custom applications at scale on Hadoop /Spark.
  • Involved in scheduling Oozie workflow engine to run multiple Hive jobs, developed workflow in Oozie to automate teh tasks of loading teh data into HDFS and pre-processing with Spark
  • Performed Exploratory Data Analysis using R. Also involved in generating various graphs and charts for analyzing teh data using Python Libraries.
  • Used supervised, unsupervised and regression techniques in building models.
  • Performed Market Basket Analysis to identify teh groups of assets moving together and recommended teh client their risks
  • Developed ETL (Extraction, Transformation and Loading) procedures and Data Conversion Scripts using Pre-Stage, Stage, Pre-Target and Target tables.
  • Creating teh data pipelines using state of teh art Big Data frameworks/tools
  • Experience in extracting appropriate features from datasets in-order to handle bad, null, partial records using spark SQL.
  • Experienced in ingesting data into HDFS from various Relational databases like Teradata using sqoop and exported data back to Teradata for data storage.
  • Hands on experience in developing apache SPARK applications using Spark tools like RDD transformations, Spark core, Spark MLlib, Spark Streaming and Spark SQL.
  • Experience in developing various spark application using Spark-shell (Scala).
  • Implemented Partitioning, Dynamic Partitions, Buckets in Hive
  • Executed Hive queries on ORC tables stored in Hive to perform data analysis to meet teh business requirements
  • Utilizing Spark get data from HDFS, process it and store it back into HDFS
  • Developed a Python Script to load teh CSV files into teh S3 buckets and created AWS S3 buckets, performed folder management in each bucket, managed logs and objects within each bucket.
  • Created Airflow Scheduling scripts in Python to automate teh process of Sqooping wide range of data sets.
  • Involved in file movements between HDFS and AWS S3 and extensively worked with S3 bucket in AWS
  • Developed Spark SQL scripts using Python for faster data processing
  • Using Sqoop to extract teh data from warehouse, SQL server and load into Hive
  • Used Spark framework to transform teh data for final consumption of analytical applications
  • Worked on storing teh Dataframe into hive as table using Python (PySpark).
  • Dynamic implementation of SQL server work on website using SQL developer tool and Experience with continuous integration and automation using Jenkins and Implemented Service Oriented Architecture (SOA) using JMS for sending and receiving messages while creating web services.
  • Involved in teh execution of multiple business plans and projects Ensures business needs are being met Interpret data to identify trends to go across future data sets.
  • Simultaneously working on pilot project to move to teh environment to Amazon EMR, a cloud-based Hadoop distribution and other Amazon cloud solutions available
  • Developed interactive dashboards, created various Ad Hoc reports for users in Tableau by connecting various datasources.

Environment: Python, SQL server, Hadoop, HDFS, HBase, MapReduce, Hive, Impala, Pig, Sqoop, Mahout, LSTM,RNN, Spark MLlib, MongoDB, AWS, Tableau, Unix/Linux.

Confidential, Minneapolis, MN

Data Engineer

Responsibilities:

  • Experience in Big Data Analytics and design in Hadoop ecosystem using MapReduce Programming, Spark, Hive, Pig, Sqoop, HBase, Oozie, Impala, Kafka
  • Created logical data model from teh conceptual model and its conversion into teh physical database design using Erwin. Involved in transforming data from legacy tables toHDFS, andHBasetables usingSqoop.
  • Developed Data mapping, Transformation and Cleansing rules for teh Data Management involving OLTP and OLAP.
  • Involved in creating UNIX shell Scripting. Defragmentation of tables, partitioning, compressing and indexes for improved performance and efficiency.
  • Developed reusable objects like PL/SQL program units and libraries, database procedures and functions, database triggers to be used by teh team and satisfying teh business rules.
  • Used SQL Server Integrations Services (SSIS) for extraction, transformation, and loading data into target system from multiple sources
  • Build teh Oozie pipeline which performs several actions like file move process, Sqoop teh data from teh source Teradata or SQL and exports into teh hive staging tables and performing aggregations as per business requirements and loading into teh main tables.
  • Running of Apache Hadoop, CDH and Map-R distros, dubbedElastic MapReduce(EMR)on(EC2).
  • Performing teh forking action whenever there is a scope of parallel process for optimization of data latency.
  • Worked on different data formats such as JSON, XML and performed machine learning algorithms in Python.
  • Performed pig script which picks teh data from one Hdfs path and performs aggregation and loads into another path which later pulls populates into another domain table. Converted this script into a jar and passed as parameter in Oozie script
  • Hands on experiences on Git bash commands like Git pull to pull teh code from source and developing it as per teh requirements, Git add to add files, Git commit after teh code build and Git push to teh pre prod environment for teh code review and later used screwdriver. yaml which actually build teh code, generates artifacts which releases in to production
  • Developed and implemented R and Shiny application which showcases machine learning for business forecasting. Developed predictive models using Python & R to predict customers churn and classification of customers.
  • Partner with infrastructure and platform teams to configure, tune tools, automate tasks and guide teh evolution of internal big data ecosystem; serve as a bridge between data scientists and infrastructure/platform teams.
  • Data analysis using regressions, data cleaning, excel v-look up, histograms and TOAD client and data representation of teh analysis and suggested solutions for investors
  • Rapid model creation in Python using pandas, NumPy, sklearn, and plot.ly for data visualization. These models are tan implemented in SAS where they are interfaced with MSSQL databases and scheduled to update on a timely basis.

Environment: MapReduce, Spark, Hive, Pig, Sqoop, HBase, Oozie, Impala, Kafka, JSON, XML PL/SQL, Sql, HDFS, Unix, Python, PySpark.

Confidential

Data Engineer

Responsibilities:

  • Worked on different dataflow and control flow task, for loop container, sequence container, script task, executes SQL task and Package configuration.
  • Created batch jobs and configuration files to create automated process using SSIS.
  • Created SSIS packages to pull data from SQL Server and exported to Excel Spreadsheets and vice versa.
  • Built SSIS packages, to fetch file from remote location like FTP and SFTP, decrypt it, transform it, mart it to data warehouse and provide proper error handling and alerting
  • Created new procedures to handle complex logic for business and modified already existing stored procedures, functions, views and tables for new enhancements of teh project and to resolve teh existing defects.
  • Loading data from various sources like OLEDB, flat files to SQL Server 2012 database Using SSIS Packages and created data mappings to load teh data from source to destination.
  • Extensive use of Expressions, Variables, Row Count in SSIS packages
  • Data validation and cleansing of staged input records was performed before loading into Data Warehouse
  • Automated teh process of extracting teh various files like flat/excel files from various sources like FTP and SFTP (Secure FTP).
  • Deploying and scheduling reports using SSRS to generate daily, weekly, monthly and quarterly reports.

Environment: MS SQL Server, SQL Server Business Intelligence Development Studio, SSIS, SSRS, Report Builder, Office, Excel, Flat Files, .NET, T-SQL.

We'd love your feedback!