Senior Big Data Engineer Resume
Weehawken, NJ
SUMMARY
- Over 8+ years of diversified experience in Software Design & Development. Experience as Big Data Engineer solving business use cases for several clients. Experience in the field of software with expertise in backend applications.
- Experienced working with various Hadoop Distributions (Cloudera, Hortonworks, Map R, Amazon EMR) to fully implement and leverage new Hadoop features.
- Experience in Implement frameworks to import and export data from Hadoop to RDBMS.
- Deep knowledge of troubleshooting and tuning Spark applications and Hive scripts to achieve optimal performance.
- Experience in manipulating/analysing large datasets and finding patterns and insights within structured and unstructured data.
- Strong experience with ETL and/or orchestration tools (e.g. Talend, Oozie, Airflow)
- Experience in using Teradata ETL tools and utilizes such as BTEQ, MLOAD, FASTLOAD, TPT, and Fast Export.
- Experience setting up AWS Data Platform - AWS CloudFormation, Development End Points, AWS Glue, EMR and Jupyter/Sagemaker Notebooks, Redshift, S3, and EC2 instances
- Experience in developing Spark Applications using Spark RDD, Spark - SQL and Data frame APIs.
- Worked with real-time data processing and streaming techniques using Spark streaming and Kafka.
- Experience in moving data into and out of the HDFS and Relational Database Systems (RDBMS) using Apache Sqoop.
- Replaced existing MR jobs and Hive scripts with Spark SQL & Spark data transformations for efficient data processing.
- Experience in Developing Spark applications using Spark - SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns
- Experience developing Kafka producers and Kafka Consumers for streaming millions of events per second on streaming data
- Experience in developingcustomUDFsfor Pig and Hive to in corporate methods and functionality of Python/Java intoPig LatinandHQL(HiveQL) and Used UDFs from Piggybank UDF Repository.
- Experience on Migrating SQL database to Azure Data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and Controlling and granting database access and Migrating On premise databases to Azure Data lake store using Azure Data factory
- Good understanding of the Data modelling (Dimensional & Relational) concepts like Star-Schema Modelling, a Schema Modelling, Fact and Dimension tables.
- Experience in writing complex SQL queries, creating reports and dashboards.
- Proficient in using Unix based Command Line Interface.
- Used Informatica Power Center for (ETL) extraction, transformation and loading data from heterogeneous source systems into target database
- Implemented frameworks to extract, transform, load data from various sources.
- Deep knowledge of troubleshooting and tuning Spark applications and Hive scripts to achieve optimal performance
- Database design, modeling, migration and development experience in using stored procedures, triggers, cursor, constraints and functions. Used My SQL, MS SQL Server, DB2, and Oracle
- Experience working with NoSQL database technologies, including MongoDB, Cassandra and HBase.
- Expertise in working with HIVE data warehouse infrastructure-creating tables, data distribution by implementing Partitioning and Bucketing, developing and tuning the HQL queries.
- Experience with Software development tools such as JIRA, Play, GIT.
- Experienced in using Agile methodologies including extreme programming, SCRUM and Test-Driven Development (TDD)
TECHNICAL SKILLS
Big Data Tools: Hadoop, HDFS, Map Reduce, Spark, Airflow, Nifi, HBase, Hive, Pig, Sqoop, Kafka, Oozie, Zookeeper
Operating System: Windows, Unix, Sun Solaris
Programming Languages: Python, SQL, PL/SQL, Scala, and UNIX
Methodologies: RAD, JAD, System Development Life Cycle (SDLC), Agile
Cloud Platform: AWS (Amazon Web Services), Microsoft Azure
Cloud Management: Amazon Web Services (AWS)- EC2, EMR, S3, Redshift, EMR, Lambda, Athena
Data Modeling Tools: Erwin Data Modeler, ER Studio v17
OLAP Tools: Tableau, SSAS, Business Objects, and Crystal Reports 9
Databases: Oracle 12c/11g, Teradata R15/R14.
ETL/Data warehouse Tools: Informatica 9.6/9.1, and Tableau.
PROFESSIONAL EXPERIENCE
Confidential, Weehawken, NJ
Senior Big Data Engineer
Responsibilities:
- Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, Oracle, MongoDB, T-SQL, and SQL Server usingPython.
- Used Airflow for scheduling the Hive, Spark and MapReduce jobs.
- Developed Spark/Scala, Python for regular expression (regex) project in the Hadoop/Hive environment with Linux/Windows for big data resources.
- Data sources are extracted, transformed and loaded to generate CSV data files with Python programming and SQL queries.
- Analyzing SQL scripts and designed the solution to implement using PySpark
- Export tables from Teradata to HDFS using Sqoop and build tables in Hive.
- Loaded and transformed large sets of structured, semi structured and unstructured data usingHadoop/Big Data concepts.
- Use SparkSQL to load JSON data and create Schema RDD and loaded it into Hive Tables and handled structured data using SparkSQL.
- Worked on SQL Server concepts SSIS (SQL Server Integration Services), SSAS (Analysis Services) and SSRS (Reporting Services). Using Informatica & SSIS, SPSS, SAS to extract transform & load source data from transaction systems.
- Developed reusable objects like PL/SQL program units and libraries, database procedures and functions, database triggers to be used by the team and satisfying the business rules.
- Involved with writing scripts in Oracle, SQL Server and Netezza databases to extract data for reporting and analysis and Worked in importing and cleansing of data from various sources like DB2, Oracle, flat files onto SQL Server with high volume data
- Developed Spark code using Scala and Spark-SQL/Streaming for faster processing of data.
- Develop RDD's/Data Frames in Spark using and apply several transformation logics to load data from Hadoop Data Lakes.
- Utilized Apache Spark with Python to develop and execute Big Data Analytics and Machine learning applications, executed machine Learning use cases under Spark ML and Mllib.
- Developed Spark Streaming job to consume the data from the Kafka topic of different source systems and push the data into HDFS locations.
- Converting Hive/SQL queries into Spark transformations using Spark RDDs and Pyspark
- Filtering and cleaning data using Scala code and SQL Queries
- Troubleshooting errors in Hbase Shell/API, Pig, Hive and MapReduce.
- Implemented Installation and configuration of multi-node cluster on Cloud using Amazon Web Services (AWS) onEC2.
- Using Python in spark to extract the data from Snowflake and upload it to Salesforce on Daily basis.
- Worked with Hadoop ecosystem and Implemented Spark using Scala and utilized Data frames and Spark SQL API for faster processing of data.
- Designed and Developed Real Time Stream Processing Application using Spark, Kafka, Scala and Hive to perform Streaming ETL and apply Machine Learning.
- Use python to write a service which is event based using AWS Lambda to achieve real time data to One-Lake (A Data Lake solution in Cap-One Enterprise).
- Used Talend for Big Data Integration using Spark and Hadoop.
- Responsible for analyzing large data sets and derive customer usage patterns by developing new MapReduce programs using Java.
- Perform structural modifications using MapReduce, Hive and analyze data using visualization/reporting tools (Tableau).
- Developing Spark programs with Python, and applied principles of functional programming to process the complex structured data sets.
- Designed Kafka producer client using Confluent Kafka and produced events into Kafka topic.
- Responsible for gathering requirements, system analysis, design, development, testing and deployment and
- Worked with Hadoop infrastructure to storage data in HDFS storage and use Spark / HIVE SQL to migrate underlying SQL codebase in AWS.
- Subscribing the Kafka topic with Kafka consumer client and process the events in real time using spark.
- Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregation on the fly to build the common learner data model and persists the data in HDFS.
Environment: Hadoop, Spark, Scala, Hbase, Hive, Python, PL/SQL AWS, EC2, S3, Lambda, Auto Scaling, Cloud Watch, Cloud Formation, IBM Info sphere, DataStage, MapReduce, Oracle12c, Flat files, TOAD, MS SQL Server database, XML files, Cassandra, MongoDB, Kafka, MS Access database, Autosys, UNIX, Erwin.
Confidential, Vernon Hills, IL
Big Data Engineer
Responsibilities:
- Developed thefeatures,scenarios,step definitionsforBDD (Behavior Driven Development)andTDD (Test Driven Development)usingCucumber, Gherkinandruby.
- Involved in all the steps and scope of the project reference data approach to MDM, have created a Data Dictionary and Mapping from Sources to the Target in MDM Data Model.
- Experience managing Azure Data Lakes (ADLS) and Data Lake Analytics and an understanding of how to integrate with other Azure Services.
- Wrote production level Machine Learning classification models and ensemble classification models from scratch using Python and PySpark to predict binary values for certain attributes in certain time frame.
- ImplementedSQOOPfor large dataset transfer between Hadoop and RDBMS.
- Data visualization:Pentaho, Tableau, D3. Have knowledge of Numerical optimization, Anomaly Detection and estimation, A/B testing, Statistics, and Maple. Have big data analysis technique using Big data related techniques i.e.,Hadoop, Map Reduce, NoSQL, Pig/Hive, Spark/Shark, MLlibandScala, NumPy, SciPy, Pandas, scikit-learn.
- Work on data that was a combination of unstructured and structured data from multiple sources and automate the cleaning usingPython scripts.
- Creating Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform, and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.
- Designed end to end scalable architecture to solve business problems using various Azure Components like HDInsight, Data Factory, Data Lake, Storage and Machine Learning Studio.
- Developed multipleMapReducejobs to perform data cleaning and pre-processing.
- Used SQL Server Integrations Services (SSIS) for extraction, transformation, and loading data into target system from multiple sources
- Designed both 3NF data models for OLTP systems and dimensional data models using star and snowflake Schemas.
- Created and maintained SQL Server scheduled jobs, executing stored procedures for the purpose of extracting data from Oracle into SQL Server. Extensively used Tableau for customer marketing data visualization
- Used ApacheSpark Data frames, Spark-SQL, Spark MLlibextensively and developing and designing POC's using Scala, Spark SQL and MLlib libraries.
- Transforming business problems into Big Data solutions and define Big Data strategy and Roadmap. Installing, configuring, and maintaining Data Pipelines
- Built a new CI pipeline. Testing and deployment automation with Docker, Swamp, Jenkins and Puppet. Utilized continuous integration and automated deployments with Jenkins and Docker.
- UtilizedSpark, Scala, Hadoop, HBase, Cassandra, MongoDB, Kafka, Spark Streaming, MLlib, Pythonand utilized the engine to increase user lifetime by 45% and triple user conversations for target categories.
- Improve fraud prediction performance by using random forest and gradient boosting for feature selection withPython Scikit-learn.
- Designed and developed architecture for data services ecosystem spanning Relational, NoSQL, and Big Data technologies.
- Developed JSON Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the SQL Activity. Build an ETL which utilizes spark jar inside which executes the business analytical model
- Data Integrationingests, transforms, and integrates structured data and delivers data to a scalable data warehouse platform using traditional ETL (Extract, Transform, Load) tools and methodologies to collect of data from various sources into a single data warehouse.
- Performed all necessary day-to-day GIT support for different projects, Responsible for design and maintenance of the GIT Repositories, and the access control strategies.
Environment: Hadoop, Kafka, Spark, Sqoop, Docker, Swamp, Azure, Azure HD Insight, Spark SQL, TDD, Spark-Streaming, Hive, Scala, pig, Azure Data Bricks, Azure Data Storage, Azure Data Lake, Azure SQL, No SQL, Impala, Oozie, Hbase, Data Lake, Zookeeper.
Confidential, New York, NY
Data Engineer
Responsibilities:
- Involved in transforming data from Mainframe tables toHDFS, andHBasetables using Sqoop.
- Visualized the results using Tableau dashboards and the Python Seaborn libraries were used for Data interpretation in deployment.
- UsedRest APIto Access HBase data to performanalytics.
- Involved in creatingHivetables, loading with data and writingHive queriesthat will run internally in MapReduce way
- Acted for bringing in data underHBaseusing HBase shell alsoHBaseclient API.
- Experienced with handling administration activations usingClouderamanager
- Created and maintained Technical documentation for launching Hadoop Clusters and for executingPigScripts.
- Worked on POC for IOT devices data, with spark.
- Automatically scale-up the EMR instances based on the data.
- DevelopedMapReduceprograms to process theAvrofiles and to get the results by performing some calculations on data and also performed map side joins.
- Imported Bulk Data intoHBaseUsing MapReduce programs.
- Involved in migrating tables fromRDBMSintoHivetables usingSQOOPand later generate visualizations using Tableau.
- Experience working with ApacheSOLRfor indexing and querying.
- Created customSOLRQuery segments to optimize ideal search matching.
- Stored the time-series transformed data from the Spark engine built on top of a Hive platform to Amazon S3 and Redshift.
- Facilitated deployment of multi-clustered environment using AWS EC2 and EMR apart from deploying Dockers for cross-functional deployment.
- Involved in writing optimizedPigScript along with developing and testingPig LatinScripts.
- Involved in collecting, aggregating and moving data from servers to HDFS usingFlume.
- Imported and Exported Data from Different Relational Data Sources like DB2, SQL Server, Teradata to HDFS usingSqoop.
- Involved in data ingestion intoHDFSusingSqoopfor full load and Flume for incremental load on variety of sources like web server,RDBMSand Data API’s.
- Collected data using Spark Streaming fromAWSS3bucket in near-real- time and performs necessary Transformations and Aggregations to build the data model and persists the data inHDFS.
- InstalledOozieworkflow engine to run multipleHiveandPigjobs which run independently with time and data availability.
- UsedSCALAto storestreaming datato HDFS and to implementSparkfor faster processing of data.
- Worked on creating theRDD's,DF's for the required input data and performed the data transformations using Spark Python.
- Migrated complex map reduce programs intoin memory Sparkprocessing using Transformations and actions.
- Designed and implemented Incremental Imports intoHivetables and writing Hive queries to run onTEZ.
- CreatedETLMapping with Talend Integration Suite to pull data from Source, apply transformations, and load data into target database.
- Designed and implemented Incremental Imports intoHivetables.
- Ingest real-time and near-real time (NRT) streaming data intoHDFSusingFlume.
- Worked with NoSQL databases likeHBasein makingHBasetables to load expansive arrangements of semi structured data.
Environment: Hadoop, Cloudera, Flume, HBase, HDFS, MapReduce, AWS, YARN, Hive, Pig, Sqoop, Oozie, Tableau, Java, Solr.
Confidential, Houston, TX
Hadoop Developer
Responsibilities:
- Importing and exporting data into HDFS from Oracle Database and vice versa using Sqoop.
- Created batch jobs and configuration files to create automated process using SSIS.
- Wrote MapReduce job using Pig Latin. Involved in ETL, Data Integration and Migration.
- Creating Hive tables and working on them using Hive QL. Experienced indefining jobflows.
- Involved in creating Hive tables, loading the data and writing hive queries that will run internally in a map reduce way. Developed a custom File System plugin for Hadoop so it can access files on Data Platform.
- Installed and configured Pig and also written Pig Latin scripts.
- Designed and implemented MapReduce-based large-scale parallel relation-learning system
- Setup and benchmarked Hadoop/HBase clusters for internal use
- Created SSIS packages to pull data from SQL Server and exported to Excel Spreadsheets and vice versa.
- Deploying and scheduling reports using SSRS to generate daily, weekly, monthly and quarterly reports.
- Loading data from various sources like OLEDB, flat files to SQL Server database Using SSIS Packages and created data mappings to load the data from source to destination.
- The custom File System plugin allows Hadoop MapReduce programs, HBase, Pig and Hive to work unmodified and access files directly.
- Extensive use of Expressions, Variables, Row Count in SSIS packages
- Data validation and cleansing of staged input records was performed before loading into Data Warehouse
Environment: Hadoop, MapReduce, Pig,MS SQL Server, SQL Server Business Intelligence Development Studio, Hive, Hbase, SSIS, SSRS, Report Builder, Office, Excel, Flat Files, .NET, T-SQL.
Confidential
Data Analyst
Responsibilities:
- Used Informatica Repository Manager for managing all the repositories (development, test & validation), was also involved in migration of folders from one repository to another.
- Involved in migrating PowerCenter folders from Development to Production Repository using Repository Manager.
- Produced a Unit Test Document, which captures the test conditions and scripts, expected/actual results.
- Created Mappings using Mapping Designer to load the data from various sources using different transformations like Aggregator, Expression, Stored Procedure, External Procedure, Filter, Joiner, Lookup, Router, Sequence Generator, Source Qualifier, and Update Strategy transformations.
- Created Mapping Parameters and Variables.
- Extracted data from sources like Oracle and Fixed width and Delimited Flat files. Transformed the data according to the business requirements and then Loaded into the Oracle.
- Modified several of the existing mappings and created several new mappings based on the user requirements.
- Ran data loads into all the environments using DAC (Data warehouse Application Console), DAC is a centralized console, providing access to the entire Siebel Data Warehouse application. It allows you to create, configure, and execute modular data warehouse applications in a parallel, high performing environment.
- Extensively involved in testing by writing some QA procedures, for testing the target data against source data.
- Developed the design document for each ETL mapping, defining the source and target tables and all the fields, transformations and the join condition, which helped the users to better understand the type of data.
- Maintained existing mappings by resolving performance issues.
Environment: Informatica Power Center, Siebel, DAC (Data warehouse Application Console),HPUNIX, Windows,Oracle 9i, SQL, PL/SQL, SQL * Loader, TOAD, Erwin.
