We provide IT Staff Augmentation Services!

Sr Big Data Engineer Resume

3.00/5 (Submit Your Rating)

Weehawken, NJ

SUMMARY

  • IT professional with 9+ years of experience offers extensive expertise in building production ready Big Data ingestion pipelines using Python, Scala, Java, Apache Spark, Apache Kafka, Akka,Hbase and SQL
  • Excellent Experience in Designing, Developing, Documenting, Testing of ETL jobs and mappings in Server and Parallel jobs using Data Stage to populate tables in Data Warehouse and Data marts.
  • Experience in usage of Hadoop distribution like Cloudera and Hortonworks.
  • Deep understanding of MapReduce with Hadoop and Spark. Good working knowledge of Big Data ecosystem like Hadoop (HDFS, Hive, Pig, Impala), Spark (SparkSQL, Spark MLlib, Spark Streaming).
  • Utilized analytical applications like SPSS, Rattle and Python to identify trends and relationships between different pieces of data, draw appropriate conclusions and translate analytical findings into risk management and marketing strategies that drive value.
  • Have experience inApache Spark, Spark Streaming, Spark SQL and NoSQLdatabases likeHBase, Cassandra, andMongoDB.
  • Capable of using AWS utilities such as EMR, S3 and cloud watch to run and monitorHadoop and Spark jobs on AWS.
  • Establishes and executes the Data Quality Governance Framework, which includes end - to-end process and data quality framework for assessing decisions that ensure the suitability of data for its intended purpose.
  • Experienced on Hadoop Ecosystem and Big Data components including Apache Spark, Scala, Python, HDFS,Map Reduce, KAFKA.
  • Enterprise information /data architect enterprise information archtechture and management.
  • Enterprise Data modelling business intelligence / analytics .
  • Worked on Snowflake Schemas and Data Warehousing
  • Good knowledge in Database Creation and maintenance of physical data models with Oracle, Teradata, Netezza, DB2, MongoDB, HBase and SQL Server databases.
  • Experienced in writing complex SQL Quires like Stored Procedures, triggers, joints, and Sub quires.
  • Interpret problems and provides solutions to business problems using data analysis, data mining, optimization tools, and machine learning techniques and statistics.
  • Experienced in fact dimensional modeling (Star schema, Snowflake schema), transactional modeling and SCD (Slowly changing dimension)
  • Worked on Microsoft azure services like HDInsight Clusters, BLOB, ADLS, Data Factory and Logic Apps and also done POC on Azure Data Bricks.
  • Experienced with JSON basedRESTfulweb services, and XML/QML basedSOAPweb services and also worked on various applications using python integrated IDEs like Sublime Text andPyCharm
  • Integrated Kafka with Spark Streaming for real time data processing.
  • Skilled in performing data parsing, data manipulation and data preparation with methods including describe data contents.
  • Hands on porting the existing on-premise Hive code migration to GCP (Google Cloud Platform) BigQuery
  • Extensive experience in Text Analytics, generating data visualizations using R, Python and creating dashboards using tools like Tableau.
  • Experience with Data Analytics, Data Reporting, Ad-hoc Reporting, Graphs, Scales, PivotTables and OLAP reporting.
  • Excellent performance in building, publishing customized interactive reports and dashboards with customized parameters including producing tables, graphs, listings using various procedures and tools such as Tableau and user-filters using Tableau.

TECHNICAL SKILLS

Big Data: Cloudera Distribution, HDFS, Yarn, Data Node, Name Node, Resource Manager, Node Manager, MapReduce, PIG, SQOOP, Kafka, Hbase, Hive, Flume, Cassandra, Spark, Storm, Scala, Impala

Programming: Python, PySpark, Scala, Java, C, C++, Shell script, Perl script, SQL, PL/SQL

Databases: Snowflake(cloud), Teradata, IBM DB2, Oracle, SQL Server, MySQL, NoSQL

Cloud Technologies: AWS, Microsoft Azure

Frameworks: Django REST framework, MVC, Hortonwork

ETL/Reporting: Ab Initio, Informatica, Tableau

Tools: PyCharm, Eclipse, Visual Studio, SQL*Plus, SQL Developer, TOAD, SQL Navigator, Query Analyzer, SQL Server Management Studio, SQL Assistance, Eclipse, Postman

Machine Learning Techniques: Linear & Logistic Regression, Classification and Regression Trees, Random Forest, Associative rules, NLP and Clustering.

Database Modelling: Dimension Modeling, ER Modeling, Star Schema Modeling, Snowflake Modeling

Visualization/ Reporting: Tableau, ggplot2, matplotlib, SSRS and Power BI

Web/App Server: UNIX server, Apache Tomcat

Operating System: UNIX, Windows, Linux, Sun Solaris

PROFESSIONAL EXPERIENCE

Confidential, Weehawken, NJ

Sr Big Data Engineer

Responsibilities:

  • Used Apache NiFi to copy data from local file system to HDP. Thorough understanding of various modules of AML including Watch List Filtering, Suspicious Activity Monitoring, CTR,CDD, and EDD.
  • Developed Spark applications using Pyspark and Spark-SQL for data extraction, transformation and aggregation from multiple file formats.
  • Used Spark Streaming to receive real time data from the Kafka and store the stream data to HDFS using Python and NoSQL databases such as HBase and Cassandra
  • Working on Docker containerizedservice to leverage the infrastructure.
  • Worked on Big data on AWS cloud services i.e. EC2, S3, EMR and DynamoDB
  • Validated the test data in DB2 tables on Mainframes and on Teradata using SQL queries.
  • Installing, configuring and maintaining Data Pipelines
  • Designed and developed Informatica Mappings and Sessions, Workflows based on business rules to load data from source Oracle tables to Teradata target tables
  • Worked on analyzing Hadoop cluster using different big data analytic tools including Flume, Pig, Hive, HBase, Oozie, Zookeeper, Sqoop, Spark and Kafka.
  • Played a key role inmigrating Cassandra, Hadoop cluster on AWS and defined different read/write strategies
  • Worked on google cloud platform (GCP) services like compute engine, cloud load balancing, cloud storage, cloud SQL, stack driver monitoring and cloud deployment manager.
  • Written multiple Map Reduce program in Java for data extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV and other compressed file formats.
  • Develop solutions to leverage ETL tools and identify opportunities for process improvements using Informatica and Python installation, configuration and deployment of product soft wares on new edge nodes that connect and contact Kafka cluster for data acquisition
  • Created scripts to readCSV, json and parquet filesfrom S3 buckets inPythonand load intoAWS S3, DynamoDB and Snowflake.
  • Files extracted from Hadoop and dropped on daily hourly basis intoS3
  • Develop Nifi workflow to pick up the data from rest API server, from data lake as well as from SFTP server and send that to Kafka broker
  • Worked on Dimensional and Relational Data Modeling using Star and Snowflake Schemas, OLTP/OLAP system, Conceptual, Logical and Physical data modeling using Erwin.
  • Automated the data processing with Oozie to automate data loading into the Hadoop Distributed File System.
  • Worked on deploying the project on the servers using Jenkins
  • Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, Oracle, Mongo DB, T-SQL, and SQL Server usingPython.
  • Designed and implemented Sqoop for the incremental job to read data from DB2 and load to Hive tables and connected to Tableau for generating interactive reports using Hive server2.
  • Defined and deployed monitoring, metrics, and logging systems on AWS.
  • Authoring Python (PySpark) Scripts for custom UDF’s for Row/ Column manipulations, merges, aggregations, stacking, data labeling and for all Cleaning and conforming tasks.
  • Using Flume, Kafka and Spark streaming to ingest real time or near real time data in HDFS.
  • Worked on Big Data Integration &Analytics based on Hadoop, SOLR, Spark, Kafka, Storm and web Methods.
  • Writing Pig Scripts to generate Map Reduce jobs and performed ETL procedures on the data in HDFS.
  • Design and implement multiple ETL solutions with various data sources by extensive SQL Scripting, ETL tools, Python, Shell Scripting and scheduling tools. Data profiling and data wrangling of XML, Web feeds and file handling using python, Unix and Sql.
  • Migrated the data from Redshift data warehouse to Snowflake.
  • Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregation on the fly to build the common learner data model and persists the data in HDFS.
  • Enhanced coverage across domains, system of records and critical data elements .
  • Provided definitions by domains to feed in to the business glossery
  • Defined stewrd data and data quality rules for the critical data elements .
  • Created high level data quality operative modern design and road map.
  • Profiled the current state and established base line metrices on the exisitingf data sets and implementated data quality rules.
  • Managed security groups on AWS, focusing on high-availability, fault-tolerance, and auto scaling using Terraform templates. Along with Continuous Integration and Continuous Deployment with AWS Lambda and AWS code pipeline.
  • Worked on analysis tool like Tableau for regression analysis, pie charts, and bar graphs.

Environment: Cloudera Manager (CDH5), Hadoop, Hive, S3, Kafka, Pyspark, HDFS, NiFi, Pig, Scrum, Git, Sqoop, Oozie. Pyspark, Informatica, Tableau, Snowflake, OLTP, OLAP, SQL Server, Python, Shell Scripting, HBase, Cassandra, Informatica, XML,Unix.

Confidential, Dania Beach - FL

Big Data Engineer

Responsibilities:

  • Created and maintained SQL Server scheduled jobs, executing stored procedures for the purpose of extracting data from Oracle into SQL Server. Extensively used Tableau for customer marketing data visualization
  • Used SQL Server Integrations Services (SSIS) for extraction, transformation, and loading data into target system from multiple sources
  • Wrote production level Machine Learning classification models and ensemble classification models from scratch using Python and PySpark to predict binary values for certain attributes in certain time frame.
  • Built real time pipeline for streaming data usingKafkaandSparkStreaming.
  • Involved inUnit Testingthe code and provided the feedback to the developers. PerformedUnit Testingof the application by usingNUnit.
  • Designed both 3NF data models for OLTP systems and dimensional data models using star and snowflake Schemas.
  • Experience in Google Cloud components, Google container builders and GCP client libraries and cloud SDK’s
  • Migration of on premise data (Oracle/ SQL Server/ DB2/ MongoDB) to Azure Data Lake and Stored (ADLS) using Azure Data Factory (ADF V1/V2).
  • Responsible for wide-ranging data ingestion using Sqoop and HDFS commands. Accumulate ‘partitioned’ data in various storage formats like text, Json, Parquet, etc.
  • Architect & implement medium to large scale BI solutions on Azure using Azure Data Platform services (Azure Data Lake, Data Factory, Data Lake Analytics, Stream Analytics, Azure SQL DW, HDInsight/Databricks, NoSQLDB).
  • Writing UNIX shell scripts to automate the jobs and scheduling cron jobs for job automation using commands with Crontab.
  • Transforming business problems into Big Data solutions and define Big Data strategy and Roadmap. Installing, configuring, and maintaining Data Pipelines
  • Developed thefeatures,scenarios,step definitionsforBDD (Behavior Driven Development)andTDD (Test Driven Development)usingCucumber, Gherkinandruby.
  • Designing the business requirement collection approach based on the project scope and SDLC methodology.
  • Creating Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform, and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.
  • Implemented Kafka producer and consumer applications on Kafka cluster setup with help of Zookeeper.
  • Used Spring Kafka API calls to process the messages smoothly on Kafka Cluster setup.
  • Involved in all the steps and scope of the project reference data approach to MDM, have created a Data Dictionary and Mapping from Sources to the Target in MDM Data Model.
  • Data warehouse experience in Star Schema, Snowflake Schema, Slowly Changing Dimensions (SCD) techniques etc.
  • Developed various Mappings with the collection of all Sources, Targets, and Transformations using Informatica Designer
  • Responsible for designing and developing data ingestion from Kroger using Apache NiFi/Kafka.
  • Used ApacheSpark Data frames, Spark-SQL, Spark MLLibextensively and developing and designing POC's using Scala, Spark SQL and MLlib libraries.
  • Data Integrationingests, transforms, and integrates structured data and delivers data to a scalable data warehouse platform using traditional ETL (Extract, Transform, Load) tools and methodologies to collect of data from various sources into a single data warehouse.
  • Monitored cluster health by Setting up alerts using Nagios and Ganglia
  • Experience managing Azure Data Lakes (ADLS) and Data Lake Analytics and an understanding of how to integrate with other Azure Services. Knowledge of USQL
  • Performed all necessary day-to-day GIT support for different projects, Responsible for design and maintenance of the GIT Repositories, and the access control strategies.

Environment: Spark-Streaming, Hive, Scala, Hadoop, Kafka, Spark, Sqoop, Docker, Spark SQL, TDD, pig, NoSQL, Impala, Oozie, Hbase, Data Lake, Zookeeper, Snowflake, Azure, Unix/Linux Shell Scripting,Python, PyCharm, Informatica, Informatica PowerCenter Linux,, Shell Scripting, Git

Confidential

Data Engineer

Responsibilities:

  • Created various complex SSIS/ETL packages to Extract, Transform and Load data
  • Used OozieScheduler system to automate the pipeline workflow and orchestrate the map reduces jobs that extract the data on a timely manner
  • Integrated Kafka with Spark Streaming for real time data processing
  • Managed security groups on AWS, focusing on high-availability, fault-tolerance, and auto scaling using Terraform templates. Along with Continuous Integration and Continuous Deployment with AWS Lambda and AWS code pipeline.
  • Worked publishing interactive data visualizations dashboards, reports /workbooks on Tableau and SAS Visual Analytics.
  • Developed SSRS reports, SSIS packages to Extract, Transform and Load data from various source systems
  • Defined facts, dimensions and designed the data marts using the Ralph Kimball's Dimensional Data Mart modeling methodology using Erwin
  • Used Hive SQL, Presto SQL and Spark SQL for ETL jobs and using the right technology for the job to get done.
  • Defined and deployed monitoring, metrics, and logging systems on AWS.
  • Developed code to handle exceptions and push the code into the exception Kafka topic.
  • Was responsible for ETL and data validation using SQL Server Integration Services.
  • Implementing and Managing ETL solutions and automating operational processes.
  • Optimizing and tuning the Redshift environment, enabling queries to perform up to 100x faster for Tableau and SAS Visual Analytics.
  • Developed Python scripts to automate data sampling process. Ensured the data integrity by checking for completeness, duplication, accuracy, and consistency
  • Applied various machine learning algorithms and statistical modeling like decision tree, logistic regression, Gradient Boosting Machine to build predictive model using scikit-learn package in Python
  • Wrote various data normalization jobs for new data ingested into Redshift.
  • Created ad hoc queries and reports to support business decisions SQL Server Reporting Services (SSRS).
  • Analyze the existing application programs and tune SQL queries using execution plan, query analyzer, SQL Profiler and database engine tuning advisor to enhance performance.
  • Created Entity Relationship Diagrams (ERD), Functional diagrams, Data flow diagrams and enforced referential integrity constraints and created logical and physical models using Erwin
  • Implemented Work Load Management (WML) in Redshift to prioritize basic dashboard queries over more complex longer-running adhocqueries. This allowed for a more reliable and faster reporting interface, giving sub-second query response for basic queries.

Environment: SQL Server, Erwin, Kafka, Python, MapReduce, Oracle, AWS, Redshift, Informatica RDS, NOSQL, MySQL, PostgreSQL.

We'd love your feedback!