Sr. Data Engineer Resume
Parlin, NJ
SUMMARY
- Experience as Software developer with strong emphasis on Data Engineering and Data Analysis using Big Data tools and application development using HADOOP framework and related technologies such as HDFS, MapReduce, HIVE, PIG, STORM, YARN, OOZIE, SQOOP, Airflow and Zookeeper and includes working experience in Spark Core, Spark SQL, Spark Streaming, Scala, Python and Kafka.
- Hands - on experience in designing and implementing data engineering pipelines and analyzing data using AWS stack like AWS EMR, AWS Glue, EC2, AWS Lambda, Athena, Redshift, Scoop and Hive.
- Hands-on experience in programming using Python, Scala, Java and SQL.
- Experience gathering customer requirements, writing test cases, and partnering with developers to ensure full understanding of internal/external customer needs.
- Experienced in developing production ready spark applications using Spark RDD APIs, Data frames, Spark-SQL and Spark-Streaming API's.
- Expertise in writing end to end Data processing Jobs to analyze data using MapReduce, Spark, and Hive.
- Experience with Apache Spark ecosystem using Spark-Core, SQL, Data Frames, RDD's, and knowledge on Spark MLLib.
- Experienced in data manipulation using python for loading and extraction as well as with python libraries such as NumPy, SciPy and Pandas for data analysis and numerical computations.
- Experience on Migrating SQL database to Azure data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and controlling and granting database access and Migrating On premise databases to Azure Data Lake store using Azure Data factory.
- Design and develop Spark applications using Pyspark and Spark-SQL for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage pattern.
- Experience in Work on AWS Databases like Elastic Cache (Memcached & Redis) and NoSQL databases Cassandra & MongoDB for database performance tuning & data modeling.
- Experience in Azure Cloud Services (PaaS & IaaS), Azure Synapse Analytics, SQL Azure, Data Factory, Azure Analysis services, Application Insights, Azure Monitoring, Key Vault, Azure Data Lake.
- Extract, Transform and Load (ETL) source data into respective target tables to build Data Marts.
- Conducted Gap Analysis, created Use Cases, workflows, screen shots and Power Point presentations for various Data Applications.
- Active involvement in all scrum ceremonies - Sprint Planning, Daily Scrum, Sprint Review and Retrospective meetings and assisted Product owner in creating and prioritizing user stories.
- Involved in best practices forCassandra, migrating applications toCassandradatabase from the legacy platform for Choice.
- Hands on experience in Apache Spark creating RDDs and Data Frames, applying Transformations and Actions and converting RDDs to Data Frames.
- Experienced in developing MapReduce programs using Apache Hadoop for Big Data workloads.
- Good understanding of XML methodologies (XML, XSL, XSD) including SOAP.
- Used Spark CassandraConnector to load data to and fromCassandra and analyze the data with Apache Spark.
- Experienced process-oriented Data Analyst having excellent analytical, quantitative, and problem-solving skills using SQL, MicroStrategy, Advanced Excel, Python.
- Proficient in writing unit testing code using Unit Test/PyTest and integrating the test code with the build process.
- Used Pythonscripts to parse XML and JSON reports and load the information in database.
- Experienced with version control systems like Git, GitHub, Bitbucket to keep the versions and configurations of the code organized.
TECHNICAL SKILLS
Hadoop Distributions: AWS EMR and Azure Data Factory.
Languages: Scala, Python, Py Spark, Python, Hive QL, PL/SQL, JSON, XML.
Cloud platform: AWS: Amazon EC2, S3, EBS, RDS, SNS, Athena, Glue, Lambda, EMR, Redshift, DynamoDB Azure: Azure Cloud Services (PaaS & IaaS), Azure Synapse Analytics, SQL Azure, Data Factory, Azure Analysis services, Application Insights, Azure Monitoring, Key Vault, Azure Data Lake, Azure HDInsight.
Reporting and ETL Tools: Tableau, Power BI, Apache Nifi, Druid
CI/CD Tools: Jenkins
Designing Tools: UML, Visio
IDEs: Eclipse, NetBeans
Java Technologies: JSP, JDBC, Servlets, Junit
Web Technologies: XML, HTML, JavaScript, jQuery, JSON
Databases: Oracle, SQL Server, Teradata, Cassandra, Mongo DB
Big Data Technologies: Hadoop, HDFS, Hive, Pig, Oozie, Sqoop, Spark, Snowflake, Machine Learning, Pandas, NumPy, Seaborn, Impala, Zookeeper, Flume, Airflow, Data Bricks, Kafka
Ticketing Tools: JIRA, Service Now, Confluence.
Version Control: Git, Bitbucket
Operating Systems: UNIX, LINUX, Ubuntu, Windows.
Scheduling Tools: Control-M, Active Batch, Zena.
SDLC Methodologies: Agile, Scrum, Waterfall
Others: Putty, WinSCP, Data Lake
PROFESSIONAL EXPERIENCE
Confidential - Parlin, NJ
Sr. Data Engineer
Responsibilities:
- Worked closely with multiple teams to gather requirements and maintain relationships with those that are heavy users of data for analytics. Used AWS Redshift, S3, Spectrum and Athena services to query large amounts of data stored on S3 to create a Virtual Data Lake without having to go through the ETL process.
- Designed and implemented big data ingestion pipelines to ingest multi-TB data from various data source using Kafka, Spark streaming including data quality checks, transformation, and stored as efficient storage formats Performing data wrangling for a variety of downstream purposes such as analytics using PySpark.
- Worked on importing metadata into Hive using Python and migrated existing tables and the data pipeline from Legacy to AWS cloud (S3) environment and wrote Lambda functions to run the data pipeline in the cloud.
- Utilized Spark Scala API to implement batch processing of jobs
- Converted Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
- Developed Spark scripts using Python on AWS EMR for Data Aggregation, Validation and Ad Hoc querying.
- Experience working for EMR cluster in AWS cloud and working with S3.
- Used broadcast variables in spark, effective & efficient Joins, transformations and other capabilities for data processing. Utilized Spark in Memory capabilities, to handle large datasets.
- Converted Hive/SQL queries into Spark Transformations using Spark RDDs and Scala and involved in using SQOOP for importing and exporting data between RDBMS and HDFS.
- Optimizing and tuning the Redshift environment, enabling queries to perform up to 100 xs faster for Tableau and SAS Visual Analytics.
- Experience building reusable ETL components using Postgres and snowflake.
- Worked extensively on writing triggering Snowpipe, Snowflake data loads automatically using Amazon SQS (Simple Queue Service) notifications for an S3 bucket.
- Connected to Amazon Redshift through Tableau to extract live data for real time analysis.
- UsedSpark-SQLto LoadJSONdata and createSchema RDDand loaded it intoHiveTables and handled Structured data usingSparkSQL.
- Scheduled Airflow DAGs to run multiple Hive and Pig jobs, which independently run with time and data availability and Performed Exploratory Data Analysis and Data Visualizations using Python, and Tableau.
Environment: Hadoop/Bigdata Ecosystem (Spark, Kafka, Hive, HDFS, Sqoop, Oozie, Cassandra, MongoDB), AWS (S3, AWS Glue, Redshift, RDS, Lambda, Athena, SNS, SQS), Oracle, snowflake, Docker, Git, SQL Server, Python 3.x, Pyspark, Teradata, Tableau, Quick sight, Data warehousing.
Confidential - Foster city, CA
Sr.Data Engineer
Responsibilities:
- Worked closely with stake holders to understand business requirements to design quality technical solutions that align with business and IT strategies and comply with the organization's architectural standards.
- Developed multiple applications required for transforming data across multiple layers of Enterprise Analytics Platform and implement Big Data solutions to support distributed processing using Big Data technologies.
- Responsible for data identification and extraction using third-party ETL and data-transformation tools or scripts. (e.g., SQL, Python)
- Worked on migration of data from On-prem SQL server to Cloud databases (Azure Synapse Analytics (DW) & Azure SQL DB).
- Developed and managed Azure Data Factory pipelines that extracted data from various data sources, transformed it according to business rules, using python scripts that utilized Pyspark and consumed APIs to move data into an Azure SQL database.
- Created a new data quality check framework project in Python that utilized pandas.
- Implemented source control and development environments for Azure Data Factory pipelines utilizing Azure Repos.
- Created Hive/Spark external tables for each source table in the Data Lake and written Hive SQL and Spark SQL to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Designed and developed ETL & ETL frameworks using Azure Data Factory and Azure Data Bricks.
- Flattening and transforming huge amounts of nested data in parquet and delta forms using Spark SQL and the newest join optimization methods, then loading them into Hive, Delta Lake, and Snowflake tables.
- Created generic data bricks NOTEBOOKs for performing data cleansing.
- Created Azure Data factory pipelines to refactor on-prem SSIS packages into Data factory pipelines.
- Working with Azure BLOB and Data Lake storage for loading data into Azure SQL Synapse (DW).
- Ingested and transformed source data using Azure Data flows and Azure HDInsight.
- Loaded the data into the patient analytics database on the Snowflake platform.
- Created Azure Functions to ingest data at regular intervals.
- Created Data Bricks notebooks for performing complex transformations and integrated them as activities in ADF pipelines.
- Written complex SQL queries for data analysis and extraction of data in required format.
- Created Power BI DataMart’s and reports for various stakeholders in the business.
- Created CI/CD pipelines using Azure DevOps.
- Enhanced the functionality of existing ADF pipeline by adding new logic to transform the data.
- Worked on Spark jobs for data preprocessing, validation, normalization, and transmission.
- Optimized code and configurations for performance tuning of Spark jobs.
- Worked with unstructured and semi structured data sets to aggregate and build analytics on the data.
- Work independently with business stakeholders with strong emphasis on influencing and collaboration.
- Daily participation in Agile based Scrum team with tight deadlines.
Environment: Azure Synapse Analytics, Azure Data Factory, Azure Data bricks, Azure Synapse Studio, Snowflake, Hadoop, SQL Server, Power BI, Oracle 12c/11g, SQL scripting, PL/SQL, Python, Unix Shell, Jira, Confluence.
Confidential - Framingham, MA
Sr. Big Data Engineer/Data Analyst
Responsibilities:
- Involved in writing Spark applications using Python to perform various data cleansing, validation, transformation, and summarization activities according to the requirement.
- Developed multiple POCs using Pyspark and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL/Teradata and developed code in reading multiple data formats on HDFS using Pyspark.
- Worked on AWS Cloud to convert all on premise, existing processes and databases to AWS Cloud.
- Design and Develop ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.
- Used AWS Redshift, S3, Spectrum and Athena services to query large amounts of data stored on S3 to create a Virtual Data Lake without having to go through the ETL process.
- Developed a pyspark job to load the CSV files into the S3 buckets and createdAWS S3 Buckets, performed folder management in each bucket, managed logs and objects within each bucket.
- Involved in file movements between HDFS and AWS S3 and extensively worked with S3 bucket in AWS.
- Developed a daily process to do incremental import of data from DB2 and Teradata into Hive tables using Sqoop.
- Worked on importing metadata into Hive using Python and migrated existing tables and the data pipeline from Legacy to AWS cloud (S3) environment and wrote Lambda functions to run the data pipeline in the cloud.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Extensively worked with Partitions, Dynamic Partitioning, bucketing tables in Hive, designed both Managed and External tables, also worked on optimization of Hive queries.
- Designed, developed and created ETL(Extract, Transformand Load)packagesusing Python to load data into Data warehouse tools (Teradata) from databases such as Oracle SQL Developer, MS SQL Server.
- Utilized inbuilt Python module JSON to parse the member data which is in JSON format using json. loads or json.dumps and loads into a database for reporting.
- Used Pandas API to put the data as time series and tabular format for central timestamp data manipulation and retrieval during various loads in the DataMart.
- Worked on bash scripting to automate the Python jobs for day-to-day administration.
- Performed data extraction and manipulation over large relational datasets using SQL,Python, and other analytical tools.
- Extensively worked with Teradata utilities like BTEQ, Fast Export, Fast Load, Multi Load to export and load Claims & Callers data to/from different source systems including flat files.
Environment: Spark, Python, Hadoop, Hive, S3, RDS, EMR, EC2, SNS, Lambda, Athena, Step Functions, Jenkins, Foundry, Git.
Confidential, Fort Lauderdale, FL
Hadoop Developer
Responsibilities:
- Experience in supporting and managing Hadoop Clusters using Hortonworks distributions by deploying it on AWS cloud.
- Collected aggregated large amount of web log data from different sources such as web servers, mobile and network devices using Apache Kafka.
- Ingestion framework was developed in python Big Data technologies with data stores such as DynamoDB, Cassandra.
- Creating the RDD’s, Data frames for faster execution and performing data transformations and actions using Spark.
- Developed optimal strategies for distributing the web log data over the cluster.
- Implemented Hive Generic UDF's to in corporate business logic into Hive Queries.
- Configuring Spark Streaming to receive real time data from the Kafka for high speed data processing and Store the stream data to HDFS.
- Used Scala to read text data, CSV data, image data from HDFS, S3 and Hive
- Worked on Spark SQL for faster execution of Hive queries using Spark SQL Context.
- Implemented complex big data with a focus on collecting, parsing, managing, analyzing, and visualizing large sets of data to turn information into business insights using multiple platforms in the Hadoop ecosystem.
- The developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
- Involved in source system analysis, data analysis, and data modeling to ETL (Extract, Transform and Load).
- Written Spark programs to model data for extraction, transformation, and aggregation from multiple file-formats including XML, JSON, CSV& other compressed file formats.
- Imported data from the structured data source into HDFS using Sqoop incremental imports.
- Created Hive tables, partitions and implemented incremental imports to perform ad-hoc queries on structured data.
- Build Hive tables using list partitioning and hash partitioning and created Hive Generic UDF's to process business logic with HiveQL.
- Developed SQL scripts using Spark for handling different data sets and verifying the performance over Map Reduce jobs.
- Supported MapReduce Programs that are running on the cluster and Wrote MapReduce jobs using JavaAPI.
- Designed unit test Data models and applications for data analytics solutions on streaming data
Environment: Hortonworks, HDFS, Hive, Sqoop, Oozie, Storm, Scala 2.11.8, Spark 2.0, Spark SQL, Spark streaming, Python, Kafka, GitHub, Kerberos, AWS, Amazon S3, Amazon EC2, Amazon EBS, Tableau.
Confidential
BI Developer / Data Analyst
Responsibilities:
- As a Data Visualization Consultant supporting Risk Consulting Team in Model Documentation and Data Visualizations for the Data Science Team which develops Credit and Market Risk Predictive Models for the Banking clients.
- Worked as a Tableau Desktop Developer focusing on developing high-end visualizations driven by data coming in from various data sources including Flat files, SQL Server, and MS Excel.
- Developed various stories and dashboards using multiple data sources and by blending them on a single worksheet in Tableau Desktop version.
- Involved extensively in dashboards development such as creating Tableau Extracts, Tableau Connectors (Live and Extract), formatting and report operations (sorting, filtering, ranking, Top-N Analysis).
- Worked extensively with Advance analysis Actions, Calculations, Parameters, Background images, Maps, Trend Lines, Statistics, and Log Axes. Groups, hierarchies in Tableau.
- Designed, Deployed, and Integrated and Maintained MDM systems.
- Used advanced data visualization and representation techniques in Tableau to provide an easy-to-understand interface for end users to quickly identify key areas within their data.
- Trained end users in various parts of the world to effectively use Tableau.
- Create and enhance the standard of client communications like company reports & screen presentations in PowerPoint, various Reports in MS Excel/Tableau through corporate template as per specification.
- Engaged in the design and production of visual communication materials like charts, financial presentations based on PowerPoint/Think cell/Tableau for use of Confidential Leadership Team.
Environment: Tableau, MS SQL Server, MS Excel, Flat files, MS PowerPoint
Confidential
Data Analyst
Responsibilities:
- Participated in testing of procedures and Data utilizing, PL/SQL to ensure integrity and quality of Data in Data warehouse.
- Worked to ensure high levels of Data consistency between diverse source systems including flat files, XML and SQL Database.
- Developed and run ad-hoc Data queries from multiple database types to identify system of records, Data inconsistencies, and Data quality issues.
- Developed complex SQL statements to extract the Data and packaging/encrypting Data for delivery to customers.
- Provided business intelligence analysis to decision-makers using an interactive OLAP tool
- Created T/SQL statements (select, insert, update, delete) and stored procedures.
- Involved in defining the source to target Data mappings, business rules and Data definitions.
- Ensured the compliance of the extracts to the Data Quality Center initiatives
- Metrics reporting, Data mining and trends in helpdesk environment using Access
- Worked on SQL Server Integration Services (SSIS) to integrate and analyze data from multiple heterogeneous information sources.
- Built reports and report models using SSRS to enable end user report builder usage.
- Created Excel charts and pivot tables for the Ad-hoc Data pull.
- Created Column Store indexes on dimension and fact tables in the OLTP database to enhance read operation.
Environment: SQL, PL/SQL, T/SQL, XML, OLAP, SSIS, SSRS, Excel, ER win.
