Data Engineer Resume
Charlotte, NC
SUMMARY
- 7+ years of experience in implementing various Big Data/ Cloud Engineering, Snowflake, Data Warehouse, Data Modelling, Data Mart, Data Visualization, Reporting, Data Quality, Data virtualization and Data Science Solutions.
- Experience in Data transformation, Data mapping from source to target database schema, Data Cleansing procedures
- Deep knowledge and strong deployment experience in Hadoop and Big Data ecosystems - HDFS, MapReduce, Spark, Pig, Sqoop, Hive, Oozie, Kafka, zookeeper, and HBase.
- Expertise in working with Hive data warehouse infrastructure-creating tables, data distribution by implementing Partitioning and Bucketing, developing and tuning the HQL queries.
- Strong experience in tuning Spark applications and Hive scripts to achieve optimal performance.
- Developed Spark Applications using Spark RDD, Spark-SQL and Data frame APIs.
- Strong experience building end-to-end data pipelines on the Hadoop platform.
- Developed Simple to complex MapReduce streaming jobs using Python language that are implemented using Hive and Pig.
- Capable of processing large sets of structured, semi-structured, and unstructured data and supporting systems application architecture.
- Expertise in OLTP/OLAP System Study, Analysis and E-R modelling, developing Database Schemas like Star schema and Snowflake schema used in relational, dimensional and multidimensional modelling
- Experience in creating separate virtual data warehouses with difference size classes in AWS Snowflake
- Hands-on experience in bulkloading & unloadingdata into Snowflake tables using COPY command
- Experience with data transformations utilizing SnowSQL in Snowflake
- Developed Spark Applications that can handle data from various RDBMS(MySQL, Oracle Database) and Streaming sources.
- Solid understanding of AWS, Redshift, S3, EC2 and Apache Spark, Scala process, and concepts, configuring the servers for auto scaling and elastic load balancing
- Hands on experience in machine learning, big data, data visualization, R and Python development, Linux, SQL, GIT/GitHub
- Experienced in python data manipulation for loading and extraction as well as with python libraries such as NumPy, SciPy and Pandas for data analysis and numerical computations
- Extensive working experience with Python including Scikit-learn, SciPy, Pandas, and NumPy developing machine learning models, manipulating and handling data
- Extensive experience in Text Analytics, developing different Statistical Machine Learning, Data Mining solutions to various business problems and generating data visualizations using R, Python
- Experience in extracting, transforming and loading (ETL) data from spreadsheets, database tables and other sources using Microsoft SSIS and Informatica.
- Developed mapping spreadsheets for (ETL) team with source to target data mapping with physical naming standards, data types, volumetric, domain definitions, and corporate meta-data definitions.
- Performeddatavisualization and Designed dashboards with Tableau, and generated complex reports, including charts, summaries, and graphs to interpret findings to team and stakeholders
- Developed Snow pipes for continuous injection of data using event handler from AWS (S3 bucket)
- Design and developed end-to-end ETL process from various source systems to Staging area, from staging to Data Marts and data load
- Strong Understanding of dimensional and relational modelling techniques. Well versed normalization/ Denormalization techniques for optimum performance in relational and dimensional DB.
TECHNICAL SKILLS
Operating Systems: Unix, Linux(Ubuntu, CentOS), Mac OS, OpenSUSE, Windows 2003/2008/2012/ XP/7/8/9X/NT/Vista
Hadoop Ecosystem/Distributions: HDFS, MapReduce, Yarn, Oozie, Zookeeper, Job Tracker, Task Tracker, Name Node, Data Node, Cloudera, Horton works
Big Data Ecosystem: Hadoop, Spark, MapReduce, YARN, Hive, SparkSQL, Impala, Pig, Sqoop, HBase, Flume, Oozie, Zookeeper, Avro, Parquet, Maven, Snappy, Hue
Data Ingestion: Sqoop, Flume, NiFi, Kafka
Cloud Computing Tools: Snowflake, SnowSQL, AWS, Databricks, GCP, Azure data lake services, Amazon EC2
NoSQL Databases: HBase, Cassandra, MongoDB, CouchDB, Apache, Hadoop HBase
Programming Languages: Python( Jupyter Notebook, PyCharm IDE), R, Java, C, Scala, SQL, PL/SQL, VBScript and Shell Scripts, XML, HTML, Visual Basic 6.0, Fox Pro, SAS
Frameworks: MVC, Struts, Spring, Hibernate
Web Technologies: HTML 5, CSS 3, XML, JavaScript, Maven Spring 4, Spring MVC, JSP, Angular JS, Ajax, jQuery, XSP, WSDL, JSON
Scripting Languages: Bash, Pearl, Python, R Language
Databases: Snowflake Cloud DB, Oracle, MySQL, Teradata 12/14, DB2 10.5, MS Access, SQL Server 2000/2005/2008/2012, PostgreSQL 9.3, Sybase ASE 11.9.2, Netezza, AmazonRDS
SQL Server Tools: SQL Server Management studio, Enterprise Manager, Query Analyzer, Profiler, Export and Import ( DTS )
IDE: IntelliJ, Eclipse, Visual Studio, IDLE
Web Services: Restful, SOAP, O9iAS, Oracle Form Server, Weblogic 8.1/10.3, Web Sphere MQ 6.0
Packages and Tools: MS-Office, TOAD, SQL Developer & Navigator, Share point portal server, Visual Source Safe, SVN, TFS, BTEQ
Methodologies: Agile, Scrum, Iterative Development, Waterfall Model, UML, Design Patterns, UML
ETL/ Data: Tensor flow, Data API, PySpark, Pervasive Cosmos Business Integrator ( Data Junction Tool), CTRL-M, Data Stage, Informatica Power Center 9.6.1/9.5/8.6.1/8.1/7.1, Talend, Pentaho, Microsoft SSIS, Data Stage 7.5, Ab Initio
OLAP Tools: MS Analysis Services, Business Objects & Crystal Reports 9, MS SQL Analysis Manager, DB2 OLAP, Cognos Powerplay
Warehousing and Modelling/ Architect Tools: Erwin 7.3&9.5 ( Dimensional Data Modelling, Relational DM, Star Schema, Snow-flake, Fact and Dimensional Tables, Physical and Logical DM, Canonical Modelling), Visio 6.0, ER/Studio, Rational System Architect, IBM Infosphere DA, MS Visio Professional, DTM, DTS 2000, SSIS, SSAS
Reporting / BI Tools: MS Excel, Tableau, Tableau Server and Reader, Power BI, QlikView, SAP Business objects, Crystal Reports, SSRS, Splunk
Utilities/ Tools: Bugzilla, QuickTestPro 9.2, Selenium, Quality Center, Test Link, TWS, Documentum, Tortoise SVN, Putty, WIN SCP, Log4J, Junit,, GIT, Jasper Reports, Jenkins, Eclipse, TOMCAT, NetBeans, SVN, SOAPUI
PROFESSIONAL EXPERIENCE
Data Engineer
Confidential - Charlotte, NC
Responsibilities:
- Created Hive tables for loading and analysing data.
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and extracted data from MYSQL into HDFS vice-versa using Sqoop.
- Used Spark API over Cloudera Hadoop Yarn to perform analytics on data in Hive.
- Designed and implemented an ETL framework using Scala and Python to load data from multiple sources into Hive and from Hive to Vertica
- Used HBase on top of HDFS as a non-relational database.
- Loaded the data into Spark RDD, perform advanced procedures like text analytics and processing using in-memory data Computations capabilities of Spark using Scala to generate the Output response.
- Developed Scala scripts using both Data frames/SQL and RDD/MapReduce in Spark for Data Aggregation, queries and writing data back into the OLTP system through Sqoop.
- Handled large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and others during the ingestion process itself.
- Implemented Partitions, Buckets, and developed Hive query to process the data and generate the data cubes for visualizing.
- Optimizing existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames and Pair RDD.
- Used Spark Streaming APIs to perform necessary transformations and actions on the fly for building the common learner data model which gets the data from Kafka in near real-time and Persists into Cassandra.
- Extracted Fingerprint image Data stored on local network to Conduct Exploratory Data analysis (EDA), Cleaning and organize. Ran NFIQ algorithm to ensure data quality by collecting the high score images. Finally Created histograms to compare distributions of different datasets.
- Transformed the image dataset to protocol buffers, serialized and finally stored inside TFrecord data format.
- Loaded the data in GPU and achieved Half Precision FP16 training on Nvidia Titan RTX and Titan V GPU for TensorFlow 1.14.
- Setup alertingand monitoring using Stackdriverin GCP
- Optimized TFRecord data ingestion pipeline using tf.Data API and made them scalable by streaming over network, thus enabling training of models with Datasets which were bigger than CPU memory.
- Worked extensively on AWS Components such as Elastic Map Reduce (EMR)
- Loaded data using AWS Glue
- Used Athena for data analytics.
- Worked with the data science team in automating and productionalizing various models like logistic regression, k-means using Spark MLlib.
- Created various reports using Tableau based on requirements with the BI team.
- Used DataStax Spark-Cassandra connector to load data into Cassandra and used CQL to analyze data from Cassandra tables for quick searching, sorting, and grouping.
- Used Spark API over Cloudera Hadoop Yarn to perform analytics on data in Hive
Environment: Hadoop Yarn, Spark Core, Spark Streaming, Spark SQL, Spark MLlib, HBase, Scala, Python, Kafka, Hive, Sqoop, Amazon AWS, Athena, Cassandra, Tableau, Cloudera, MySQL, Linux, MapReduce.
Big Data Engineer
Confidential - Seattle, WA
Responsibilities:
- Extensive hands on experience with Big Data Engineer Stack including HDFS, MapReduce, Sqoop, Hive, Pig, HBase, Oozie, Flume, Kafka, Zookeeper, and spark
- Experience with NoSQL Databases like HBase as well as other eco systems like Zookeeper, Oozie, Impala, Strom, Spark-Streaming/SQL, Kafka, Flume
- Developed Hive, Bash scripts for source data validation and transformation.
- Experience in Converting Hive/SQL Queries into Spark transformations using Java and experience in ETL development using Kafka, flume and Sqoop.
- Designed and implemented an ETL framework to load data from multiple sources into Hive and from Hive into Teradata.
- Developed Spark Code and Spark-SQL/Streaming for faster testing and processing of data.
- Good experience in Hive Data Warehousing concepts like Static/ Dynamic Partitioning, Bucketing, Managed, and External tables, Join operations on tables.
- Worked on loading CSV/TXT/AVRO/PARQUET files using Scala/ Java language in Spark Framework and Process the data by creating Spark Data Frame and RDD and save the file in Parquet format in HDFS.
- Used Spark-SQL to Load JSON data and create SchemaRDD and loaded it into Hive Tables and handled Structured data using Spark SQL.
- Load D-Stream data into Spark RDD and do in memory data Computation to generate Output response.
- Well Versed with Major Hadoop distributions, Cloudera and Horton Works. Having experience on Eclipse, NetBeans IDEs.
- Used Python and R scripting by implementing machine algorithms to predict data and forecast data for better results
- Experience in developing packages in R studio with a shiny interface
- Developed spark jobs to process all the information and specify the passion points and email promotion for each user. Used Sqoop, FTP, APIs, SQS, On S3 copy to pull data to HDFS.
- Wrote Data Pipeline that fetches Adobe Omniture data which is routed to S3 using SQS every hour.
- Generated and injected Vehicle IOT data to AWS IOT platform using python
- Implemented AWS lambda architectural model for handling end-to-end real time and batch analytic loads.
- Published IOT Data to Kafka Stream, Consumed by a spark module to perform predictive analysis using “ Random Forest “ Machine learning algorithm.
- Utilized Redshift to storing the processed records and implemented batch scripts for continuous learning.
- Built Elastic search cluster in integration with Kibana for publishing real time dashboards for maintenance data.
- Handled real-time data using Kafka.
- Transferred all the data from history jobs to the Azure with HDInsight installed
- Involved in designing and deploying multi-tier applications using all the AWS services like (EC2, Route53, S3, RDS, Dynamo DB, SNS, SQS, IAM) focusing on high-availability, fault tolerance, and auto-scaling in AWS Cloud Formation
Environment: HDFS, Azure, Teradata, Scala, MapReduce, Spark, Hive, Scala, NiFi, MySQL, Oozie, Kafka, Shell Scripting, Cloudera, MongoDB, JSON, Avro, XML, Parquet.
Big Data Developer
Confidential - Murray, UT
Responsibilities:
- Design the Hive table Structure and collect/ load the necessary data required for the development, Understanding the business requirements and build necessary workflows in PySpark.
- Working on data Pre-processing and feature engineer to enable strategy teams for developing the ML models smoothly.
- Good Experience in working with different file formats such as RC, ORC, Parquet, Avro, Sequence, Text File, and CSV.
- Involved in Building scalable distributed data lake system for confidential real time and batch analytical needs
- Involved in designing, reviewing, optimizing data transformation processes using apache storm.
- Experience in Job Management using Fair Scheduling and developed job processing scripts using Control-M workflow.
- Exporting of result set from HIVE to MySQL using Sqoop export tool for further processing.
- Automation of all the jobs starting from pulling the Data from different Data Sources like MySQL and pushing the result dataset to Hadoop Distributed File System and running MR, PIG, and Hive jobs using Kettle and Oozie (Work Flow management)
- Written Storm topology to emit data into Cassandra DB.
- Written Storm topology to accept data from Kafka producer and process the data
- Expereinced in handling large data sets using partitions, Spark in memory capabilities, Broadcasts in spark, effective efficient joins, transformation and other during ingestion process itself.
- Worked on a POC to Compare processing time for impala with apache hive for batch applications to implement the former in project.
- Worked extensively on AWS Components such as Elastic Map Reduce (EMR)
- Created data pipeline for different events of ingestion, aggregation and load consumer response data in AWS S3 bucket into Hive external tables in HDFS location to serve as feed for Tableau dashboards.
- Built models using Python and Pyspark to predict probability of attendance for various campaigns and events
- Used DataStax Spark-Cassandra connector to load data into Cassandra and used CQL to analyze data from Cassandra tables for quick searching, sorting, and grouping.
- Used Spark API over Cloudera Hadoop Yarn to perform analytics on data in Hive
Environment: Hadoop, MapReduce, Spark, Pig, Hive, Sqoop, Oozie, HBase, Zookeeper, Kafka, Spark streaming, Flume, S Storm, Impala, Cassandra,, MySQL, Windows, Unix.
Data Modeler/ Analyst
Confidential
Responsibilities:
- Experienced in Data modeler performing business area analysis and logical and physical data modelling using Erwin and data warehouse/ data mart applications as well as operational applications enhancements and new development. Data warehouse/ data marts design was implemented using Ralph Kimball methodology.
- Highly maintained the stage and production conceptual, logical, and physical data models along with related documentation for a large data warehouse project. This included confirming migration of data models from oracle designer to Erwin and updating the data models to correspond to the existing database structures.
- Created a dimensional logical model with approximately 10 facts, 30 dimensions with 500 attributes using Erwin.
- Excellent SQL programming skills and developed Stored Procedures, Triggers, Functions, Packages using SQL/PL SQL. Perfomance tuning and query optimization techniques in transactional and data warehouse environemnts.
- Involved with DBA group to create Best-Fit Physical Data Model from the logical Data model using Forward engineering using Erwin tool.
- Enforced referential integrity in the OLTP data model for consistent relationship between tables and efficient database design.
- Conducted design walk through sessions with business intelligence team to ensure that reporting requirements are met for the business.
- Developed Data Mapping, Data Governance, and Transformation and cleansing rules for the Master Data Management Architecture involving OLTP, ODS.
- Served as a member of a development team to provide business data requirements analysis services, producing logical and physical data models using Erwin.
- Extensively used ETL methodology for supporting data extraction, transformations and loading processing, in a complex EDW using informatica.
- Written Complex SQL queries for validating the data against different kinds of reports generated by business objects XIR2.
- Performing data management projects and fulfilling ad-hoc requests according to user specifications by utilizing data management software programs and tools like Perl, Toad, MS Access, Excel and SQL.
- Wrote PL/SQL statement and stored procedures in Oracle for extracting as well as writing data.
- Worked very close with Data Architectures and DBA team to implement data model changes in database in all environments.
- In depth analyses of data report was prepared weekly, biweekly, monthly using MS Excel, SQL&UNIX.
Environment: Erwin 9/8/r7.0, Rational Requisite Pro, Data Stage, DB2 UDB, SQL Server, Sybase, Windows NT, Linux, Sun Solaris, T-SQL, ER Studio, SQL, SQL Server 2000/2005, Rational Rose, Crystal Reports 9, Windows XP, Oracle 8i, Windows XP/NT/2000, SQL, PL/SQL Developer 5.1, Sun Solaris8.0, Erwin, ER-Studio.
