We provide IT Staff Augmentation Services!

Sr. Data Engineer Resume

5.00/5 (Submit Your Rating)

San, JosE

SUMMARY

  • Over 7+ years of experience in Information Technology with emphasis on Business Requirements Analysis, Application Design, Development, Coding, testing, implementation, and maintenance of client/server Data Warehouse using Hadoop, HDFS, Horton works, MapReduce, and Hadoop Ecosystem.
  • Experience in the Big data platform having extensive firsthand experience in the Apache Hadoop ecosystem and enterprise application development. Good knowledge of extracting the models and trends from the raw data collaborating with the data science team.
  • Extensively worked on Hadoop ecosystem experience in ingestion, storage, querying, processing, and analysis of big data
  • Experience inproviding solutions for Big Datausing Hadoop, Spark, HDFS, Map Reduce, YARN, Kafka, Pig, Hive, Sqoop, HBase, Oozie, Zookeeper, Cloud era Manager, Horton works.
  • Hands - on experience with importing and exporting data from Relational databases to HDFS, Hive, and HBase using Sqoop.
  • Strong experience working withAmazon cloud serviceslike EMR, Redshift, DynamoDB, Lambda, Athena, Glue, S3, API Gateway, RDS, CloudWatch for efficient processing of Big Data.
  • Experienced in processing real-time data using Kafka 0.10.1 producers and stream processors and implemented stream process using Kinesis and data landed into data lake S3.
  • Good experience onAzure cloudcomponents likeHD Insight,Data bricks,Data Lake,Blob Storage,Data Factory,Storage Explorer,SQL DB,SQL DWH,Cosmos DB.
  • Experience in building data pipelines usingAzure Data Factory,Azure Databricks, and loading data toAzure Data Lake, Azure SQL Database,Azure SQL Data warehouse, and controlling and granting database access.
  • Proven expertise in employing techniques for Supervised and Unsupervised (Clustering, Classification, PCA, Decision trees, KNN, SVM) learning, Predictive Analytics, Optimization Methods, and Natural Language Processing (NLP), Time Series Analysis.
  • Experience in Machine Learning Regression Algorithms like Simple, Multiple, Polynomial, SVR (Support Vector Regression), Decision Tree Regression, Random Forest Regression.
  • UsedPandas, Numpy, Scipy, Scikit-learn, NLTKinPythonfor scientific computing and data analysis.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDD and PySpark concepts.
  • ConfiguredSpark StreamingwithKafkato clean and aggregate real-time data.
  • Designed and developed spark pipelines to ingest real-time event-based data from Kafka and other message queue systems and processed huge data with spark batch processing into data warehouse hive.
  • Good understanding ofSpark Architecture,MPP Architecture, includingSparkCore,SparkSQL,Data Frames,Spark Streaming, Driver Node, Worker Node, Stages, Executors, and Tasks.
  • Experience in the development ofT-SQL,Oracle PL/SQL Scripts, Stored Procedures, and Triggers for business logic implementation.
  • Experience with industry-standard methodologies likeWaterfall, Agile, andScrummethodology within the Software Development Life Cycle.
  • Extensively worked with Version Control Systems like Bitbucket, GIT, SVN, GitHub.

TECHNICAL SKILLS

Big Data Stack: Hadoop, Spark, MapReduce, Hive, Pig, YARN, Sqoop, Kafka, Impala, Storm

Cloud Services: S3, Kinesis, VPC, EC2, EMR, Redshift, Dynamo DB, RDS, IAM, Lambda, API Gateway, Athena, Cloud Watch, Azure HD Insight, Data bricks, Data Lake, Blob Storage, Azure Data Factory (ADF), SQL, DB, Azure Synapse, Cosmos DB.

Programming Languages: Python, Scala, PL/SQL, C, C++and JAVA

Relational Databases: Oracle, MySQL, SQL Server, Postgre SQL, Teradata, Snowflake.

Databases: Mongo DB, Cassandra, HBase

Machine Learning Algorithms: Linear & Logistic Regression, Decision Trees, Random Forest, K-Means Clustering, Support Vector Machines, Gradient Boost Machines, Neural Networks.

Libraries: Pandas, NumPy, Scikit-learn, Matplotlib, Tweepy, Seaborn, TensorFlow, Keras, MLlib, Boto3.

Data Analysis Skills: Data Cleaning, Data Visualization, Feature Selection, Pandas.

Web technologies: Django, Flask, HTML/CSS, JavaScript

BI Tools: Tableau, Tableau server & Reader, OBIEE, QlikView, SAP Business Intelligence, Amazon Redshift, or Azure Data Warehouse.

Operating Systems: Unix, Linux, Windows

Version Control Systems: Bitbucket, GIT, SVN, GitHub

PROFESSIONAL EXPERIENCE

Confidential, San Jose

Sr. Data Engineer

Responsibilities:

  • Written ETL jobs using spark data pipelines to process data from a different source to transform data to multiple targets.
  • Created streams using Spark and processed real-time data into RDDs & data frames and created analytics using PySpark SQL.
  • Developed incremental and complete load Python processes to ingest data into Elastic Search from oracle database.
  • Developed a reconciliation process to make sure the Elastic Search index document count matches to source records.
  • Developed REST services to write data into Elastic Search index using Python Flask specifications.
  • Implemented Installation and configuration of the multi-node cluster on Cloud using Amazon Web Services (AWS) on EC2.
  • Designed Redshift based data delivery layer for business intelligence tools to operate directly on AWS S3.
  • Created Data Pipeline using Processor Groups and multiple processors using Apache NiFi for Flat File, RDBMS as part of a POC using Amazon EC2.
  • Developed Pig Latin scripts for replacing the existing legacy process to the Hadoop and the data is fed to AWS S3.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster processing of data.
  • Developed data pipelines to consume data from Enterprise Data Lake (MapReduce, Hadoop distribution - Hive tables/HDFS) for the analytics solution.
  • Implemented the Big Data solution using Hadoop, Hive, and Informatica to pull/load the data into the HDFS system.
  • Pulled the data from the data lake (HDFS) and messaging the data with various RDD transformations. Build Hadoop solutions for big data problems using MR1 and MR2 in YARN.
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting on the dashboard.
  • Integrated Apache Storm with Kafka to perform web analytics and to perform click stream data from Kafka to HDFS.
  • Monitoring the functioning of big data and messaging systems like Hadoop, Kafka, Kafka Mirror makers to ensure they operate at their peak performance at all times.
  • Involved in PL/SQL query optimization to reduce the overall run time of stored procedures.
  • Utilized Oozie workflow to run Pig and Hive Jobs extracted files from Mongo DB through Sqoop and placed in HDFS and processed.
  • Continuously tuned Hive UDF's for faster queries by employing partitioning and bucketing.
  • Implemented partitioning, dynamic partitions, and buckets in Hive.
  • Used Flume to collect, aggregate, and store the weblog data from different sources like web servers, mobile, and network devices, and pushed to HDFS.
  • Supported in setting up the QA environment and updating configurations for implementing scripts with Pig, Hive, and Sqoop.

Confidential, Dallas

Big Data Engineer

Responsibilities:

  • Designed and developed Hadoop Data repository/Lake to serve as common data pipeline for internal applications and advanced analysis. This includes designing and developing scripts for pre-processing data from source to HDFS, HDFS to HIVE/HBASE, and transformation of data using Spark Scala/Python scripts for complex data processing
  • Involved in design and development of several big data implementations such as Patient Account receivables efficiency System, Claims Denial Discovery system, DRG analysis Engine, billing coding efficiency, Customer service productivity etc. in the Hadoop platform.
  • Worked on designing HIVE tables to both manage and use externally with data pipeline and interactive querying.
  • Have extensively used Apache Spark features such as RDD operations (mapping, merging, combining, aggregation of data, vectorization of data etc.) and Data frames and datasets for transformation, enrichment of data, data storage operations, applying descriptive statistics, and aggregation of data.
  • Have extensively worked on developing Spark Scala and Python scripts for data ingestion, transformations, building data pipeline for Data scientists, Data Analysis and Business Analysts.
  • Used Scoop API to perform data ingestion to raw zone.
  • Worked on a big data development platform consisting of 124 nodes in Map distribution Cluster.
  • Worked on NLP to process unstructured healthcare survey/feedback data using HIVE, PYSpark and Pandas.
  • Created and worked Sqoop jobs with incremental load to populate Hive External tables
  • Scripts were written for distribution of query for performance test jobs in Amazon Data Lake.
  • Developed optimal strategies for distributing the web log data over the cluster importing and exporting the stored web log data into HDFS and Hive using Sqoop.

Confidential, Atlanta, GA

Big Data Engineer

Responsibilities:

  • Worked onBig Data technologieslikeApache Hadoop, MapReduce, Shell Scripting, and Hive.
  • Involved in all phases ofSDLCusingAgileand participated in daily scrum meetings with cross teams.
  • Wrote complexHive queriesto extract data from heterogeneous sources(Data Lake)and persist the data intoHDFS.
  • Createddata integrationand technical solutions forAzure Data Lake Analytics, Azure Data Lake Storage, Azure Data Factory, Azure SQL databases, andAzure SQL Data Warehousefor providing analytics.
  • Created linked services to connect toAzure Storage,on-premisesSQL Server,andAzure HD Insight.
  • ConfiguredAzure SQL databasewithAzurestorage Explorer and withSQL server.
  • Involved in all phases ofdata mining, data collection, data cleaning, developing models, validation, and visualization.
  • Installed and configuredHadoop ecosystemlikeHBase, Flume, Pig, and Sqoop.
  • Designed and developBig Data analyticsolutions on aHadoop-based platform and engage clients in technical discussions.
  • Installed, Configured, and Maintained theHadoop clusterfor application development and Hadoop ecosystem components likeHive, Pig, HBase, Zookeeper, and Sqoop.
  • Worked onHive queriesto categorize data of different wireless applications and security systems.
  • CreatedHive queriesand tables that helped a line of business identify trends by applying strategies on historical data before promoting them to production.
  • Responsible for loading and transforming huge sets ofstructured, semi-structured,andunstructured data.
  • Extensively involved in writingPL/SQL, stored procedures, functions, and packages.
  • Involved in Data Architecture, Data profiling, Data analysis, data mapping, and Data architecture artifacts design.
  • Worked withNoSQLdatabases likeHBasein creating tables to load large sets ofsemi-structureddata coming from source systems.
  • DevelopedMapReducejobs inScalaforData Cleansingand Analyzing Data inImpala.

Confidential

Python Developer

Responsibilities:

  • Set up and builtAWSinfrastructure with various services available by writing cloud formation templates (CFT) in json and yaml.
  • Developed Cloud Formation scripts to build EC2 on demand
  • With the help of IAM created roles, users and groups and attached policies to provide minimum access to the resources.
  • Updating the bucket policy with IAM role to restrict the access to user.
  • ConfiguredAWSIdentity Access Management (IAM) Group and users for improved login authentication.
  • Created topics in SNS to send notifications to subscribers as per the requirement.
  • Involved in full life cycle of the project from Design, Analysis, logical and physical architecture modeling, development, Implementation, testing.
  • Moving data from Oracle to HDFS using Sqoop
  • Data profiling on critical tables from time to time to check for the abnormalities
  • Created Hive Tables, loaded transactional data from Oracle using Sqoop and Worked with highly unstructured and semi structured data.
  • Developed MapReduce (YARN) jobs for cleaning, accessing, and validating the data.

We'd love your feedback!