Data Engineer Resume
Dallas, TX
SUMMARY
- Having 5+ years of IT experience in Analysis, design, development, Testing and Deployment of Distributed and cloud applications by using Apache Spark, Hadoop, Hive, Sqoop, Kafka and AWS Services like EMR, S3 and RedShift.
- Worked with Big Data distributions like Cloudera (CDH) and Hortonworks (HDP) and have an in - depth understanding of Hadoop Architecture including YARN and various components such as HDFS, Resource Manager, Node Manager, Name Node, Data Node.
- Expertise in AWS cloud services like EC2, VPC, S3, EMR, RedShift, CloudWatch and Lambda functions.
- Hands-on experience with major components in Hadoop Ecosystem like Map Reduce, HDFS, YARN, Hive, Sqoop, Oozie and Flume.
- Experienced in working with Datasets, Spark-SQL, Data Frames, RDD’s handling large data frames using Partitions, Spark in-Memory capabilities, Effective & efficient Joins.
- Experience with Apache Spark’s Core, Spark SQL, DStreams and Structured Streams.
- Experienced in scheduling and orchestrating of Spark and Hadoop jobs and data ingestion to hive using workflows like Oozie
- Experience with complex Data processing pipelines, including ETL and Data ingestion dealing with unstructured and semi-structured Data.
- Hands on experience in Avro, Parquet, ORC files, Dynamic Partitions, bucketing for best practice and performance improvement, worked on different Compression Codec's (GZIP, SNAPPY, BZIP).
- Expertise in working with Hive data warehouse tool-creating tables, data distribution by implementing partitioning and bucketing, writing and optimizing the HiveQL queries.
- Experience in loading the datafrom the different datasources like Teradata into HDFS using TDCH Teradata connector and load into partitioned Hive tables.
- Experience in migrating data by using SQOOP from HDFS to Relational Database System and vice-versa according to client's requirements.
- Experience in working with different scripting technologies like Python, UNIX shell scripts.
- Hands-on experience across all stages of Software Development Life Cycle (SDLC) including business requirement analysis, data mapping, build, unit testing, systems integration, UAT and Prod.
- Worked on Agile methodology by adhering to standards and techniques and prepared code documentation, also documented the problems and provided solutions.
- Very good communication skills, interpersonal skills, and problem-solving skills, explore/adapt to new technologies with ease and a good team member.
TECHNICAL SKILLS
Big Data Ecosystem: HDFS, Yarn, MapReduce, Spark, Kafka, Kafka Connect, Hive, Sqoop, HBaseFlume, Oozie, Zookeeper
Cloud Environments: AWS Glue, EMR, S3, Redshift
Operating Systems: Linux, Windows
Languages: Python
Databases: Oracle, SQL Server, MySQL, HBase, RedShift, DynamoDB
Version Control Tools: SVN, GIT, Bitbucket
Automation: Step Functions, Control-M, Oozie
CI/CD Tools: Jenkins, Maven
Monitoring Tools: AWS cloud watch
Scripting Languages: bash/Shell scripting, Linux/Unix.
Reporting: Tableau
PROFESSIONAL EXPERIENCE
Confidential, Dallas, TX
Data Engineer
Environment: Hadoop, HDFS, Spark, Yarn, Hive, Sqoop, Python, Kafka, Tableau, SQL Server, Shell Scripting, Oozie, Linux
Responsibilities:
- Developed Spark jobs to create data frames from the source system, process, and analyze the data based on business requirements.
- Involved in enhancing the existing ETL pipeline for better data migration with reduced data issues.
- Wrote Spark SQL queries to design the solutions and implemented them using PySpark.
- Implemented static Partitioning, Dynamic partitioning, and Bucketing in Hive using internal and external tables.
- Experienced in developing Spark scripts for data analysis in python.
- Scheduling workflow to orchestrate multiple Hive and Spark jobs. Wrote shell scripts to automate the jobs in UNIX.
- Worked on monitoring, scheduling, and authoring Data Pipelines using Oozie.
- Creating Hive tables, loading with data, and writing Hive queries, which will invoke MapReduce jobs in the backend.
- Utilized the Apache Hadoop environment by Cloudera. Monitoring and Debugging Spark jobs which are running on a Spark cluster using Cloudera Manager.
Confidential, TX
Data Engineer
Environment: Spark, Hive, AWS Glue, Athena, Redshift, Pyspark, Informatica, Oracle, S3, Python, SparkSQL, Jenkins, Maven.
Responsibilities:
- ImplementedSparkusing Python (PySpark) andSparkSQLfor faster processing of data.
- Develop ETL Processes in AWS to migrate data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.
- Create external tables with partitions using Hive, AWS Athena and Redshift.
- Worked with AWS S3 data and AWS jobs to transform data to a format that optimizes query performance for Athena.
- Used AWS Glue for the data transformation, validation, and data cleansing.
- Develop script to create external tables and updated partitioning information on daily basis
- Migrating of On-Prem existing Spark jobs into AWS jobs.
- Optimize the Spark joins by using various techniques such as Broadcast join, Bucket and skew join.
- Implemented a real-time data pipeline using Spark structured streams or discrete streams to ingest event data into Hive from Kafka.
- Worked on different file formats like ORCFILE, Parquet and Avro and used different Compression Codecs GZIP, SNAPPY, LZO.
- Developed data ingestion, preprocess, post ingestion transformation from various data sources like Oracle and SQL Server using Sqoop and loaded data into Hive as ORC tables.
- Experience to use various spark api’s like RDD and DataFrame and created data ingestion pipeline by applying transformations.
- Implemented new project builds framework usingJenkins,Mavenas build frameworks.
Confidential, TX
Data Engineer
Environment: Hive, Spark, Teradata, Pyspark, HDFS, YARN, Sqoop, Hue, Linux, Python, Jira, Mysql, Oracle.
Responsibilities:
- Worked on Hadoop Bigdata technologies like Hive, Spark for data processing.
- Developed code to import data from Teradata into HDFS and created Hive views on data in HDFS using Spark in Python.
- Worked extensively with Sqoop for importing and exporting the data from HDFS to Relational Database systems and vice-versa loading data into HDFS.
- Developed Spark transformations on RDD’s and Data frames from python using pyspark.
- Developed complex queries in Hive to apply transformations.
- Created various hive partitioned and bucketed tables with different formats like Avro and Orc with schema as per requirement.
- Enabled speedy reviews and first mover advantages by using Oozie to automate data loading into the Hadoop Distributed File System.
- Used Hadoop file system commands to manage and operate files located in HDFS.
- Worked on Yarn commands to manage applications submitted into Yarn.
- Involved in converting Hive/HQL queries into Spark transformations using Spark data frames using python and Spark Sql.
- Monitoring YARN applications. Troubleshoot and resolve cluster related system problems.
- Primary responsibilities included building scalable distributed data solutions using the Hadoop ecosystem.
- Optimized MapReduce Jobs to use HDFS efficiently by using various compression mechanisms.
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS, and Extracted the data from Teradata into HDFS using Sqoop.
- Loaded datasets from Teradata to HDFS and Hive on daily basis.
- Analyzed large amounts of data sets to determine the optimal way to aggregate and report on it.
