Big Data Engineer Resume
Atlanta, GA
SUMMARY
- 6+ years’ experience in Big Data and Hadoop.
- 1.5+ years hands on serving as Linux Systems Administrator.
- Engineer, develop, implement and administer data solutions on prem and on cloud.
- Use Apache Flume and Kafka for collecting, aggregating, and moving data from various sources.
- Build and configure virtual environments in the cloud to support Enterprise Data.
- Design and deploy ELK clusters (Elasticsearch, Logstash, Kibana, Zookeeper).
- Experience with multiple terabytes of data stored in AWS S3 using Elastic Map Reduce (EMR) and Redshift for processing.
- Hands on with AWS tools Redshift, Kinesis, S3, EC2, EMR, DynamoDB, Elasticsearch, Athena, Firehose, Lambda).
- Create Hive Managed and Unmanaged tables with partition and bucket in Hive and loaded data into Hive.
- Developed data queries using HiveQL and optimized the Hive queries.
- Create structured data from the pool of unstructured data using Spark.
- Utilize Spark to optimize ETL jobs to reduce memory and storage consumption.
- Use Spark SQL and DataFrame API extensively to build Spark applications.
- Experienced working on CQL (Cassandra Query Language) for retrieving data present in Cassandra cluster by running queries in CQL.
- Clearly document Big Data systems, procedures, governance and policies.
- Participate in design, development and system migration of high - performance metadata-driven data pipeline with Kafka and Hive.
- Good knowledge in Cluster coordination services through Kafka.
- Extensive Experience streaming data with Kafka.
- Experienced in Cloudera and Hortonworks Hadoop distributions.
- Experienced in Java, Scala and Python programming languages.
- Extend HIVE core functionality by using custom User Defined Function's (UDF), User Defined Table-Generating Functions (UDTF) and User Defined Aggregating Functions (UDAF) for Hive.
TECHNICAL SKILLS
Programming/Scripting: Scala, Python, SQL, Hive QL, Shell Scripting, Java, MySQL
Data Visualization: Kibana, Tableau, PowerBi
ETL Pipelines: Elasticsearch, Elastic MapReduce, ELK Stack (Elasticsearch, Logstash, Kibana), NiFi
Hadoop and Big Data: Apache Flume, Kerberos, Yarn, Cluster Management, Cluster Security, Zookeeper, Oozie, Airflow, Snowflake
Databases/Datastores: Hadoop HDFS, NoSQL, Cassandra, HBase, MongoDB, MySql, MSSql, Oracle, Dynamo
Hadoop Distributions and Cloud: Hadoop, Cloudera Hadoop (CDH), Hortonworks Hadoop (HDP)
Amazon Cloud Platform (AWS): AWS RDS, AWS EMR, AWS Redshift, AWS S3, AWS Lambda, AWS Kinesis, AWS ELK, AWS Cloud, AWS IAM Formation
Other Cloud Platforms: Azure
PROFESSIONAL EXPERIENCE
Big Data Engineer
Confidential, Atlanta, GA
Responsibilities:
- Created PySpark streaming job to receive real time data from Kafka.
- Defined Spark data schema and set up a development environment inside the cluster.
- Processed data with a natural language toolkit to count important words and generated word clouds.
- Started and configured master and slave nodes for Spark.
- Designed Spark Python job to consume information from S3 Buckets using Boto3.
- Set up cloud compute engine in managed and unmanaged mode and SSH key management.
- Worked in virtual machines to run pipelines on a distributed system.
- Utilized a cluster of multiple Kafka brokers to handle replication needs and allow for fault tolerance
- Created a pipeline to gather data using PySpark, Kafka and HBase
- Worked on the Spark Snowflake connector to read and write data from Snowflake table to Spark.
- Used Spark streaming to receive real-time data using Kafka.
- Worked with unstructured data and parsed out the information by Python built-in functions.
- Configured a Python API Producer file to ingest data from the Slack API using Kafka for real-time processing with Spark.
- Managed Hive connection with tables, databases and external tables
- Install Hadoop using Terminal and setting the configurations.
- Formatted the response from Spark jobs to data frames using a schema containing News Type, Article Type, Word Count, and News Snippet to parse JSONs.
- Interacted with data residing in HDFS using PySpark to process the data.
- Configured Linux on multiple Hadoop environments setting up Dev, Test, and Prod clusters within the same configuration.
- Handled HDFS Monitoring job status and life of the DataNodes according to specs.
- Installed spark and PySpark library in terminal using CLI in bootstrapping steps
- Used DynamoDB to store metadata and logs.
- Programmed Python classes to stack information from Kafka to DynamoDB according to the ideal model.
- Provided connections to different Business Intelligence tools to the tables in the data warehouse such as Tableau and Power BI.
AWS Big Data Engineer
Confidential, Chicago, IL
Responsibilities:
- Configured, deployed, and automated instances on AWS, and Data centers.
- Applied EC2, Cloud Watch, Cloud Formation, and managed security groups on AWS.
- Created automated Python scripts to convert data from different sources and to generate ETL pipelines.
- Created Hive external tables and designed information models in Hive.
- Processed Terabytes of information on real time using spark streaming.
- Applied Hive optimization techniques such as partitioning, bucketing, map join, and parallel execution.
- Implemented solutions for ingesting data from various sources and processed the Data-at-Rest utilizing Big Data technologies such as Hadoop, Map Reduce Frameworks, HBase, and Hive.
- Programmed scripts to extract data from different databases and schedule Oozie workflows to execute daily tasks.
- Converted HiveQL/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
- Produced distributed query agents to perform distributed queries against Hive.
- Loaded data from different sources such as HDFS and HBase into Spark data frames and implemented in-memory data computation to generate the output response.
- Monitored Amazon DB and CPU Memory using Cloud Watch.
- Used Spark SQL to realize quicker results compared to Hive throughout information analysis.
- Implemented usage of Amazon EMR for processing Big Data across Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service (S3) AWS Redshift.
- Executed ELK (Elastic Search, Log Stash, Kibana) stack in AWS to gather and investigate the logs created by the website.
- Wrote streaming applications with Spark Streaming/Kafka.
- Developed DBC/ODBC connectors between Hive and Spark for the transfer of the newly populated data frames from MSSQL.
- Executed Hadoop/Spark jobs on AWS EMR using programs and data stored in S3 Buckets.
- Used Datastax Spark Cassandra Connector to extract and load data to/from Cassandra.
Hadoop Engineer
Confidential
Responsibilities:
- Installed and configured Hadoop HDFS and developed multiple jobs in Java for data cleaning and preprocessing.
- Developed Map/Reduce jobs using Java for data transformations.
- Developed different components of systems’ Hadoop processes that involved Map Reduce and Hive.
- Developed data pipeline using Sqoop, MR, and Hive to extract the data from weblogs and store the results for downstream consumption.
- Developed Hive queries and UDFS to analyze/transform the data in HDFS.
- Designed and Implemented Partitioning (Static, Dynamic), Buckets in Hive.
- Worked in a team to develop an ETL pipeline that involved extraction of Parquet serialized files from S3 and persisted them in HDFS.
- Used Sqoop to efficiently transfer data between databases and HDFS and used Flume to stream the log data from servers.
- Used Zookeeper and Oozie for coordinating the cluster and scheduling workflows.
- Configured Yarn capacity scheduler to support various business SLAs.
- Implemented and maintained security with LDAP and Kerberos as designed for cluster.
- Coordinates with monitors cluster upgrade needs, and monitors cluster health and builds proactive tools to look for anomalous behaviors.
- Worked with cluster users to ensure efficient resource usage in the cluster and alleviate multi-tenancy concerns.
- Migrated ETL processes from Oracle to Hive to test the easy data manipulation.
- Wrote HiveQL scripts to perform trend analysis on Big Data log data.
- Utilized Sqoop to extract data back to relational databases for business reporting.
- Created Hive tables, and loaded data and wrote Hive queries.
- Debugged and identified issues reported by QA with the Hadoop jobs by configuring a local file system.
- Implemented Flume to import streaming data logs and aggregating the data to HDFS.
- Used Cloudera Manager for installation and management of single-node and multi-node Hadoop Cluster.
Linux Systems Administrator
Confidential, Fort Mill, SC
Responsibilities:
- Installed, configured, monitored, and administered Linux servers.
- Installed, deployed, and managed Linux RedHat Enterprise, CentOS, Ubuntu, and installed patches and packages for Red Hat Linux Servers.
- Configured and installed RedHat and Centos Linux Servers on virtual machines and bare metal installations.
- Worked with the DBA team for database performance issues, network related issues on LINUX/UNIX servers and with vendors regarding hardware related issues.
- Monitored CPU, memory, hardware and software including raid, physical disk, multipath, filesystems, and networks using Nagios monitoring tool.
- Hosted servers using Vagrant on Oracle virtual machines.
- Automated daily tasks using bash scripts while documenting the changes in the environment and in each server, analyzing the error logs, user logs and /var/log messages.
- Created and modified users and groups with root permissions.
- Administered local and remote servers using the SSH on a daily basis.
- Created and maintained Python scripts for automating build and deployment processes.
- Utilized Nagios-based open-source monitoring tools to monitor Linux Cluster nodes.
- Created users, managed user permissions, maintained user and file system quotas, and installed and configured DNS.
- Adhered to industry standards by securing systems, directory and file permissions, groups and supporting user account management along with the creation of users.
- Performed kernel and database configuration optimization such as I/O resource usage on disks.
- Analyzed and monitored log files to troubleshoot issues.
