We provide IT Staff Augmentation Services!

Big Data Engineer Resume

2.00/5 (Submit Your Rating)

Atlanta, GA

SUMMARY

  • 6+ years’ experience in Big Data and Hadoop.
  • 1.5+ years hands on serving as Linux Systems Administrator.
  • Engineer, develop, implement and administer data solutions on prem and on cloud.
  • Use Apache Flume and Kafka for collecting, aggregating, and moving data from various sources.
  • Build and configure virtual environments in the cloud to support Enterprise Data.
  • Design and deploy ELK clusters (Elasticsearch, Logstash, Kibana, Zookeeper).
  • Experience with multiple terabytes of data stored in AWS S3 using Elastic Map Reduce (EMR) and Redshift for processing.
  • Hands on with AWS tools Redshift, Kinesis, S3, EC2, EMR, DynamoDB, Elasticsearch, Athena, Firehose, Lambda).
  • Create Hive Managed and Unmanaged tables with partition and bucket in Hive and loaded data into Hive.
  • Developed data queries using HiveQL and optimized the Hive queries.
  • Create structured data from the pool of unstructured data using Spark.
  • Utilize Spark to optimize ETL jobs to reduce memory and storage consumption.
  • Use Spark SQL and DataFrame API extensively to build Spark applications.
  • Experienced working on CQL (Cassandra Query Language) for retrieving data present in Cassandra cluster by running queries in CQL.
  • Clearly document Big Data systems, procedures, governance and policies.
  • Participate in design, development and system migration of high - performance metadata-driven data pipeline with Kafka and Hive.
  • Good knowledge in Cluster coordination services through Kafka.
  • Extensive Experience streaming data with Kafka.
  • Experienced in Cloudera and Hortonworks Hadoop distributions.
  • Experienced in Java, Scala and Python programming languages.
  • Extend HIVE core functionality by using custom User Defined Function's (UDF), User Defined Table-Generating Functions (UDTF) and User Defined Aggregating Functions (UDAF) for Hive.

TECHNICAL SKILLS

Programming/Scripting: Scala, Python, SQL, Hive QL, Shell Scripting, Java, MySQL

Data Visualization: Kibana, Tableau, PowerBi

ETL Pipelines: Elasticsearch, Elastic MapReduce, ELK Stack (Elasticsearch, Logstash, Kibana), NiFi

Hadoop and Big Data: Apache Flume, Kerberos, Yarn, Cluster Management, Cluster Security, Zookeeper, Oozie, Airflow, Snowflake

Databases/Datastores: Hadoop HDFS, NoSQL, Cassandra, HBase, MongoDB, MySql, MSSql, Oracle, Dynamo

Hadoop Distributions and Cloud: Hadoop, Cloudera Hadoop (CDH), Hortonworks Hadoop (HDP)

Amazon Cloud Platform (AWS): AWS RDS, AWS EMR, AWS Redshift, AWS S3, AWS Lambda, AWS Kinesis, AWS ELK, AWS Cloud, AWS IAM Formation

Other Cloud Platforms: Azure

PROFESSIONAL EXPERIENCE

Big Data Engineer

Confidential, Atlanta, GA

Responsibilities:

  • Created PySpark streaming job to receive real time data from Kafka.
  • Defined Spark data schema and set up a development environment inside the cluster.
  • Processed data with a natural language toolkit to count important words and generated word clouds.
  • Started and configured master and slave nodes for Spark.
  • Designed Spark Python job to consume information from S3 Buckets using Boto3.
  • Set up cloud compute engine in managed and unmanaged mode and SSH key management.
  • Worked in virtual machines to run pipelines on a distributed system.
  • Utilized a cluster of multiple Kafka brokers to handle replication needs and allow for fault tolerance
  • Created a pipeline to gather data using PySpark, Kafka and HBase
  • Worked on the Spark Snowflake connector to read and write data from Snowflake table to Spark.
  • Used Spark streaming to receive real-time data using Kafka.
  • Worked with unstructured data and parsed out the information by Python built-in functions.
  • Configured a Python API Producer file to ingest data from the Slack API using Kafka for real-time processing with Spark.
  • Managed Hive connection with tables, databases and external tables
  • Install Hadoop using Terminal and setting the configurations.
  • Formatted the response from Spark jobs to data frames using a schema containing News Type, Article Type, Word Count, and News Snippet to parse JSONs.
  • Interacted with data residing in HDFS using PySpark to process the data.
  • Configured Linux on multiple Hadoop environments setting up Dev, Test, and Prod clusters within the same configuration.
  • Handled HDFS Monitoring job status and life of the DataNodes according to specs.
  • Installed spark and PySpark library in terminal using CLI in bootstrapping steps
  • Used DynamoDB to store metadata and logs.
  • Programmed Python classes to stack information from Kafka to DynamoDB according to the ideal model.
  • Provided connections to different Business Intelligence tools to the tables in the data warehouse such as Tableau and Power BI.

AWS Big Data Engineer

Confidential, Chicago, IL

Responsibilities:

  • Configured, deployed, and automated instances on AWS, and Data centers.
  • Applied EC2, Cloud Watch, Cloud Formation, and managed security groups on AWS.
  • Created automated Python scripts to convert data from different sources and to generate ETL pipelines.
  • Created Hive external tables and designed information models in Hive.
  • Processed Terabytes of information on real time using spark streaming.
  • Applied Hive optimization techniques such as partitioning, bucketing, map join, and parallel execution.
  • Implemented solutions for ingesting data from various sources and processed the Data-at-Rest utilizing Big Data technologies such as Hadoop, Map Reduce Frameworks, HBase, and Hive.
  • Programmed scripts to extract data from different databases and schedule Oozie workflows to execute daily tasks.
  • Converted HiveQL/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
  • Produced distributed query agents to perform distributed queries against Hive.
  • Loaded data from different sources such as HDFS and HBase into Spark data frames and implemented in-memory data computation to generate the output response.
  • Monitored Amazon DB and CPU Memory using Cloud Watch.
  • Used Spark SQL to realize quicker results compared to Hive throughout information analysis.
  • Implemented usage of Amazon EMR for processing Big Data across Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service (S3) AWS Redshift.
  • Executed ELK (Elastic Search, Log Stash, Kibana) stack in AWS to gather and investigate the logs created by the website.
  • Wrote streaming applications with Spark Streaming/Kafka.
  • Developed DBC/ODBC connectors between Hive and Spark for the transfer of the newly populated data frames from MSSQL.
  • Executed Hadoop/Spark jobs on AWS EMR using programs and data stored in S3 Buckets.
  • Used Datastax Spark Cassandra Connector to extract and load data to/from Cassandra.

Hadoop Engineer

Confidential

Responsibilities:

  • Installed and configured Hadoop HDFS and developed multiple jobs in Java for data cleaning and preprocessing.
  • Developed Map/Reduce jobs using Java for data transformations.
  • Developed different components of systems’ Hadoop processes that involved Map Reduce and Hive.
  • Developed data pipeline using Sqoop, MR, and Hive to extract the data from weblogs and store the results for downstream consumption.
  • Developed Hive queries and UDFS to analyze/transform the data in HDFS.
  • Designed and Implemented Partitioning (Static, Dynamic), Buckets in Hive.
  • Worked in a team to develop an ETL pipeline that involved extraction of Parquet serialized files from S3 and persisted them in HDFS.
  • Used Sqoop to efficiently transfer data between databases and HDFS and used Flume to stream the log data from servers.
  • Used Zookeeper and Oozie for coordinating the cluster and scheduling workflows.
  • Configured Yarn capacity scheduler to support various business SLAs.
  • Implemented and maintained security with LDAP and Kerberos as designed for cluster.
  • Coordinates with monitors cluster upgrade needs, and monitors cluster health and builds proactive tools to look for anomalous behaviors.
  • Worked with cluster users to ensure efficient resource usage in the cluster and alleviate multi-tenancy concerns.
  • Migrated ETL processes from Oracle to Hive to test the easy data manipulation.
  • Wrote HiveQL scripts to perform trend analysis on Big Data log data.
  • Utilized Sqoop to extract data back to relational databases for business reporting.
  • Created Hive tables, and loaded data and wrote Hive queries.
  • Debugged and identified issues reported by QA with the Hadoop jobs by configuring a local file system.
  • Implemented Flume to import streaming data logs and aggregating the data to HDFS.
  • Used Cloudera Manager for installation and management of single-node and multi-node Hadoop Cluster.

Linux Systems Administrator

Confidential, Fort Mill, SC

Responsibilities:

  • Installed, configured, monitored, and administered Linux servers.
  • Installed, deployed, and managed Linux RedHat Enterprise, CentOS, Ubuntu, and installed patches and packages for Red Hat Linux Servers.
  • Configured and installed RedHat and Centos Linux Servers on virtual machines and bare metal installations.
  • Worked with the DBA team for database performance issues, network related issues on LINUX/UNIX servers and with vendors regarding hardware related issues.
  • Monitored CPU, memory, hardware and software including raid, physical disk, multipath, filesystems, and networks using Nagios monitoring tool.
  • Hosted servers using Vagrant on Oracle virtual machines.
  • Automated daily tasks using bash scripts while documenting the changes in the environment and in each server, analyzing the error logs, user logs and /var/log messages.
  • Created and modified users and groups with root permissions.
  • Administered local and remote servers using the SSH on a daily basis.
  • Created and maintained Python scripts for automating build and deployment processes.
  • Utilized Nagios-based open-source monitoring tools to monitor Linux Cluster nodes.
  • Created users, managed user permissions, maintained user and file system quotas, and installed and configured DNS.
  • Adhered to industry standards by securing systems, directory and file permissions, groups and supporting user account management along with the creation of users.
  • Performed kernel and database configuration optimization such as I/O resource usage on disks.
  • Analyzed and monitored log files to troubleshoot issues.

We'd love your feedback!