We provide IT Staff Augmentation Services!

Sr. Aws Big Data Engineer Resume

2.00/5 (Submit Your Rating)

Phoenix, AZ

SUMMARY

  • Over 10+ yearsof experience in IT Industry in the Big data platform having extensive hands - on experience in Apache Hadoop ecosystem and enterprise application development.
  • Extensive experience and actively involved in Requirements gathering, Analysis, Design, Developing and Code Reviews, Unit and Integration Testing.
  • Hands on experience on Data Analytics Services such as Athena, Glue Data Catalog & Quick Sight
  • Working experience with large scale Hadoop environments build and support including design, configuration, installation, performance tuning and monitoring.
  • Experienced in analyzing data using HiveQL, HBase and custom Map Reduce programs in Java.
  • Experience in developing Microservices with Spring boot using Java and Akka framework using Scala.
  • Hands on experience in installing, configuring, and using Hadoop ecosystem components like Hadoop Map Reduce, HDFS, HBase, Oozie, Hive, Sqoop, Pig, Storm, Kafka, Zookeeper, Yarn and Spark.
  • Involved in converting Hive/SQL queries into Spark transformations (RDD & DataFrames) using Pyspark.
  • Experience in writing Spark application using Python and Scala
  • Experience in developing pipelines geared for scalability, performance, easy to maintain, and creating monitoring and alert systems.
  • Worked with different file formats (AVRO, JSON, ORC, Parquet) and has the knowledge about data compression techniques (LZO, Bzip2 and Snappy).
  • Good knowledge in using ApacheNiFito automate the data movement between different Hadoop systems.
  • Hands-on experience withAmazon EC2, Amazon S3, Amazon RDS, VPC, IAM, Amazon Elastic Load Balancing, Auto Scaling, CloudWatch, SNS, SQSand other services of the AWS family.
  • Extensive expertise using the core Spark APIs and processing data on a EMR cluster
  • Implemented cost control strategies on AWS by selecting appropriate services to design and deploy an application based on given requirements.
  • Experience in NoSQL data stores (HBase, Accumulate and Mongo DB)
  • Experience in AWS cloud integration with Amazon Elastic MapReduce (EMR), Amazon Cloud Compute (EC2) and Amazon's Simple Storage Service (S3).
  • Strong proficiency with Hadoop ecosystem, utilizing tools both on premises and on cloud platforms.
  • Good knowledge on extracting the models and trends from the raw data collaborating with the data science team.
  • Strong experience in complete project life cycle (design, development, testing and implementation) of Client Server and Web applications.
  • Experience on agile methodologiesScrum.
  • Ability to adapt to evolving technology, strong sense of responsibility and accomplishment.
  • Ability to meet deadlines and handle multiple tasks, flexible in work schedules and possess good communication skills

TECHNICAL SKILLS

Hadoop Eco System: HDFS, Map Reduce, Pig, Hive, HBase, Mahout, Falcon, Oozie, Accumulo, Zookeeper, YARN, Spark, Kafka, Flume and Sqoop

Programming Languages: C, C++, Java, Python and Scala.

Technologies and Tools: JSP, Java Bean, Servlets, Spring, Hibernate, Maven, JDBC, JPA1.0, EJB3.0, Amazon Cloud (S3, EC2, EMR, Glue, Athena, Lambda and RDS).

Web Technologies: HTML, JavaScript, AJAX, XML, CSS, JQuery, Perl, VB Script.

Application Servers: WebLogic8.1/9.1/10.x, Web-Sphere5.x/6.x/7.x, Glass Fish Server 2.x, JBoss 4.x/5.x.

Web Servers: Apache Tomcat 4.0/ 5.5, Java Web Server 2.0.

Operating Systems: Windows 10/8.1/7/XP, UNIX, Linux (Red Hat/Ubuntu/CentOS)

Databases: SQL, PL/SQL, Oracle 9i/10g, MYSQL, Microsoft Access, SQLServer, No SQL (HBASE, MongoDB).

IDE’s: Eclipse, Net Beans, Jupyter Notebook, Visual Studio and IntelliJ IDEA

Platforms: Windows XP/NT/9x/2000, MS-DOS, UNIX /LINUX/Solaris/AIX

Distribution: Cloudera and Hortonworks

Version Control: Win CVS, VSS, PVCS, Subversion, GIT and Git-Lab

PROFESSIONAL EXPERIENCE

Confidential, Phoenix, AZ

Sr. AWS Big Data Engineer

Responsibilities:

  • Analyzed and processed complex data sets using advanced querying, visualization and analytics tools.
  • Involved in designing and deploying multi-tier applications using the AWS services like (EC2, Route53, S3, RDS, Dynamo DB, SNS, SQS, IAM) focusing on high-availability, fault tolerance, and auto-scaling in AWS Cloud Formation
  • Worked on ETL Migration services by developing and deploying AWS Lambda functions for generating a serverless data pipeline which can be written to Glue Catalog and can be queried from Athena.
  • Programmed in Hive, Spark SQL, and Python to streamline the incoming data and build the data pipelines to get the useful insights, and orchestrated pipelines
  • Extensive expertise using the core Spark APIs and processing data on a EMR cluster
  • Supporting Continuous storage in AWS using Elastic Block Storage, S3, Glacier. Created Volumes and configured Snapshots for EC2 instances
  • Used Data Frame API in Python for converting the distributed collection of data organized into named columns, developing predictive analytic using Apache Spark APIs.
  • Developed REST based Scala service to pull data from Elasticsearch dashboard, Splunk and Atlassian Jira.
  • Developed PySpark scripts using both Data frames/SQL/Data sets and RDD/MapReduce in Spark for Data Aggregation, queries and writing data back into OLTP system through Sqoop.
  • Developed native Scala/Java library using Jsch to remotely execute Auto Logs.
  • Developed Hive queries to pre-process the data required for running the business process
  • Experience in using Terraform to create Infrastructure as Code on AWS
  • Collaborated with data science team and e-commerce team to successfully deploy and integrate the models.
  • Engineered an automated ETL pipeline for data ingestion and feature engineering using AWS Sagemaker and configured job failure alerts and notify SNS.
  • Manage code repository using Git to ensure integrity of code base is maintained at all times
  • Ensured system architecture met business requirements, constantly worked with different teams to ensure every aspect of architecture is beneficial to the company

Confidential, Atlanta, GA

Data Engineer/Spark Developer

Responsibilities:

  • Developed Spark Streaming applications in Java and Scala for data loading and transformation.
  • Created a framework using spark streaming and Kafka to process data in Real time which feeds data to APIs.
  • Developed Microservices based on Restful web service using Akka actors and Akka-Http framework in Scala which handles high concurrency and high volume of traffic.
  • Involved in Analyzing data from different sources like Teradata, MySQL and Sqooping data into Hive using Sqoop.
  • Written UDF, UDTF and UDAF in spark to implement business logic on data.
  • Created a Real time data pipelines and frameworks with Kafka, Spark streaming and loading data to Hbase.
  • Worked on Nifi processors for creating data pipelines to copy data from JMS MQ to Kafka topics and processing data in between like JSON to XML conversion.
  • Monitored Spark Web UI, DAG scheduler and Yarn resource manager UI to optimize queries and performance in spark.
  • Worked on different file formats like Parquet, Avro files and ORC file formats.
  • Experience in AWS architecture (ELB, EC2, S3, RDS, CloudFormation, Security Groups), DevOps practices, Unix (RedHat), networking, web services architecture are essential for this hands-on role.
  • Migrated data pipeline jobs from on-premise to AWS environment of SVOC project.
  • Developed AWS Lambda functions to consume events from Kafka (MSK) and write to SQS queue and S3.
  • Written Integration test cases using Cucumber framework for test-driven development, created step definitions to test API and data pipeline functionality
  • Worked with Site reliability and Hadoop admin team to resolve production issues.
  • Designed table structure for different projects in HBase, Hive according to the design doc and mapping requirements.

Confidential, Dallas, TX

Big Data Engineer

Responsibilities:

  • Extensive experience programming Spark jobs to process complex data formats (structured, semi-structured and unstructured data)
  • Experience in handling real time data using Kafka and Spark streaming
  • Experience loading external data to Hadoop environments using tools like Sqoop and Flume
  • Scripting experience in PySpark, which involves cleansing and transformation of data
  • Experience working with very large data sets, knows how to build programs that leverage the parallel capabilities of Hadoop with loading data to Hive
  • Used Hive optimization techniques during joins and best practices in writing hive scripts.
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting
  • Hands on experience in writing custom UDF and serde functions in hive
  • Implemented Batch Analytics leveraging Spark data frames and Spark SQL.
  • Designed and implemented real time analytics platform using Spark Streaming and Structured Streaming to ingest billions of events and process them at scale and store the aggregated results in NoSQL stores.
  • Developed custom Spark UDFs using PySpark.
  • Optimizing of existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames and Pair RDD's.
  • Extensive Experience in creating and managing EMR Clusters and EC2 instances.
  • Experience in processing the data through S3 buckets, managing AMI policies of S3.
  • Hands on experience on managing SNS, SQS alerts and triggering them programmatically
  • Interacted with data scientists and industry experts to understand how data needs to be converted, loaded and presented.
  • Worked with both technical and business-oriented teams as a part of daily activity

Confidential, Minneapolis, MN

Hadoop Developer

Responsibilities:

  • Configured different topologies for the PriceIndexing Storm cluster and deployed them on regular basis.
  • Consuming the data form Kafka queue using storm and applied rules engine to determine drive sales or drive margin.
  • Implemented Sqoop to connect to the DB2 and move the data to HDFS and created Hive tables.
  • Developed job processing scripts using Oozie workflow.
  • Involved in designing and implementation of Hadoop.
  • Developed integration code for accessing Oracle and DB2 databases.
  • Involving in day-to-day standups and working closely with clients and BA’s.
  • Developed unit test cases using Jmockit framework and automated the scripts.
  • Worked in Agile environment, which uses Version one to maintain the story points.
  • Experience in setting up a cluster environment on OpenStack servers for storm, Kafka and zookeeper.
  • Maintained different cluster security settings and involved in creation and termination of multiple cluster environments.

Confidential, Boston, MA

Hadoop Developer

Responsibilities:

  • Launching and Setup of HADOOP related tools on AWS, which includes configuring different components of HADOOP.
  • Experience in Using Sqoop to connect to the Sql Server or Oracle database and move the pivoted data to Hive tables and stored in Avro files.
  • Managed the Hive database, which involves ingest and index of data.
  • Expertise in exporting the data from Avro files and indexing the documents in sequence or serde file format.
  • Hands on experience in writing custom UDF’s and also custom input and output formats.
  • Scheduling the Hive jobs using Oozie and falcon process files.
  • Developed Map Reduce jobs to store the data in to HBase tables.
  • Involved in design and architecture of custom Lucene storage handler.
  • Configured and Maintained different topologies in storm cluster and deployed them on regular basis.
  • Understanding of Ruby scripts used to generated yaml files.
  • Experience in working with BI reporting tools to create dashboards.
  • Involved in GUI development using JavaScript and AngularJS and Guice.
  • Developed Unit test case using Jmockit framework and automated the scripts.
  • Worked in Agile environment, which uses Jira to maintain the story points and Kanban model.
  • Involved in implementing Kerberos secured environment for Hadoop cluster.
  • Hands on experience on maintaining the builds in Bamboo and resolved the build failures in Bamboo.

Environment: Hadoop, Big data, Hive, Hbase, Sqoop, Accumulo, Oozie, Falcon, HDFS, Map Reduce, Jira, Bit bucket, Maven, Bamboo, J2EE, Guice, AngularJS, Jmockit, Lucene, Storm, Ruby, Unix, Sql, AWS (Amazon Web Services).

Confidential, Houston, TX

Java/J2EE Developer

Responsibilities:

  • Used Hibernate ORM tool as persistence Layer - using the database and configuration data to provide persistence services (and persistent objects) to the application.
  • Responsible for developing DAO layer using Spring MVC and configuration XML’s for Hibernate and to also manage CRUD operations (insert, update, and delete).
  • Implemented Dependency injection of spring framework.
  • Developed reusable services using BPEL to transfer data.
  • Created JUnit test cases, and Development of JUnit classes.
  • Configured log4j to enable/disable logging in application.
  • Developed Rich user interface using HTML, JSP, AJAX, JSTL, Java Script, JQuery and CSS.
  • Implemented PL/SQL queries, Procedures to perform data base operations.
  • Wrote UNIX Shell scripts and used UNIX environment to deploy the EAR and read the logs.
  • Implemented Log4j for logging purpose in the application.

Environment: Java, Jest, SOA Suite 10g (BPEL), Struts, Spring, Hibernate, Web services (JAX-WS), JMS, EJB, Web logic 10.1 Server, JDeveloper, Sql Developer, HTML, LDAP, Maven, XML, CSS, JavaScript, JSON, SQL, PL/SQL, Oracle, JUnit, CVS and UNIX/Linux.

We'd love your feedback!