We provide IT Staff Augmentation Services!

Hadoop Spark/scala Developer Resume

4.00/5 (Submit Your Rating)

Austin, TX

SUMMARY:

  • Around 8 years of extensive work experience as a Big Data Hadoop Ecosystem, Deep understanding of Hadoop Architecture, Spark Execution engine, Hive data warehousing and NO - SQL databases.
  • Good experience in Hadoop infrastructure which includes Map Reduce, Pig, Hive, Sqoop, Flume, Hbase, HDFS, Spark.
  • In depth understanding/knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and MapReduce concepts.
  • Good knowledge on Data Warehousing, ETL development, Distributed Computing, and large scale data processing.
  • Good team player and ability to work in fast paced environment.
  • Well experienced in data transformation using MapReduce, Hive and Pig scripts for different types of file formats.
  • Strong knowledge of HDFS Architecture and its basic components such as Name node, Data node, Job tracker, Task tracker and Map Reduce concepts.
  • Worked on HBase to load and retrieve data for real time processing.
  • Hands on experience in writing Pig Latin Scripts and Pig commands and also Hive queries.
  • Worked with Apache Spark which provides fast and general engine for large data processing integrated with functional programming language Scala.
  • Experienced in writing Map Reduce programs using Java to perform transformations on different data sets using Mapper and Reducer tasks .
  • Hands on experience with Scala 2.10 or higher, SBT, ScalaTest, ScalaCheck
  • Extensively worked on Hive. Created tables and wrote numerous queries.
  • Hands on experience on Hadoop 2.5 and YARN configurations and major components in Hadoop Ecosystem including Hive, Hbase, Sqoop, Flume and knowledge in Mapper/Reduce/HDFS Framework.
  • Hands on experience with MapReduce, Pig, Programming Model, Installation and Configuration of Hadoop, HBase, Hive, Pig, Sqoop and Flume using Linux commands.
  • Strong knowledge on creating and monitoring Hadoop cluster on VM, CDH3, CDH4 Cloudera Manager on Linux, Ubuntu OS.
  • Worked with BI tools like Tableau for report creation and further analysis from the front end. Connecting Hive using Tableau and generating Bar chart etc based on business requirement.
  • Hands on experience with SPARK to handle the streaming data and SCALA for the batch processing and spark streaming data.
  • Experience in querying RDBMS such as MYSQL and SQL Server by using SQL for data integrity.
  • Certified administrator for Cloudera Manager and Hortonworks, Ambari.
  • Experience in Database Design and Development using Relational Databases (Oracle, MYSQL, DB2, MySQL Server 2005/2008) and NoSQL Databases (MongoDB, Cassandra, HBase, DynamoDB)
  • Experience in developing service components using JDBC.
  • Extensive knowledge in using SQL queries for back-end database analysis.
  • Experience using Sqoop to import data into HDFS from RDBMS and vice-versa.
  • Good experience in using reporting tools like Tableau and creating reports for hive data.
  • Expert in understanding Operating Systems like Linux, Unix, Windows.
  • Ability in using of Java, XML, Shell Scripts, Python, for monitoring and to automate the build and deploying process.
  • Processing this data using Spark Streaming API with Scala .
  • Strong Problem Solving and Analytical skills and abilities to make Balanced and Independent Decisions.
  • Ability to work with onsite and offshore team members .

TECHNICAL SKILLS:

Big Data Hadoop Technologies: Hadoop, HDFS, Map Reduce, Hive, Sqoop, Pig, HBase, Spark, Flume, Kafka, Oozie.

Platforms: GNU/LINUX (Ubuntu and Debian), Windows.

NO SQL Databases: HBase, Cassandra.

Programming Languages: C, C++, Python, HTML and Java.

Database Languages: Mysql, Sql Server, No Sql.

Scripting Languages: Shell Scripting, JavaScript, XML, HTML, Scala

Web Tools/Technologies: HTML, XML, CSS.

CM Tools: Chef, Puppet and Ansible.

Business Intelligence Tools: Tableau, Splunk.

PROFESSIONAL EXPERIENCE:

Confidential, Austin, TX

Hadoop Spark/Scala Developer

Responsibilities:

  • Developed Simple to complex Map/reduce streaming jobs using Java language that are implemented Using Hive and Pig.
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and Extracted the data from Oracle into HDFS using Sqoop.
  • Using Spark Streaming to bring all credit card transactions in the Hadoop environment.
  • Analyzed the data by performing Hive queries (HiveQL) and running Pig scripts (Pig Latin) to study customer behavior.
  • Was involved in data modelling through the process of No-SQL database implementation .
  • Involved in source system analysis, data analysis, data modeling to ETL (Extract, Transform and Load) Worked extensively in creating MapReduce jobs to power data for search and aggregation.
  • Implemented Spark using Scala and Spark SQL for faster testing and processing of data.
  • Worked extensively with Sqoop for importing and exporting the data from HDFS to Relational Database systems/mainframe and vice-versa. Loading data into HDFS.
  • Extensively used Pig for data cleansing.
  • Worked on migrating MapReduce programs into Spark transformations using Spark and Scala, initially done using Python (PySpark).
  • Develop Hive queries for the analysts and responsible for loading data from UNIX file systems to HDFS. Installed and configured Hive and also written Hive UDFs.
  • Developed work flow in Oozie to automate the tasks of loading the data into HDFS and per-processing with Pig.
  • Developed high-performance distributed queuing system, Scala, Redis, Akka, closure, MQ messaging, Json.
  • Proficient in Spark, Spark Streaming, Spark SQL and DataFrames.
  • Developed, with another consultant, prototype of mobile video app for mass consumer use. Pre-launch confidentiality precludes further detail. Scala, Android.
  • Involved in the database migrations to transfer data from one database to other and complete virtualization of many client applications.
  • Developed multiple POCs using Scala and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL/Tera data.
  • Supports and assist QA Engineers in understanding, transaction or not.
  • Willing to take up challenges and interested to work on different kinds of applications and emerging technologies with in the IBM Mainframe world.
  • Analyzed data using Hadoop components Hive and Pig and created tables in hive for the end users.
  • Analyzed the SQL scripts and designed the solution to implement using Scala.
  • Developing reports using Tableau ETL.

Environment: HDFS, Hive, HBase, Spark, Spark-SQL, Kafka, Hive, Pig, MapReduce, Scala, Java (jdk1.6), Trifacta, MySQL, Eclipse, Tableau, MongoDB, Cassandra, Linux, ETL, Cloudera, Shell Scripting, Docker, Cheff, Puppet, SQL Developer.

Confidential . Charlotte, NC

Hadoop Developer

Responsibilities:

  • Responsible for gathering all required information and requirements for the project.
  • Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
  • Real time streaming the data using Spark with Kafka.
  • Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS using Scale.
  • Performed analysis on implementing Spark using Scala and wrote spark sample programs using PySpark.
  • Worked on debugging, performance tuning of Hive & Pig Jobs.
  • Involved in loading data from LINUX file system to HDFS. Importing and exporting data into HDFS using Sqoop and Kafka.
  • Developed Scala scripts using both Data frames/SQL/Data sets and RDD/MapReduce in Spark for Data Aggregation, queries and writing data back into OLTP system through Sqoop.
  • Experience working on processing unstructured data using Spark and Hive.
  • Involved in writing custom Pig Loaders and Storage classes to work with a variety of data formats such as JSON, Compressed CSV, etc.
  • Launched 3 new products (Knox, Spark, Zeppelin) in Hortonworks data platform to gain and extend market share.
  • Automated all the jobs for pulling data from FTP server to load data into Hive tables using Oozie workflows.
  • Gained knowledge in NoSQL database with Cassandra and MongoDB.
  • Experience in Agile Programming and accomplishing the tasks to meet deadlines.
  • Exported the result set from Hive to MySQL using Shell scripts.
  • Actively involved in code review and bug fixing for improving the performance.

Environment: Hadoop, HDFS, Flume, Hive, Pig, Scoop, Spark, Kafka, JSON, Map Reduce, HBase, HortonWorks, Scala, Oozie, Cassandra, MongoDB, MySQL, Docker, Puppet, Cheff, GitHub, LINUX, JSP, CSS, JavaScript, spring, Java and XML

Confidential, Alabama

Jr. Big Data Developer

Responsibilities:

  • Installed, configured, upgraded, and applied patches and bug fixes for Prod, Lab and Dev Servers.
  • Installed/Configured/Maintained Hadoop clusters in Dev/Lab/Pre-Prod/Prod environments.
  • Install, configure and administer HDFS, Hive, Pig, HBase Sqoop, Spark and Yarn.
  • Involved in upgrading Cloudera Manager Upgrade from Cloudera Manager 5.5 to Cloudera Manager 5.6.
  • Worked on analyzing Hadoop cluster using different big data analytic tools including Pig, Hive, and MapReduce.
  • Involved in loading data from LINUX file system to HDFS.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Worked on processing unstructured data using Pig and Hive.
  • Performed Hadoop streaming jobs to process terabytes of xml format data.
  • Involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs.
  • Developed Pig Latin scripts to extract data from the web server output files to load into HDFS.
  • Extensively used Pig for data cleansing.
  • Created and maintained Technical documentation for launching HADOOP Clusters and for executing Hive queries and Pig Scripts.
  • Involved in setting up alerts in Cloudera Manager for the monitoring health and performance of Hadoop Clusters.
  • Actively involved in code review and bug fixing for improving the performance.

Environment: Hadoop, HDFS, Pig, Hive, MapReduce, Sqoop, Yarn, Linux, Cloudera, Big Data, Java APIs, Java collection, JSP Servlets, SQL.

Confidential

Java Developer

Responsibilities:

  • Involved in the complete SDLC software development life cycle of the application from requirement analysis to testing.
  • Developed the modules based on struts MVC Architecture.
  • Developed The UI using JavaScript, JSP, HTML, and CSS for interactive cross browser functionality and complex user interface.
  • Created Business Logic using Servlets, Session beans and deployed them on WebLogic server.
  • Used MVC struts framework for application design.
  • Developed EJBs, JSPs and Java Components for the application using Eclipse.
  • Prepared the Unit test plans and the integrated test plans.
  • Created complex SQL Queries, PL/SQL Stored procedures, Functions for back end.
  • Prepared the Functional, Design and Test case specifications.
  • Involved in writing Stored Procedures in Oracle to do some database side validations.
  • Performed unit testing, system testing and integration testing
  • Developed Unit Test Cases. Used JUnit for unit testing of the application.
  • Provided Technical support for production environments resolving the issues, analyzing the defects, providing and implementing the solution defects. Resolved more priority defects as per the schedule.
  • Used Eclipse IDE for all coding in Java, Servlets and JSPs.
  • Used Flex Styles and CSS to manage the Look and Feel of the application.

Environment: core Java, J2EE, JDBC, Java 1.4, Servlets, JSP, Struts, Hibernate, Web services, RESTful services, SOAP, WSDL, Design Patterns, MVC, HTML, JavaScript 1.2, WebLogic 8.0, XML, Junit, Oracle 10g, My Eclipse.

Confidential

Jr. Java Developer

Responsibilities:

  • Analyzing and preparing the requirement Analysis Document.
  • Deploying the Application to the JBOSS Application Server.
  • Implemented Web Service using SOAP protocol using Apache Axis.
  • Requirement gatherings from various parties involved in the project
  • Study OAuth/JWT/TOTP/SAML series protocol for SSO solution
  • Used to J2EE and EJB to handle the business flow and Functionality.
  • Involved in the complete SDLC of the Development with full system dependency.
  • Actively coordinated with deployment manager for application production launch.
  • Provide Support and update for the period under warranty.
  • Monitoring of test cases to verify actual results against expected results.
  • Carrying out Regression testing to track the problem tracking.

Environment: Java, J2EE, EJB, UNIX, XML, Work Flow, JMS, JIRA, Oracle, JBOSS, Soap

We'd love your feedback!