We provide IT Staff Augmentation Services!

 spark/scala Developer Resume

5.00/5 (Submit Your Rating)

Auburn Hills, MI

SUMMARY:

  • Adept at implementing E2E solutions on Big Data using Hadoop frame work, executed and designed big data solutions on multiple distribution systems like Cloudera (CDH3 & CDH4 ), Hortonworks.
  • Strong knowledge of HDFS Architecture and its basic components such as Name node, Data node, Job tracker, Task tracker and Map Reduce concepts.
  • Extensively worked on Implementing and optimizing Hadoop/MapReduce algorithms for Big Data analytics.
  • Experienced in writing Map Reduce programs using Java to perform transformations on different data sets using Maper and Reducer tasks.
  • Worked with join patterns and implemented Maper side joins and Reducer side joins using Map Reduce.
  • Hands on experience in Sequence files, RC files, Combiners, Counters, Dynamic Partitions, Bucketing for best practice and performance improvement.
  • Developed multiple MapReduce jobs to perform data cleaning and preprocessing.
  • Implemented Ad - hoc query using Hive to perform analytics on structured data and designed HIVE queries, Pig scripts to perform data analysis, data transfer and table design.
  • Strong knowledge in writing Hive UDF, Generic UDF's to in corporate complex business logic into Hive Queries.
  • Experienced in optimizing Hive queries by tuning configuration parameters.
  • Involved in designing the data model in Hive for migrating the ETL process into Hadoop and wrote Pig Scripts to load data into Hadoop environment.
  • Having experience in developing a data pipeline using Kafka to store data into HDFS.
  • Experienced in using Apache Flume to collect the logs and error messages across the cluster.
  • Implemented SQOOP for transferring large dataset transfer onto HDFS from RDBMS.
  • Used Cassandra CQL with Java API’s to retrieve data from Cassandra tables.
  • Expertise in performing real time analytics on HDFS using HBase.
  • Experience in composing shell scripts to dump the shared information from MySQL servers to HDFS.
  • Worked with Oozie and Zoo-keeper to coordinate flow of jobs in the cluster.
  • Good knowledge in writing Spark application using Python and Scala.
  • Implemented pre-defined operators in spark such as map, flat Map, filter, reduceByKey, groupByKey, aggregateByKey and combineByKey etc.
  • Used Scala SBT to develop Scala coded spark projects and executed using spark-submit.
  • Experience in understanding the security requirements for Hadoop and integrated with Kerberos.
  • Worked with different file formats (AVRO, JSON, ORC, Parquet) and has the knowledge about data compression techniques (LZO, Bzip2 and Snappy).
  • Worked on Talend Open Studio and Talend Integration Suite.
  • Experience in Building Web-based, Enterprise level and standalone application using JSP, Struts, Spring, Hibernate, JSF, Restful Web services.
  • Familiarity working with popular frameworks likes Struts, Hibernate, Spring MVC and AJAX.
  • Good understanding of XML methodologies (XML, XSL, XSD) including Web Services and SOAP.
  • Good in using version control like GITHUB and SVN and hands-on building tool Maven and continuous integration like Jenkins.
  • Good understanding of all aspects of Testing such as Unit, Regression, Agile, White-box, Black-box.
  • Adept knowledge and working experience in Agile, waterfall and Test-Driven Development (TDD) methodologies.
  • Ability to work with onsite and offshore team members.
  • Able to work on own initiative, highly proactive, self-motivated commitment towards work and resourceful.
  • Strong debugging and critical thinking ability with good understanding of frameworks advancement in methodologies and strategies.

TECHNICAL SKILLS:

Big Data Ecosystems: Hadoop, MapReduce, HDFS, Zookeeper, Hive,Pig, Sqoop, Oozie, Flume, Yarn, Spark

Database Languages: SQL, PL/SQL, Oracle

Programming Languages: Java, Scala

Frameworks: Spring, Hibernate, JMS

Scripting Languages: JSP, Servlets, JavaScript, XML, HTML, Python

Web Services: RESTful web services

Databases: RDBMS, HBase, Cassandra

IDE: Eclipse, IntelliJ

Platforms: Windows, Linux, Unix

Application Servers: Apache Tomcat, Web Sphere, Web logic, JBoss

Methodologies: Agile, Waterfall

ETL Tools: Talend

PROFESSIONAL EXPERIENCE:

Confidential, Auburn Hills, MI

Spark/Scala Developer

Responsibilities:

  • According to the business requirement and data researcher’s strategy, Analyzed and suggested system architecture for the cluster.
  • Developed Kafka Producers and Consumers as per the required specifications.
  • Used Kafka for log accumulation like gathering physical log documents from servers and transferred them onto HDFS for handling purpose.
  • Involved in Configuring Spark streaming to ingest real-time information from Kafka to store them onto HDFS.
  • Extracted Real time feed using Kafka and Spark Streaming and converted it to RDD and processed data into Data Frame to save the data as Parquet format in HDFS.
  • Using spark Transformations and Actions performed data cleansing on the input data.
  • Implemented Spark using Scala and Spark SQL for faster testing and processing of data.
  • Involved in developing a linear regression model to predict a customer’s behavior with the help of transactions data, developed using spark with Scala API.
  • Involved in performing the analytics and visualization for the data from the logs for anomaly detection and the probability of future occurrence’s using regressing models.
  • Worked extensively on spark and MLlib to develop a regression model for log information.
  • Experienced in using the spark application master to monitor the spark jobs and capture the logs for the spark jobs.
  • Involved in writing custom Map-Reduce programs using java API and pig Latin for data processing.
  • Developed shell scripts to generate the HIVE create statements and loaded into the table.
  • The HIVE tables are created as per requirement were Internal or External tables defined with appropriate static, dynamic partitions and bucketing, for efficiency.
  • With large sets of structured, semi structured data, developed Hive queries for the analysts.
  • Optimized Hive QL/ pig scripts by using execution engine like Tez, Spark.
  • Implemented Cassandra using DataStax Java API and has a very good understanding of Cassandra cluster mechanism including replication strategies, snitch, gossip, consistent hashing and consistency levels.
  • Used WEB HDFS REST API to make the HTTP GET, PUT, POST and DELETE requests from the webserver to perform analytics on the data lake.
  • Integrated Maven build and designed workflows to automate the build and deploy process.

Environment:: Hadoop, Hive, HDFS, HPC, WEBHDFS, WEBHCAT, Spark, Spark-SQL, KAFKA, Java, Scala, Web Server’s, Maven Build and SBT build.

Confidential, St Louis, MO

Hadoop Developer

Responsibilities:

  • Developed MapReduce programs to parse and filter the raw data and partitioned tables in Greenplum.
  • Created Hive queries that helped market analysts spot emerging trends by comparing incremental data with Greenplum reference tables and historical metrics.
  • Responsible for creating Hive tables, loading the structured data resulted from MapReduce jobs into the tables and writing hive queries to further analyze the logs to identify issues and behavioral patterns.
  • Involved in running MapReduce jobs for processing millions of records.
  • Built reusable Hive UDF libraries for business requirements which enabled users to use these UDF's in Hive Querying.
  • Developed a data pipeline using Flume for real-time data ingestion and Sqoop to transfer data into HDFS on regular basis.
  • Experienced in migrating Hive QL into Impala to minimize query response time.
  • Responsible for Data Modeling as per our requirement in HBase and for managing and scheduling Jobs on a Hadoop cluster using Oozie jobs.
  • Installed and configured Hadoop MapReduce and developed multiple MapReduce jobs in java for data cleansing and preprocessing.
  • Worked on Spark SQL and Data frames for faster execution of Hive queries using Spark Sql Context.
  • Performed analysis on implementing Spark using Scala and wrote spark sample programs using PySpark.
  • Used Solr Search & MongoDB for querying and storing data.
  • Created UDFs to calculate the pending payment for the given residential or small business customers quotation data and used in Pig and Hive Scripts.
  • Deployed and built the application using Maven.
  • Experience in managing and reviewing Hadoop log files.
  • Experienced in moving data from Hive tables into HBase for real time analytics on Hive tables.
  • Handled importing of data from various data sources, performed transformations using Hive. (External tables, partitioning).
  • Responsible for data modeling in MongoDB to load data which is coming as structured and unstructured data.
  • Unstructured files like XML's, JSON files are processed using custom built Java API and pushed into MongoDB.
  • Wrote test cases in MR unit for unit testing of Mapreduce Programs.
  • Extensively worked on User Interface for few modules using JSPs, JavaScript and Ajax
  • Created Business Logic using Servlets, Session beans and deployed them on Web logic server.
  • Involved in templates and screens in HTML and JavaScript.
  • Developed the XML Schema and Web services for the data maintenance and structures.
  • Provided Technical support for production environments resolving the issues, analyzing the defects, providing and implementing the solution defects
  • Built and deployed Java applications into multiple Unix based environments and produced both unit and functional test results along with release notes.

Environment: HDFS, MapReduce, Hive, Pig, Cloudera, Impala, Oozie, Greenplum, MongoDB, HBase, Sqoop, Flume, Maven, Python, Cloud Manager, JDK, J2EE, Struts, JSP, Servlets, Solr, WebSphere, HTML, XML, JavaScript, MR unit.

Confidential, San Francisco, CA

Hadoop Developer

Responsibilities:

  • Developed Pig Scripts, Pig UDFs and Hive Scripts, Hive UDFs to analyze HDFS data.
  • Used Sqoop to export data from HDFS to RDBMS.
  • Having experience on Hadoop eco system components HDFS, MapReduce, Hive, Pig, Sqoop and HBase.
  • Expertise with web based GUI architecture and development using HTML, CSS, AJAX, JQuery, AngularJS, and JavaScript.
  • Developed Map Reduce programs for some refined queries on big data.
  • Involved in loading data from UNIX file system to HDFS.
  • Used Pig as ETL tool to do transformations, event joins and some pre-aggregations before storing the data onto HDFS.
  • Extracted the data from Databases into HDFS using Sqoop
  • Handled importing of data from various data sources, performed transformations using Hive, PIG and loaded data into HDFS.
  • Used PIG predefined functions to convert the fixed width file to delimited file.
  • Used HIVE join queries to join multiple tables of a source system and load them into Elastic Search Tables.
  • Manage and review Hadoop log files. Implemented lambda architecture as a solution.
  • Involved in analysis, design, testing phases and responsible for documenting technical specifications.
  • Adept at understanding Partitions, bucketing concepts managed and created external tables in Hive to optimize performance.
  • Expertise in writing Hadoop Jobs for analyzing data using HiveQL (Queries), Pig Latin (Data flow language), and custom MapReduce programs in Java.
  • Experienced in running Hadoop streaming jobs to process terabytes data in Hive.
  • Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
  • Created reports for the BI team using Sqoop to import data into HDFS and Hive.

Environment: CDH, Hadoop, HDFS, MapReduce, Yarn, Hive, PIG, Oozie, Sqoop, Linux, Shell scripting, Java, SBT, Amazon S3, JIRA, Git Stash, HDFS, Eclipse, SQL, Oracle 11g.

Confidential, Pasadena, CA

Jr. Big Data Developer

Responsibilities:

  • Worked on analyzing Hadoop cluster using different big data analytic tools including Pig, Hive, and MapReduce.
  • Involved in loading data from LINUX file system to HDFS.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Worked on processing unstructured data using Pig and Hive.
  • Performed Hadoop streaming jobs to process terabytes of xml format data.
  • Involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs.
  • Developed Pig Latin scripts to extract data from the web server output files to load into HDFS.
  • Extensively used Pig for data cleansing.
  • Implemented SQL, PL/SQL Stored Procedures.
  • Worked on debugging, performance tuning of Hive & Pig Jobs.
  • Implemented test scripts to support test driven development and continuous integration.
  • Worked on tuning the performance of Pig queries.
  • Created and maintained Technical documentation for launching HADOOP Clusters and for executing Hive queries and Pig Scripts.
  • Actively involved in code review and bug fixing for improving the performance.

Environment: Hadoop, HDFS, Pig, Hive, MapReduce, Sqoop, LINUX, Cloudera, Big Data, Java APIs, Java collection, SQL.

Confidential

Java Developer

Responsibilities:

  • Played an active role in the team by interacting with welfare business analyst/program specialists and converted business requirements into system requirements.
  • Developed and deployed UI layer logics of sites using JSP.
  • Struts (MVC) is used for implementation of business model logic.
  • Worked with Struts MVC objects like Action Servlet, Controllers, and validators, Web Application Context, Handler Mapping, Message Resource Bundles and JNDI for look-up for J2EE components.
  • Developed dynamic JSP pages with Struts.
  • Developed the XML data object to generate the PDF documents and other reports.
  • Used Hibernate, DAO, and JDBC for data retrieval and medications from database.
  • Messaging and interaction of Web Services is done using SOAP and REST
  • Developed JUnit Test cases for Unit Test cases and as well as System and User test scenarios
  • Worked with Restful web services to enable interoperability.

Environment: core Java, J2EE, JDBC, Java 1.4, Servlets, JSP, Struts, Hibernate, Web services, RESTful services, SOAP, WSDL, Design Patterns, MVC, HTML, JavaScript 1.2, WebLogic 8.0, XML, Junit, Oracle 10g, My Eclipse.

Confidential,

Jr. Java Developer

Responsibilities:

  • Analyzing and preparing the requirement Analysis Document.
  • Deploying the Application to the JBOSS Application Server.
  • Implemented Web Service using SOAP protocol using Apache Axis.
  • Requirement gatherings from various parties involved in the project
  • Study OAuth/JWT/TOTP/SAML series protocol for SSO solution
  • Used to J2EE and EJB to handle the business flow and Functionality.
  • Involved in the complete SDLC of the Development with full system dependency.
  • Actively coordinated with deployment manager for application production launch.
  • Provide Support and update for the period under warranty.
  • Monitoring of test cases to verify actual results against expected results.
  • Carrying out Regression testing to track the problem tracking.

Environment: Java, J2EE, EJB, UNIX, XML, Work Flow, JMS, JIRA, Oracle, JBOSS, Soap

We'd love your feedback!