We provide IT Staff Augmentation Services!

Sr.spark Developer Resume

4.00/5 (Submit Your Rating)

Washington, Dc

PROFESSIONAL SUMMARY:

  • Around 9 years of IT experience in Analysis, design, development, implementation, maintenance and support with experience in developing strategic methods for deploying big data technologies to efficiently solve Big Data processing requirement.
  • Around 4 years of experience on BIG DATA using HADOOP framework and related technologies such as HDFS, HBASE, MapReduce, HIVE, PIG, FLUME, OOZIE, SQOOP, and ZOOKEEPER.
  • Experience in data analysis using HIVE, PIG LATIN, HBASE and custom Map Reduce programs in Java.
  • Experience in writing custom UDFs in JAVA and SCALA for HIVE and PIG TO EXTEND THE FUNCTIONALITY.
  • Experience with Cloudera and Horton works distributions.
  • Around 2 year 3 months’ experience on SPARK, SCALA, DATA FRAMES and KAFKA.
  • Developed analytical components using KAFKA, SCALA, SPARK and SPARK STREAM.
  • Experience in working with Flume to load the log data from multiple sources directly into HDFS.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems (RDBMS) and from RDBMS to HDFS.
  • Involved in creating HDINSIGHT cluster in MICROSOFT AZURE PORTAL also created EVENTSHUB and AZURE SQL DATABASES.
  • Worked on a clustered Hadoop for Windows Azure using HDInsight and HORTONWORKS Data Platform for Windows.
  • Built real time pipeline for streaming data using EVENTSHUB/MICROSOFT AZURE Queue and SPARK STREAMING.
  • Spark Streaming collects this data from EVENTSHUB in near - real-time and performs necessary transformations and AGGREGATION on the fly to build the common learner data model and persists the data in AZURE DATABASE.
  • Worked on the SPARK SQL and SPARK STREAMING modules of Spark extensively and Used SCALA to write code for all Spark use cases.
  • Used DATAFRAME API in Scala for converting the distributed collection of data organized into named columns.
  • Exploring with the Spark for improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, SPARK-SQL, DATA FRAME, PAIR RDD'S and YARN.
  • Experienced in managing Hadoop Cluster using HORTONWORKS AMBARI.
  • I have experienced with WINDOWS VISUAL STUDIO, AZ copy, JUPYTER NOTEBOOK, BLOB STORAGE, PUTTYI have been experience with AWS, AZURE, EMR and S3.
  • Good knowledge in ELASTIC MAPREDUCE and setting up environments on Amazon AWS EC2 instances.
  • Implemented Hadoop based data warehouses, INTEGRATED HADOOP with ENTERPRISE DATA WAREHOUSE systems.
  • Extensive experience in Data Ingestion, In-Stream data processing, BATCH ANALYTICS and Data PERSISTENCE STRATEGY.
  • Experience in Object Oriented Analysis Design (OOAD) and development of software using UML Methodology, good knowledge of J2EE design patterns and Core Java design patterns.
  • Experience in designing both time driven and data driven automated workflows using Oozie.
  • Experience in writing UNIX shell scripts.
  • Experience working with JDK 1.7, JAVA, J2EE, JDBC, ODBC, JSP, JAVA ECLIPSE, JAVA BEANS, EJB, SERVLETS, MS SQL SERVER.
  • Experience in J2EE technologies like Struts, JSP/Servlets, and spring.
  • Good Exposure on scripting languages like JAVASCRIPT, ANGULAR JS, JQUERY and XML.
  • Debugging the code and BUILDING JAVADOC for the backend. Performing unit testing of the classes with Junit framework.
  • Experience in all stages of SDLC (Agile, Waterfall), writing Technical Design document, Development, Testing and Implementation of Enterprise level Data mart and Data warehouses.
  • Extensive experience working IN ORACLE, DB2, SQL SERVER and My SQL database.
  • Experience in software testing, JMETER, JUNIT, MOCKITO, Regression testing, defect tracking and management using Quality Center.
  • Developed unit test cases using proprietary framework which is similar to JUNIT.
  • DELIVERY ASSURANCE - QUALITY FOCUSED & PROCESS ORIENTED:
  • Ability to work in high-pressure environments delivering to and managing stakeholder expectations
  • Application of structured methods to: Project Scoping and Planning, risks, issues, schedules and deliverables.
  • Strong analytical and Problem solving skills.
  • Good Inter personnel skills and ability to work as part of a team. Exceptional ability to learn and master new technologies and to deliver outputs in short deadlines

TECHNICAL SKILLS:

Technology Windows\: Azure, Spark, Hadoop Ecosystem/J2SE/J2EE/JDK1.7,1.8 / Data base

Operating Systems: Windows Vista/XP/NT/2000/ LINUX (Ubuntu, Cent OS), UNIX

DBMS/Databases: DB2, My SQL, PL/SQL

Programming Languages: C, C++, Core Java, scala,XML, JSP/Servlets, Struts, Spring, HTML, JavaScript, jQuery, Web services, Xml.

Big Data Ecosystem: HDFS, Map Reducing, Oozie, Hive, Pig, Sqoop, Flume, Zookeeper, Kafka and Hbase.

Methodologies: Agile, Water Fall

NOSQL Databases: Hbase

Version Control Tools: SVN, CVS

ETL Tools: IBM data stage 8.1, Informatica

WORK EXPERIENCE:

SR.SPARK DEVELOPER

Confidential

RESPONSIBILITIES:

  • Developed Fully automated Configuration driven data pipeline using spark,hive,Hdfs,SqlServer,oozie and azure File storage to load the client data into Mirror databases. performed necessary Transformations and Aggregation on the fly to build the common learner data model and persists the data in HDFS.
  • Explored the usage of Spark for improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark SQL and Spark Yarn.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs and Scala.
  • Expertise On optimizing spark Jobs when dealing with Huge joins and data Skew.
  • Performed Different type of Optimizations in spark such as Broadcast - join,reparttion and kryo serial serialization.
  • Written oozie workflow to invoke the Jobs in predefined Intervals.
  • Expert in scheduling Oozie coordinator based on input data events it starts Oozie workflow when input data is available.
  • Good experience with Artificial intelligent tool JBoss Drools-Engine.
  • Developed Robust application with the help of Drools-engine In-order to execute complex rules on top of customer data.
  • Good experience on creating drools DRL files and kmodule.xml fiels.
  • Pretty Good knowledge On the Hartnworks administration and security things such as Apache Ranger,Knox Gateway,HighAvailabilty.
  • Performed Hadoop backup Strategy to take the backup of hive,hdfs,Hbase,oozie etc.
  • Used Hive to analyze the Partitioned and Bucketed data and compute various metrics for reporting.
  • Performed hive performance tuning aspects like Map join, cost based optimization and column level statistics.
  • Created logical view instead of tables in order to enhance the performance of hive queries.
  • Involved in developing Hive DDLS to create, alter and drop Hive tables.
  • Experience in creating UDF's, UDAF's for Hive and Pig.
  • Extensively used Pig for data cleansing and HIVE queries for the analysts.
  • Performed Benchmark between Hive and sparkSql .

ENVIRONMENT: AZURE, SPARK, HIVE, SPARK SQL, KAFKA, HORTON WORKS,JBOSS DROOLS, HIVE,PIG,OOZIE,HBASE,PYTHON, SCALA, MAVEN, JUPYTER NOTEBOOK, VISUAL STUDIO, UNIX SHELL SCRIPTING.

Hadoop / Spark Developer

Confidential, Washington, DC .

RESPONSIBILITIES:

  • Developed data pipeline using EVENTHUBS, SPARK, HIVE, PIG AND AZURE SQL DATABASE to ingest customer behavioral data and financial histories into HDINSIGHT cluster for analysis.
  • Involved in creating HDINSIGHT cluster in MICROSOFT AZURE PORTAL also created EVENTSHUB and AZURE SQL DATABASES.
  • Worked on a clustered Hadoop for Windows Azure using HDInsight and HORTONWORKS Data Platform for Windows.
  • Spark Streaming collects this data from EVENTSHUB in near-real-time and performs necessary transformations and AGGREGATION on the fly to build the common learner data model and persists the data in AZURE DATABASE.
  • Used PIG to do transformations, event joins, filter boot traffic and SOME PRE-AGGREGATIONS before storing the data onto azure database.
  • Expertise with the tools in Hadoop Ecosystem including PIG, HIVE, HDFS, YARN, OOZIE, AND ZOOKEEPER. Hadoop architecture and its components.
  • Involved in integration of Hadoop cluster with spark engine to perform BATCH and GRAPHX operations.
  • Exploring with the SPARK improving the performance and optimization of the existing algorithms in Hadoop using SPARK CONTEXT, SPARK-SQL, DATA FRAME, PAIR RDD'S, SPARK YARN.
  • I have been experienced with SPARK STREAMING to ingest data into SPARK ENGINE.
  • Import the data from different sources like EVENTHUBS, COSMOS into SPARK RDD.
  • Developed SPARK CODE using SCALA and Spark-SQL/Streaming for faster testing and processing of data.
  • Involved in converting Hive/SQL queries into SPARK TRANSFORMATIONS using Spark RDDs, and SCALA.
  • Developed multiple POCs using SCALA and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL/Teradata.
  • Worked on the SPARK SQL and SPARK STREAMING modules of Spark extensively and Used SCALA to write code for all Spark use cases.
  • Used DATAFRAME API in Scala for converting the distributed collection of data organized into named columns.
  • Involved in converting the JSON data into DATAFRAME and stored into hive tables.
  • Experienced with AZCOPY, LIVY, WINDOWS POWERSHELL and CURL to submit the spark jobs on HDINSIGHT CLUSTER.
  • Analyzed the SQL scripts and designed the solution to implement USING SCALA.
  • Developed EVENTHUBS PRODUCER application in Scala to generate events into eventhubs.
  • Analyzed the SQL scripts and designed the solution to implement using PYSPARK.
  • Used Hive to analyze the PARTITIONED AND BUCKETED data and compute various metrics for reporting.
  • Involved in developing HIVE DDLS to create, alter and drop Hive tables and storm.
  • Create scalable and high-performance web services for data tracking.
  • Involved in loading data from UNIX file system to HDFS. Installed and configured Hive and also written Hive UDFs and Cluster coordination services through Zoo Keeper.
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
  • Experienced in managing Hadoop Cluster using CLOUDERA MANAGER TOOL.
  • Involved in using HCATALOG to access Hive table metadata from Map Reduce or Pig code.
  • Computed various metrics using Java Map Reduce to calculate metrics that define user experience.
  • Involved in using SQOOP for importing and exporting data into HDFS.
  • Used Eclipse and ant to build the application. Proficient work experience with NOSQL, Monod databases also the HDFS data from Rows to Columns and Columns to Rows.
  • Involved in developing Shell scripts to orchestrate execution of all other scripts (Pig, Hive, and Map Reduce) and move the data files within and outside of HDFS.

ENVIRONMENT: AZURE HDINSIGHT, SPARK, HIVE, SPARK SQL, EVENTHUB, HORTON WORKS, SCALA IDE, PYTHON, SCALA, MAVEN, JUPYTER NOTEBOOK, VISUAL STUDIO, UNIX SHELL SCRIPTING.

Hadoop Developer

Confidential

Responsibilities:

  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Used Bash Shell Scripting, Sqoop, AVRO, Hive, Pig, Java, Map/Reduce daily to develop ETL, batch processing, and data storage functionality.
  • Used Pig to do data transformations, event joinsand some pre-aggregations before storing the data on the HDFS.
  • Exploited Hadoop MySQL-Connector to store Map Reduce results in RDBMS.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Worked on loading all tables from the reference source database schema through Sqoop.
  • Worked on designed, coded and configured server side J2EE components like JSP, AWSand JAVA.
  • Collected data from different databases(i.e. Oracle, MySQL) to Hadoop
  • Used Oozie and Zookeeper for workflow scheduling and monitoring.
  • Worked on Designing and Developing ETL Workflows using Java for processing data in HDFS/Hbase using Oozie.
  • Experienced in managing and reviewing Hadoop log files.
  • Involved in loading and transforming large sets of structured, semi structured and unstructureddata from relational databases into HDFS using Sqoop imports.
  • Working on extracting files from MySQL through Sqoop and placed in HDFS and processed.
  • Supported Map Reduce Programs those running on the cluster.
  • Cluster coordination services through Zoo Keeper.
  • Involved in loading data from UNIX file system to HDFS.
  • Created several Hive tables, loaded with data and wrote Hive Queries in order to run internally in MapReduce.
  • Developed Simple to complex MapReduce Jobs using Hive and Pig.

Environment: Apache Hadoop, AWS, Map Reduce, HDFS, Hive, Java, SQL, PIG, Zookeeper, Java (jdk1.6), Flat files, Oracle 11g/10g, MySQL, Windows NT, UNIX, Sqoop, Hive, Oozie, HBase.

JAVA Developer

Confidential

Responsibilities:

  • Involved in the analysis, design, and development and testing phases of Software Development Life Cycle (SDLC)
  • . Designed and developed framework components, involved in designing MVC pattern using Struts and spring framework.
  • Responsible for developing Use case, Class diagrams and Sequence diagrams for the modules using UML and Rational Rose.
  • Developed the Action Classes, Action Form Classes, created JSPs using Struts tag libraries and configured in Struts-config.xml, Web.xml files.
  • Involved in Deploying and Configuring applications in Web Logic Server.
  • Used SOAP for exchanging XML based messages.
  • Used Microsoft VISIO for developing Use Case Diagrams, Sequence Diagrams and Class Diagrams in the design phase.
  • Developed Custom Tags to simplify the JSP code. Designed UI screens using JSP and HTML.
  • Actively involved in designing and implementing Factory method, Singleton, MVC and Data Access Object design patterns.
  • Web services used for sending and getting data from different applications using SOAP messages. Then used DOM XML parser for data retrieval.
  • Wrote JUNIT test cases for Controller, Service and DAO layer using MOCKITO, DBUNIT.
  • Developed unit test cases using proprietary framework which is similar to JUNIT.
  • Used JUnit framework for unit testing of application and ANT to build and deploy the application on WebLogic Server.

Environment: Java, J2EE, JDK1.7,JSP, Oracle, VSAM, Eclipse, HTML, Junit, MVC, ANT, WebLogic.

We'd love your feedback!