We provide IT Staff Augmentation Services!

Spark/hadoop Developer Resume

3.00/5 (Submit Your Rating)

Long Island, NY

SUMMARY:

  • Around 4 years of experience in IT industry, played major role in implementing, developing and maintenance of various Web Based applications using Java and Big Data Ecosystem.
  • Over 3 years of strong end to end experience in Hadoop development using different Big Data tools.
  • Strong knowledge of Hadoop Architecture and Daemons such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce concepts.
  • Expertise in importing and exporting data into HDFS and hive using Sqoop and vice versa.
  • Experience in writing Map Reduce programs in Java.
  • Experienced in optimizing Hive Queries by tuning configuration parameters.
  • Involved in designing the data model in Hive for migrating the ETL process into Hadoop and wrote PIG Scripts to load data into Hadoop Environment.
  • Expert in data cleansing operations using Pig Latin transformations.
  • Hands on experience in NOSQL databases like HBase, Cassandra, MongoDB.
  • Used Cassandra CQL with Java API to retrieve data from Cassandra tables.
  • Experience in developing spark jobs using Scala.
  • Good Knowledge in Java, Unix shell scripting, Linux, JIRA, SQL Developer.
  • Extensive knowledge on Data ingestion, data processing, Batch analytics.
  • Experience in indexing data using Lucence libraries in Apache Solr.

TECHNICAL SKILLS:

Big Data Ecosystems: Hadoop, Map Reduce, HDFS, Zookeeper, Hive, Pig, Sqoop, Oozie, Flume, Yarn,Spark.

DB Languages: SQL,PL/SQL, Oracle

Programming Languages: Java, C, and Scala.

Frameworks: Spring.

Scripting Languages: JSP & Servlets, JavaScript, Python

Web Services: Restful

Databases: RDBMS, HBase, Cassandra,MongoDB

Tools: Eclipse, Net Beans.

Platforms: Windows, Linux, Unix

Application Servers: Apache Tomcat.

Methodologies: Agile, Waterfall

PROFESSIONAL EXPERIENCE:

Confidential, Long Island, NY

Spark/Hadoop Developer

Responsibilities:

  • Involved in building data pipelines that extract, classify, merge and deliver new insights on the data.
  • Working with Talend Open Studio to extract the data from Marketo.
  • Experienced in using springml package to pull the data from salesforce and store in Azure blob containers.
  • Experienced in doing the restful calls using scala and load the data into spark dataframes.
  • Involved in spinning up the cluster and Azure SQLDB in Azure cloud environment.
  • Experienced in deploying the spark applications on Microservices.
  • Experienced in using AWS lamda and AWS API Gateway to perform the API calls between the Microservices and Data Store.
  • Experienced in spinning up the HDInsight cluster from the available resources on the Azure.
  • Used spark to populate hive tables and pull the data from blob storage and perform various transformation on the data and push it to the CDL.

Environment:: Hadoop, Hive, HDFS, Azure, sql, salesforce, spark, spark - sql, marketo, scala, webserver’s, Maven Build and SBT build.

Confidential, Phoenix, AZ

Spark/Hadoop Developer

Responsibilities:
  • Analyze and define researcher’s strategy and determine system architecture and requirement to achieve goals.
  • Formulate strategic plans for component development to sustain future project objectives.
  • Used various spark transformations like map, reducebykey, filter to clean the input data.
  • Involved in writing custom Map-Reduce programs using java API for data processing.
  • Integrated Maven build and designed workflows to automate the build and deploy process.
  • Involved in developing a linear regression model to predict a continuous measurement of for an observation of particular patient this is developed using spark with scala API.
  • Worked extensively on spark and MLlib to develop a regression model for cancer data.
  • The hive tables are created as per requirement were Internal or External tables defined with appropriate static, dynamic partitions and bucketing, intended for efficiency.
  • Load and transform large sets of structured, semi structured data using hive.
  • Involved in setting up the pipeline for Bigdata genomics for different variant calling techniques using spark and scala.
  • Used spark and spark-sql to read the parquet data and create the tables in hive using the scala API.
  • Involved in making code changes for a module in Tumor simulator for processing across the cluster using spark-submit.
  • Involved in performing the analytics and visualization for the data from the tumor simulator to find the number of sub clones, population etc.
  • Used D3.js to create phylogenetic tree and radial tree from the json data that is generated from the hive queries.
  • Used WEBHDFS rest api to make the HTTP get, put, post and delete requests from the webserver to perform analytics on the data lake.
  • Worked on a POC to perform sentiment analysis of twitter data using spark-streaming.
  • Worked on high performance computing (HPC) to simulate tools required for the genomics pipeline.
  • Involved in setting up a test website on the webserver to execute various tools from the high performing computers and Hadoop cluster.

Environment:: Hadoop, Hive, HDFS, HPC, WEBHDFS, WEBHCAT, spark, spark-sql, java, scala, webserver’s, Maven Build and SBT build.

Confidential, Phoenix, AZ

Hadoop Developer

Responsibilities:
  • Involved in the architecture of the project.
  • Experienced in decommissioning the legacy systems.
  • Worked extensively on HIVE, SQOOP, SHELL, PIG and PYTHON.
  • Used SQOOP to move the structured data from AS400DB2, Oracle and SQL.
  • Used AXWAY to FTP the OPTIM files to move to Hadoop and create tables on top of the data.
  • Scheduled CRON JOB to schedule the shell scripts.
  • Experienced in handling VSAM files in mainframe to move them to Hadoop.
  • Used INFORMATICA to handle Occurs and Redefines of the VSAM files and loaded the data into Oracle.
  • Experienced in writing HIVE JOIN Queries.
  • Used PIG predefined functions to convert the fixed width file to delimited file.
  • Used python to read the AVRO file.
  • Developed shell scripts to perform the incremental loads.
  • Used HIVE join queries to join multiple tables of a source system and load them into Elastic Search Tables.
  • Experienced in moving data from MULTIMEMBER’S to Hadoop.
  • Involved in data migration from one cluster to another cluster.
  • Analyze Cassandra database and compare it with other open-source NoSQL databases to find which one of them better suites the current requirement.
  • Implemented service layer on top of Cassandra using core Java, Datastax Java API and Restful API.
  • Experienced in Cassandra database configurations.
  • Used Oozie to schedule the workflows to perform shell action and hive actions.
  • Experienced in writing the workflows and defining the job.properties file for Oozie.
  • Experienced in managing Hadoop Jobs and logs of all the scripts.
  • Experienced in Data Validations and gathering the requirements from the business.

Environment:: Hadoop, Hive, HDFS, Pig, Sqoop, Oracle10g, SQL, Linux, Mainframe, AS400DB2, Optim.

Confidential

Software Engineer

Responsibilities:
  • Migrating the needed data from MySQL in to HDFS using Sqoop and importing various formats of flat files into HDFS .
  • Developed the Sqoop scripts in order to make the interaction between Hive and MySQL Database.
  • Mainly worked on Hive queries to categorize data of different wireless applications and security systems.
  • Involved in various phases of Software Development Life Cycle (SDLC) such as requirements gathering, analysis, design and development.
  • Extensively used Core Java, Servlets, JSP and XML.
  • Worked on Action classes, Request processor, Business Delegate, Business Objects, Service classes and JSP pages
  • Developed JSP pages using Struts custom tags.
  • Used SVN for source code versioning and code repository.
  • Prepared use-case diagrams, class diagrams and sequence diagrams as part of requirement specification documentation.
  • Used basic of dependency injection and commonly used dependency injection features of spring framework, using java configuration.
  • Implemented Business components using spring core and Navigation using Spring MVC.

Environment:: Hadoop, Hive, MySQL, Sqoop, MVC, Struts, JSP, Servlets, JUnit, Apache Tomcat Server.

We'd love your feedback!