Spark/hadoop Developer Resume
Long Island, NY
SUMMARY:
- Around 4 years of experience in IT industry, played major role in implementing, developing and maintenance of various Web Based applications using Java and Big Data Ecosystem.
- Over 3 years of strong end to end experience in Hadoop development using different Big Data tools.
- Strong knowledge of Hadoop Architecture and Daemons such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce concepts.
- Expertise in importing and exporting data into HDFS and hive using Sqoop and vice versa.
- Experience in writing Map Reduce programs in Java.
- Experienced in optimizing Hive Queries by tuning configuration parameters.
- Involved in designing the data model in Hive for migrating the ETL process into Hadoop and wrote PIG Scripts to load data into Hadoop Environment.
- Expert in data cleansing operations using Pig Latin transformations.
- Hands on experience in NOSQL databases like HBase, Cassandra, MongoDB.
- Used Cassandra CQL with Java API to retrieve data from Cassandra tables.
- Experience in developing spark jobs using Scala.
- Good Knowledge in Java, Unix shell scripting, Linux, JIRA, SQL Developer.
- Extensive knowledge on Data ingestion, data processing, Batch analytics.
- Experience in indexing data using Lucence libraries in Apache Solr.
TECHNICAL SKILLS:
Big Data Ecosystems: Hadoop, Map Reduce, HDFS, Zookeeper, Hive, Pig, Sqoop, Oozie, Flume, Yarn,Spark.
DB Languages: SQL,PL/SQL, Oracle
Programming Languages: Java, C, and Scala.
Frameworks: Spring.
Scripting Languages: JSP & Servlets, JavaScript, Python
Web Services: Restful
Databases: RDBMS, HBase, Cassandra,MongoDB
Tools: Eclipse, Net Beans.
Platforms: Windows, Linux, Unix
Application Servers: Apache Tomcat.
Methodologies: Agile, Waterfall
PROFESSIONAL EXPERIENCE:
Confidential, Long Island, NY
Spark/Hadoop Developer
Responsibilities:
- Involved in building data pipelines that extract, classify, merge and deliver new insights on the data.
- Working with Talend Open Studio to extract the data from Marketo.
- Experienced in using springml package to pull the data from salesforce and store in Azure blob containers.
- Experienced in doing the restful calls using scala and load the data into spark dataframes.
- Involved in spinning up the cluster and Azure SQLDB in Azure cloud environment.
- Experienced in deploying the spark applications on Microservices.
- Experienced in using AWS lamda and AWS API Gateway to perform the API calls between the Microservices and Data Store.
- Experienced in spinning up the HDInsight cluster from the available resources on the Azure.
- Used spark to populate hive tables and pull the data from blob storage and perform various transformation on the data and push it to the CDL.
Environment:: Hadoop, Hive, HDFS, Azure, sql, salesforce, spark, spark - sql, marketo, scala, webserver’s, Maven Build and SBT build.
Confidential, Phoenix, AZ
Spark/Hadoop Developer
Responsibilities:- Analyze and define researcher’s strategy and determine system architecture and requirement to achieve goals.
- Formulate strategic plans for component development to sustain future project objectives.
- Used various spark transformations like map, reducebykey, filter to clean the input data.
- Involved in writing custom Map-Reduce programs using java API for data processing.
- Integrated Maven build and designed workflows to automate the build and deploy process.
- Involved in developing a linear regression model to predict a continuous measurement of for an observation of particular patient this is developed using spark with scala API.
- Worked extensively on spark and MLlib to develop a regression model for cancer data.
- The hive tables are created as per requirement were Internal or External tables defined with appropriate static, dynamic partitions and bucketing, intended for efficiency.
- Load and transform large sets of structured, semi structured data using hive.
- Involved in setting up the pipeline for Bigdata genomics for different variant calling techniques using spark and scala.
- Used spark and spark-sql to read the parquet data and create the tables in hive using the scala API.
- Involved in making code changes for a module in Tumor simulator for processing across the cluster using spark-submit.
- Involved in performing the analytics and visualization for the data from the tumor simulator to find the number of sub clones, population etc.
- Used D3.js to create phylogenetic tree and radial tree from the json data that is generated from the hive queries.
- Used WEBHDFS rest api to make the HTTP get, put, post and delete requests from the webserver to perform analytics on the data lake.
- Worked on a POC to perform sentiment analysis of twitter data using spark-streaming.
- Worked on high performance computing (HPC) to simulate tools required for the genomics pipeline.
- Involved in setting up a test website on the webserver to execute various tools from the high performing computers and Hadoop cluster.
Environment:: Hadoop, Hive, HDFS, HPC, WEBHDFS, WEBHCAT, spark, spark-sql, java, scala, webserver’s, Maven Build and SBT build.
Confidential, Phoenix, AZ
Hadoop Developer
Responsibilities:- Involved in the architecture of the project.
- Experienced in decommissioning the legacy systems.
- Worked extensively on HIVE, SQOOP, SHELL, PIG and PYTHON.
- Used SQOOP to move the structured data from AS400DB2, Oracle and SQL.
- Used AXWAY to FTP the OPTIM files to move to Hadoop and create tables on top of the data.
- Scheduled CRON JOB to schedule the shell scripts.
- Experienced in handling VSAM files in mainframe to move them to Hadoop.
- Used INFORMATICA to handle Occurs and Redefines of the VSAM files and loaded the data into Oracle.
- Experienced in writing HIVE JOIN Queries.
- Used PIG predefined functions to convert the fixed width file to delimited file.
- Used python to read the AVRO file.
- Developed shell scripts to perform the incremental loads.
- Used HIVE join queries to join multiple tables of a source system and load them into Elastic Search Tables.
- Experienced in moving data from MULTIMEMBER’S to Hadoop.
- Involved in data migration from one cluster to another cluster.
- Analyze Cassandra database and compare it with other open-source NoSQL databases to find which one of them better suites the current requirement.
- Implemented service layer on top of Cassandra using core Java, Datastax Java API and Restful API.
- Experienced in Cassandra database configurations.
- Used Oozie to schedule the workflows to perform shell action and hive actions.
- Experienced in writing the workflows and defining the job.properties file for Oozie.
- Experienced in managing Hadoop Jobs and logs of all the scripts.
- Experienced in Data Validations and gathering the requirements from the business.
Environment:: Hadoop, Hive, HDFS, Pig, Sqoop, Oracle10g, SQL, Linux, Mainframe, AS400DB2, Optim.
Confidential
Software Engineer
Responsibilities:- Migrating the needed data from MySQL in to HDFS using Sqoop and importing various formats of flat files into HDFS .
- Developed the Sqoop scripts in order to make the interaction between Hive and MySQL Database.
- Mainly worked on Hive queries to categorize data of different wireless applications and security systems.
- Involved in various phases of Software Development Life Cycle (SDLC) such as requirements gathering, analysis, design and development.
- Extensively used Core Java, Servlets, JSP and XML.
- Worked on Action classes, Request processor, Business Delegate, Business Objects, Service classes and JSP pages
- Developed JSP pages using Struts custom tags.
- Used SVN for source code versioning and code repository.
- Prepared use-case diagrams, class diagrams and sequence diagrams as part of requirement specification documentation.
- Used basic of dependency injection and commonly used dependency injection features of spring framework, using java configuration.
- Implemented Business components using spring core and Navigation using Spring MVC.
Environment:: Hadoop, Hive, MySQL, Sqoop, MVC, Struts, JSP, Servlets, JUnit, Apache Tomcat Server.
