We provide IT Staff Augmentation Services!

Hadoop/spark Developer Resume

3.00/5 (Submit Your Rating)

Santa Maria, CA

PROFESSIONAL SUMMARY

  • 8 years of professional IT experience in analysis, design, architecture, development, testing and implementation of Hadoop, Bigdata Technologies, Data Warehousing, and AWSon Object Oriented Programming.
  • Expertise in Hadoop, MapReduce, YARN, Hive, Pig, Sqoop, Kafka, Spark - Scala, and Spark Streaming.
  • Hands on experience with AWS components like EC2, S3, Data Pipeline, RDS, RedShift and EMR.
  • In depth knowledge on Hadoop Architecture and various components such as HDFS, Job Tracker, Name node, Data Node and MapReduce concepts. 
  • Good understanding of HDFS Designs, Daemons, HDFS High Availability (HA).
  • Good understanding and working experience on Hadoop Distributions like Cloudera.
  • Worked on NoSQL databases like MongoDB,HBase,Cassandra
  • Experience in analyzing data using HiveQL,Pig Latin and custom MapReduce programs in Java.
  • Having Good Working Expertise on handling Multi Terabytes of structured and unstructured data on significantly big Cluster Environment
  • Hands on experience in working with Flume to load the log data from multiple sources directly into HDFS.
  • Experience in analyzing large scale data to identify new analytics, insights, trends, and relationships with a strong focus on data clustering.
  • Have good knowledge on Data Center OS Mesos, and Docker.
  • Hands on experience using business intelligence tools like Tableau, Cognos and Qlik View.
  • Proficient in developing web based applications and client server distributed architecture applications in Java/J2EE technologies using Object Oriented Methodology.
  • Strong Knowledge on full Software Development life cycle -Software analysis, design, architecture, development, and maintenance.
  • Worked on relative ease with different working strategies like Agile, Waterfall, Scrum, and Test-Driven Development(TDD) methodologies.
  • Excellent experience in designing and developing Enterprise Applications for J2EE platform using Servlets, JSP, Struts, Spring, Hibernate and Web services.
  • Expertise in developing web services with XML based protocols such as SOAP and WSDL.
  • Learning new technologies will allow for more effective design and implementation.  
  • Intellectual capacity to grasp new situations quickly and problem-solving skills. 
  • Excellent communication, interpersonal skills, proactive and a good team player along with a can-do attitude and good written skills.

TECHNICAL SKILLS

Big Data Technologies: HDFS, Hive, Map Reduce, Pig, Sqoop, Oozie, Zookeeper, YARN, Spark.

Scripting Languages: Shell, Python, Scala.

Programming Languages: CoreJava, C++, C, SQL.

Front End Technologies: HTML, XHTML, CSS, XML, JavaScript, AJAX, Servlets, JSP.

Web Services: AWS EC2, S3, Data Pipeline, RDS, RedShift, EMR, Dynamo DB, SOAP, and Rest.

Application Servers: Apache Tomcat, WebLogic Server, WebSphere, JBoss.

Databases: Oracle 11g, MySQL, IBM DB2

DW & BI Tools: DataStage, Tableau, Cognos, Qlik View, and Qlik Sense.

NoSQL Databases: HBase, MongoDB, and Cassandra.

IDE: Eclipse, NetBeans, JBuilder.

Operating Systems: Linux, UNIX, MAC, Windows NT / 98 /2000/ XP / Vista, Windows 7, Windows 8.

PROFESSIONAL EXPERIENCE

Confidential, Santa Maria, CA

Hadoop/Spark Developer

Responsibilities

  • Prepared an ETL framework with the help of sqoop, pig and hive to be able to frequently bring in data from the source and make it available for consumption.
  • Worked on importing and exporting data from Oracle and DB2 into HDFSusing Sqoop. 
  • Developed data pipeline using Flume to ingest customer behavioral data and financial histories into HDFS for analysis. 
  • Scheduled a workflow to import the weekly transactions in the revenue department from RDBMS database using Oozie.
  • Built wrapper shell scripts to hold theseOozieworkflow. 
  • Developed PIG Latin scripts to transform the log data files and load processed datasets into HDFS. 
  • Used Pig as ETL tool to do transformations, event joins and some pre-aggregations before storing the data onto HDFS. 
  • Hands on experience with NoSQL databases like Cassandra for POC (proof of concept) in storing URL's and images. 
  • Developedhive UDF for functions that were not preexisting in Hive like the rank etc. 
  • Created ExternalHive tables and involved in data loading and writing Hive UDFs. 
  • Migrated iterative map reduce programs into Spark transformations using Scala.
  • Experienced in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
  • Optimizing of existing word2vec algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames, and Pair RDD's in development of Chatbot using OpenNLP and Word2Vec.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself.
  • Created concurrent access for hive tables with shared and exclusive locking that can be enabled in hive with the help of Zookeeper implementation in the cluster.
  • Build Tez source code and configured on Hive and achieved very good responsive time (<1 min) while running the huge Hive queries which used to take longer time (> 30 mins)
  • Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
  • Wrote the shell scripts to monitor the health check of Hadoop daemon services and respond accordingly to any warning or failure conditions. 
  • Developed Unit test cases using MRunit for map reduce code.

Environment:Hadoop, HDFS, MapReduce, Hive, Pig, HBase, Sqoop, Spark, Oozie, Zookeeper, RDBMS/DB, MySQL, CSV, AVRO data files.

Confidential, Atlanta, GA

Hadoop Developer

Responsibilities:

  • Responsible for loading the customer's data and event logs from Oracle database, Teradata into HDFS using Sqoop
  • Involved in initiating and successfully completing Proof of Concept on SQOOP for Pre-Processing, Increased Reliability and Ease of Scalability over traditional Oracle database.
  • End-to-end performance tuning of Hadoop clusters and Hadoop MapReduce routines against very large data sets.
  • Developed the Pig UDF'S to pre-process the data for analysis.
  • Involved in loading data from LINUX file system to HDFS.
  • Importing and exporting data into HDFS and Hive using Sqoop and Flume.
  • Proficient in using Cloudera Manager, an end to end tool to manage Hadoop operations.
  • Developed MapReduce jobs for Log Analysis, Recommendation and Analytics.
  • Wrote MapReduce jobs to generate reports for the number of activities created on a particular day, during a dumped from the multiple sources and the output was written back to HDFS
  • Reviewed the HDFS usage and system design for future scalability and fault-tolerance.
  • Installed and configured Hadoop HDFS, MapReduce, Pig, Hive, and Sqoop.
  • Wrote Pig Scripts to generate MapReduce jobs and performed ETL procedures on the data in HDFS.
  • Exported analyzed data to HDFS using Sqoop for generating reports.
  • Used MapReduce and Sqoop to load, aggregate, store and analyze web log data from different web servers.
  • Developed Hive queries for the analysts.
  • Cluster co-ordination services through Zookeeper.
  • Written the Spouts and Bolts after collecting the real stream customer data from Kafka broker to process and store into HBASE.
  • Analyze the log files and process through Flume
  • Experience in optimization of MapReduce algorithm using combiners and partitions to deliver the best results and worked on Application performance optimization.

Environment: CDH4, MapReduce, HDFS, Hive, Pig, Sqoop, Linux, XML, MySQL, MySQL Workbench,PL/SQL, SQL connector

Confidential, Chicago, IL

Hadoop/ETL Developer

Responsibilities:

  • Responsible for managing data from multiple sources.
  • Involved in the migration part of the project for 30 sources.
  • Worked with IBM data extraction application Data stage for ETL purpose to get the data on Edge Node.
  • Developed CRON jobs to write the input data files to HDFS location and Archive location.
  • Developed Map Reduce applications for the schema validation and Row Count Validation.
  • Assisted in exporting analyzed data to relational databases using Sqoop.
  • Experienced in importing and exporting data into HDFS and assisted in exporting analyzed data to RDBMS using SQOOP.
  • Once Schema and Row Count Validation is done, written MR jobs to create Avro Schema.
  • Developed Applications to convert .dat files to Avro data format.
  • Written MR jobs to create super set schema from different Avro schemas.
  • Hive tables have been created from the super set schema.
  • To run multiple MapReduce and Hive jobs installed and used Oozie Workflow engine.
  • Participated in white board sessions to get the task requirements.
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
  • Extracted the data from Teradata into HDFS using Sqoop.
  • Analyzed the data by performing Hive queries and running Pig scripts to know user behavior like shopping enthusiasts, travelers, music lovers etc.
  • Wrote REST Web services to expose the business methods to external services.
  • Exported the patterns analyzed back into Teradata using Sqoop.
  • Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
  • Installed Oozie workflow engine to run multiple Hive.

Environment: Hadoop, MapReduce, HDFS, Hive, Flume, Sqoop, Cloudera, Oozie, Data Stage, Java Cron Jobs, UNIX Scripts.

Confidential

Java/J2EEDeveloper

Responsibilities:

  • Involved in Analysis, Designing, Development and Testing phases of the application. 
  • Was involved in creation and maintenance of the backend services using Spring, Hibernate, SQLServer and Oracle.
  • Developed Web pages using JSPs with Tag libraries, HTML, and JavaScript. 
  • Writing J2EE code using Spring, hibernate to upload input CSV files for credit risk data.
  •  Implemented Dependency Injection (IOC) feature of spring framework to inject dependency into objects and AOP is used for Logging.
  • Designed and developed persistence layer build on ORM framework and developed it using Hibernate
  • Implemented various Design patterns like Business Delegate, Data Transfer Objects DTO, Service locator, Session Facade and Data Access Objects DAO patterns. 
  • Involved in writing SQL, Stored procedure, and PL/SQL for back end. Used Views and Functions at the Oracle Database end.
  • Developed various documents within the application using XML by using Eclipse as IDE tool.
  • Developed SOAP requests to interact with billing schedule system. 
  • Used Web Services (SOAP & WSDL) to exchange data between Server part and Comercia Bank. 
  • Integrating and deploying the application on WebLogic application server using ANT.
  • Developed user interfaces for presenting the expense reports, transaction details using JSP, XML, HTML, and Java Script.
  • Used Log4J for logging the application exceptions and debugging statements. 
  • Proficient in doing Object Oriented Design using UML-Rational Rose

Environment: Java, JSP, Servlets, Web Sphere Application Server, Eclipse, Java Script, Web Services (SOAP & WSDL), Microsoft VSS, Oracle, PL/SQL and JDBC.

Confidential

Java/J2EEDeveloper

Responsibilities:

  • Utilized the base UML methodologies and Use cases modeled by architects to develop the front-end interface. The class, sequence and state diagrams were developed using Rational Rose and Microsoft Visio.
  • Designed application using MVC design pattern.
  • Developed front-end user interface modules by using HTML, XML, Java AWT, and Swing.
  • Front-end validations of user requests carried out using Java Script.
  • Designed and developed the interacting JSPs and Servlets for modules like User Authentication and Summary Display.
  • Designed and developed Entity/Session EJB components for the primary modules.
  • Java Mail was used to notify the user of the status and completion of the request.
  • Developed Stored Procedures on Oracle 8i.
  • Implemented Queries using SQL (database triggers and functions).
  • JDBC was used to interface the web-tier components on the J2EE server with the relational database.
  • Performed client-side validations using JavaScript and AJAX.
  • Coded the JUnit test cases and suite for the system.
  • Project coordination with other Development teams, System managers and web master and developed good working environment

Environment: Java 1.5, Servlets, J2EE 1.4, JDBC, Oracle, 10g, PL SQL, HTML, JSP, Eclipse, UNIX

We'd love your feedback!