We provide IT Staff Augmentation Services!

Hadoop Developer Resume

5.00/5 (Submit Your Rating)

Portland, OregoN

SUMMARY:

  • Over 6 years of working experience in system requirements, analysis, design, testing, implementation, development, estimation, gathering, of various business applications using Web Technologies like (HDFS, Map Reduce, Oracle, JDK, J2EE,Apache Kafka, Java Script, Java, Angular JS, JDBC, JSP, SERVLETS, XML, EJB, Java Beans, Hibernate, and SQL Server, and Mongo - DB)
  • Used Talend Open studio for Data Integration, Extraction and Transformation.
  • Worked with most of the JDK versions for 6 years, in almost many modules.
  • More than two years of experience in installing, configuring and testing Hadoop ecosystem components. Good understanding of Cassandra , Spark Streaming and Spark SQL with Scala .
  • Good usage of Spark with Scala and Java . Inspected Test cases, Registries, Off-sets, Replicas , nodes and several other roles in KAFKA .
  • Capable of processing large sets of structured, semi-structured, unstructured data and supporting systems application architecture.
  • Extensive working knowledge on YARN and KAFKA.
  • Extensively worked on cloud platforms like Amazon web services (AWS).
  • Excellent hands on work with Hive UDF UDAF and UDTF.
  • Extensively involved in all phases of Software development life cycle including Analysis, design, development, Implementation, testing and support
  • Regular use of authentication protocols like Kerberos and LDAP.
  • Good Hands on work with Python, and very much comfortable in operating (Mac/Windows/Linux)
  • Exposure to Waterfall and Agile Software methodologies.
  • Expert in Problem solving, excellent analytical, troubleshooting and debugging skills
  • Good domain knowledge in Supply chain, Financial and Insurance domain
  • Worked on Apache FLUME distributed service
  • Worked on implementation platforms like Mahout and NO-SQL database.
  • Used various compression techniques like LZO, G-ZIP, and Snappy.
  • Experience with house automation tool AUTOMIC.
  • Good Understanding and working experience with Talend
  • Involved in writing the Pig scripts to reduce the job execution time
  • Strong experience in Web based applications design, development and implementation.
  • Excellent team player with good communication, presentation and highly motivated.

TECHNICAL SKILLS:

Big Data Ecosystems: Hadoop, Map-Reduce, HDFS, H-Base, Hive, Pig, Zookeeper, Oozie, Flume, Sqoop, Cassandra, Chukwa apache, and Talend

Programming Languages: C/ C++, Java.

Scripting Languages: JavaScript, HTML, Python, XML, JSP & Servlets, PHP and Bash

Databases: Oracle, NoSQL

UNIX Tools: Yum, RPM, Apache, red hat Linux

Tools: Eclipse, Cloudera, Horton-works, JDeveloper, J-Probe, Net-beans CVS, Ant.

Platforms: Windows(2000/XP), Linux, Solaris, AIX

Application Servers: Apache Tomcat 5.x 6.0, J-boss 4.0

Testing Tools: WSAD, RAD

Methodologies: Agile, UML, Design Patterns

PROFESSIONAL EXPERIENCE:

Confidential, Portland, Oregon

Hadoop Developer

Responsibilities:

  • Worked on Ecosystems like Oozie, Sqoop, Spark, Kafka, Flume, Pig, H - Base, Hive, and Sqoop with CDH5.
  • Worked on Hadoop cluster during pre-production stage which ranged from 20-30 nodes and sometimes extended even more in production time.
  • Worked on Spark Streaming to run sophisticated applications on Hadoop.
  • Used advanced ETL functionalities and performed job designs, business models using Talend.
  • Extensively worked on NO-SQL databases like Cassandra.
  • Extensively worked on Apache Spark, Spark streaming using Scala programming.
  • Tuned Java garbage collectors for Apache spark applications.
  • Good knowledge on Agile Methodology and the scrum process in Teradata.
  • Created HBase tables to load huge sets of unstructured, semi-structured and structured data from NoSQL, UNIX and a variety of portfolios.
  • Worked with data pipeline using Sqoop to import customer s data and historical data from data sources such as HDFS, MySQL, Oracle and Teradata and achieved good experience on Impala.
  • Worked on several miscellaneous subjects to perform several functions using Talend.
  • Worked on customizing Talend studio and studio preferences.
  • Worked together with ZOOKEEPER and KAFKA to make Kafka talk to Zookeeper by various classes.
  • Involved in Cluster coordination services through Zookeeper and adding new nodes to an existing cluster .Solved various issues like GC overhead Exception, Out of memory exception etc.
  • Worked on Data Serialization and HIVE serialization formats, which involves converting Complex objects into sequence bits by using CSV, PARQUET, JSON and AVRO formats.
  • Good Hand s on work with Hive table s creation, data loading and other Hive queries that helped in identifying customer s historical metrics comparing his current information and several other analysts.
  • Created Hive aggregator to update the Hive table after running the data profiling job.
  • Pig Scripts for joining, grouping, sorting and filtering the data.
  • Implemented Bucketing, Dynamic and regular Partitioning in Hive.
  • Worked on data analysis in Cassandra using SQL, and supported analysis in complicated queries.
  • Wrote Teradata Macros and used various Teradata analytical functions like BTEQ, Fast Load, Multi- Load etc.
  • Exported the analyzed data to Relational databases using Sqoop to generate reports and visualization for the BI team.
  • Worked on component libraries and graphical development environment in Ab Initio.
  • Involved in Requirement gathering, Create Test Plan, Constructed and executed positive/negative test cases in-order to prompt and arrest all bugs within QA environment.
  • Used AWS Data Pipeline to schedule Amazon EMR cluster to clean and process web server logs stored in Amazon S3 bucket.
  • Performed fault tolerance and deployment in-memory cache in H-Base.
  • Used Teradata utilities fast load, multi load, tpump to load data
  • Document/Develop/Capture architectural best practices for building systems on AWS.
  • Developed ETL workflow that pushes the webserver logs to Amazon S3 bucket.
  • Reviewed and Managed Hadoop log files.
  • Improved High data processing, capability and speed ups using Ab Initio.
  • Analyzed huge data sets to determine optimal way to aggregate and report on these data sets.
  • Involved in the Big-data requirements review meetings and partnered with business analysts to clarify any specific scenarios.

Environment: Hadoop, HDFS, H-Base, Sqoop, Pig, Oozie, MapReduce, Zookeeper, Hive, Oracle Cassandra, Teradata, UNIX Shell Scripting, MySql, JAVA, AWS.

Confidential, Nashville, Tennessee

Hadoop Developer

Responsibilities:

  • Analyzed Business requirements of Big Data and transformed into Hadoop centric technologies.
  • Worked on Hadoop cluster of range 8 nodes.
  • Worked on indexes, scalability and query language supporting using Cassandra
  • Worked on Data import and export from Teradata and Oracle into HDFS and Hive using Sqoop.
  • Implemented custom UDF's for Hive to achieve comprehensive data analysis, and also created several JAVA - UDF. Good Understanding with Apache KAFKA
  • Worked on Fast export, Fast Load, Multi Load and export from Teradata and Oracle into HDFS and Hive us Sqoop
  • Developed Pig Custom UDF's for performing various levels of optimization in custom input formats.
  • Worked on streaming log data into HDFS from web servers using Flume.
  • To filter data as per requirement we implemented custom interceptors for flume.
  • To identify issues and behavioral patterns we used Hive and Pig to analyze data in HDFS.
  • For optimized performance we defined static and dynamic partitions and created internal and external Hive tables.
  • Created Hive tables to store the processed results in a tabular format.
  • Wrote, tested and implemented Teradata Fast load, Multiload and Bteq scripts, DML and DDL.
  • For running advanced analytics we developed Pig Latin scripts on the data collected.
  • Extraction, processing and analysis of data is configured on daily workflow using Oozie Scheduler.
  • Designed and implemented MapReduce-based on large-scale, parallel and relation-learning system.
  • Developed Java spring and helper classes in business layer and used Junit for its testing.
  • Worked on Shell scripts and involved in performance analysis of the application and fixed problems/suggest solutions.
  • Used Eclipse Workbench where editors, perspectives, views, wizards and many worked as rich client platform.
  • Extensively used JDBC for database transactions.
  • Involved in Interface development for applications using JSP, JavaScript and Servlets.
  • Created Utility classes and Java validation classes.
  • Actively involved in Stress Testing of existing business components using Web-Logic Application Server.
  • Created Sequence and Class diagrams by using Violet integrated with eclipse.
  • For version repository we used Rational Clear Case.
  • For validation and data lookup we developed Wrapper classes and DAO classes.
  • Used sax parser to read xml to propagate values to business validation layer.
  • Implemented web service over http to expose product/ package details to front end systems.
  • Developed full stack (back-end development from the Markup, JavaScript, and Database).
  • Involved in creating various Utility, Helper and Reusable classes which are used across all modules of application.

Environment: Hadoop, MapReduce, HDFS, Hive, Pig, Java, SQL, Sqoop, Oozie, NoSQL,: JDK- 8, J2EE, JDBC, Java 1.4,Java spring, Servlets, JSP, Web services, SOAP, WSDL, UML, MVC, HTML, JavaScript 1.2, XML, My Eclipse.

Confidential

Java/BigData Developer

Responsibilities:

  • Involved in business technical issues.
  • Involved in coding, designing, debugging, documenting and maintaining applications.
  • Good understanding of Big-Data querying tools, and experience with No-SQL Database.
  • Developed web service API’s using java and java spring.
  • Developed front end GUI with Java Server Faces.
  • Good understanding of distributed computing principles.
  • Using servlets and JSPs implemented controller layer. View layer using JSPs, EL, JSTL and custom JSP tags.
  • JDBC connections between the database and the application. Implemented HTML, CSS and JavaScript for UI’s.
  • Executed test cases manually to verify expected results.
  • Implemented various aspects at Service layer using Spring AOP.
  • Used Test Director, added test categories and test details.
  • Used Object/relational-mapping (ORM) solution, Hibernate, technique of mapping data involved in doing the GAP Analysis of the use cases and requirements.
  • Worked on user/business requirements and developed System test plans.

Environment: Servlets, Java (Jdk 1.6),JSPs, HTML, Java Beans, JavaScript, CSS, JQuery, JDBC, SQL, Windows 98,Oracle 9i/10g, JSP, Servlets, J2EE, Java 1.4, C, C++, PHP, Multi-threading, JDBC, HTML, RADWSAD.

Confidential

Java Developer

Responsibilities:

  • Created/modified shell scripts for scheduling and automating tasks.
  • Used JDBC to establish connection between the database and the application.
  • Implemented controller and view modules using Servlets and JSPs respectively.
  • Wrote unit test cases using JUnit framework.
  • Involved in designing, coding, debugging, documenting and maintaining a number of applications.
  • Created the user interface using HTML, CSS and JavaScript.

Environment: JavaScript, Java spring, JUnit, JSP, Java Beans, HTML, CSS, Oracle 9i Java, Servlets.

We'd love your feedback!