We provide IT Staff Augmentation Services!

Hadoop Developer Resume

0/5 (Submit Your Rating)

OH

SUMMARY

  • Around seven years of experience in Big Data Analytics, Hadoop, Java, Database Administration, Software development expertise.
  • Experienced in the Hadoop ecosystem components like Hadoop Map Reduce, Cloudera, Hortonworks, HBase, Oozie, Hive, Sqoop, Pig, Flume, and Cassandra.
  • Experience in developing solutions to analyze large data sets efficiently.
  • In depth understanding/knowledge of Hadoop Architecture and various components such as HDFS, JobTracker, Task Tracker, NameNode, DataNode and MapReduce concepts.
  • Extensive hands on experience in writing complex MapReduce jobs, Pig Scripts and Hive data modeling.
  • Excellent understanding/knowledge of Hadoop Distributed system architecture and design principles.
  • Excellent understanding and knowledge of NOSQL databases likeMongoDB, HBase.
  • Experience in converting MapReduce applications to Spark.
  • Good working experience using Sqoop to import data into HDFS from RDBMS and vice - versa.
  • Good knowledge in using job scheduling and workflow designing tools like Oozie.
  • Experience in performance tuning the Hadoop cluster by gathering and analyzing the existing infrastructure.
  • Developed multiple POCs using PySpark and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL/Teradata.
  • Experience in Hadoop administration activities such as installation and configuration of clusters using Cloudera Manager.
  • Have good experience creating real time data streaming solutions using Apache Spark/Spark Streaming.
  • Experience in understanding the data and designing/Implementing the enterprise platforms like Hadoop Data lake and Huge Data warehouses.
  • Good exposure on usage of column-oriented NoSQL databases like HBase, Cassandra and MongoDB.
  • Extending Hive and Pig core functionality by writing custom UDFs.
  • Good understanding of Data Mining and Machine Learning techniques.

TECHNICAL SKILLS

Languages: Java (Core Java, Networking, Threads, Swing), XML, JavaScript, C, C++, Python

J2EE Technologies: J2EE, Java Mail API.

Web servers: Apache Tomcat Server, IBM Websphere Application Server 5.0/ 6.0, Weblogic application server, JBOSS4.x

Server Side: JSP, Servlets, EJB, JDBC

Frameworks/ Components: Spring, Spring Batch, Struts, Hibernate

Big Data: Hadoop 2.8.0/2.7.3/2.6.5 , Map Reduce, Hive, Pig, Spark, Sqoop, Oozie, HDFS, HBase, Lucene, Kafka, Zookeeper

Databases: SQL, MySQL, SQL Server, Oracle, DB2, Cassandra 2.1, Microsoft Access

OS: Windows, Linux, UNIX

Markup Languages: HTML, XML, DHTML

PROFESSIONAL EXPERIENCE

Confidential, OH

Hadoop Developer

Responsibilities:

  • Involved in loading data from Teradata, Oracle database into HDFS using Sqoop queries.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Developed multiple Map Reduce jobs in java for data cleaning and preprocessing.
  • Developed Map Reduce pipeline jobs to process the data and create necessary HFiles.
  • Used Spark-Streaming APIs to perform necessary transformations and actions on the fly for building the common learner data model which gets the data from Kafka in near real time and Persists into Cassandra.
  • Experienced in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
  • Optimizing of existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames and Pair RDD's.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself.
  • Involved in loading the created HFiles into HBase for faster access of large customer base without taking Performance hit.
  • Created HBase tables to store various data formats of PII data coming from different portfolios.
  • Involved in managing and reviewing Hadoop log files.
  • Responsible to manage data coming from different sources.
  • Involved in creating Pig tables, loading with data and writing Pig Latin queries which will run internally in Map Reduce way.
  • Documented the systems processes and procedures for future references.
  • Configured and Maintained different topologies in storm cluster and deployed them on regular basis.
  • Involved in writing Unix/Linux Shell Scripting for scheduling jobs and for writing pig scripts and hive QL.
  • Developed Scripts and automated data management from end to end and sync up between all the clusters.
  • Involved in creating Hive Tables, loading with data and writing Hive queries which will invoke and run MapReduce jobs in the backend.
  • Good experience with Talend open studio for designing ETL Jobs for Processing of data.
  • Assisted in performing unit testing of Map Reduce jobs using MRUnit.
  • Used Oozie Scheduler system to automate the pipeline workflow and orchestrate the map reduce jobs that extract the data on a timely manner.
  • Used Zookeeper for providing coordinating services to the cluster.
  • Worked with Hue GUI in scheduling jobs with ease and File browsing, Job browsing, Metastore management.
  • Used Maven for Project building and management.

Environment: Hadoop, Map Reduce, HDFS, Hive, Hue, Pig, HBase, Storm, Kafka, Horton Works, Teradata, Oracle 11g/10g, HBase, Oozie, Java (jdk1.6), UNIX, SVN and Zookeeper, Maven.

Confidential, Austin, TX

Hadoop Developer

Responsibilities:

  • Worked extensively on importing data using Sqoop and flume.
  • Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
  • Responsible for creating complex tables using hive and developing Hive queries for the analysts.
  • Created partitioned tables in Hive for best performance and faster querying.
  • Transportation of data to HBase using pig.
  • Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
  • Experience with professional software engineering practices and best practices for the full software development life cycle including coding standards, code reviews, source control management and build processes.
  • Involved in source system analysis, data analysis, data modeling to ETL
  • Written multiple MapReduce procedures to power data for extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV & other compressed file formats.
  • Handling structured and unstructured data and applying ETL processes.
  • Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS
  • Developed the Pig UDF'S to pre-process the data for analysis.
  • Involved in loading and transforming large sets of Structured, Semi-Structured and Unstructured data and analyzed them by running Hive queries and Pig scripts
  • Assisted in Cluster maintenance, Cluster Monitoring and Troubleshooting, Manage and review data backups and log files

Environment: Hadoop, Apache, Sqoop, Hive, Oozie, Java (jdk1.6), Flat files, Oracle 11g/10g, MySQL, Windows NT, UNIX, Zoo Keeper, Cloudera, FLUME, CentOS, Maven.

Confidential, Baskin Ridge, NJ

Hadoop Developer

Responsibilities:

  • Worked extensively in creating MapReduce jobs to power data for search and aggregation
  • Designed a data warehouse using Hive
  • Worked extensively with Sqoop for importing metadata from Oracle
  • Extensively used Pig for data cleansing
  • Created partitioned tables in Hive
  • Worked with business teams and created Hive queries for ad hoc access.
  • Evaluated usage of Oozie for Workflow Orchestration
  • Mentored analyst and test team for writing Hive Queries
  • Gained very good business knowledge on health insurance, claim processing, fraud suspect identification, appeals process etc.

Environment: Hadoop, MapReduce, HDFS, Hive, Java (jdk1.6), Hadoop distribution of Hortonworks, Oozie, Oracle 11g/10g

Confidential

Java Developer

Responsibilities:

  • Involved in the process Design, Coding and Testing phases of the software development cycle.
  • Designed use-case, sequence and class diagram (UML).
  • Developed rich web user interfaces using JavaScript (pre-developed library).
  • Created modules in Java and C++, python.
  • Developed JSP pages with Struts framework, Custom tags and JSTL.
  • Developed Servlets, JSP pages, Beans, JavaScript and worked on integration.
  • Developed SOAP/WSDL interface to exchange usage and Image and terrain information from Geomaps.
  • Developed Unit test cases for the classes using JUnit.
  • Developed stored procedures to extract data from Oracle database.
  • Developed and maintained Ant Scripts for the build purposes on testing and production environments.
  • Designed and developed user interface components using AJAX, JQuery, JSON, JSP, JSTL & Custom Tag library.
  • Involved in building and parsing XML documents using SAX parser.
  • Application developed with strict adherence to J2EE best practices.

Environment: Java, C++, Python, Ajax, JavaScript, Struts, Spring, Hibernate, SQL/PLSQL, Web Services, WSDL, Linux, Unix

Confidential

Java Developer

Responsibilities:

  • Understanding andanalyzingthe requirements.
  • Implemented server-side programs by usingServletsand JSP.
  • Designed, developedand validatedUser Interface using HTML, Java Script, XMLandCSS.
  • Implemented MVC using Struts Framework.
  • Handled the database access by implementing Controller Servlet.
  • Implemented PL/SQL stored procedures and triggers.
  • Used JDBC prepared statements to call from Servlets for database access.
  • Designed and documented of the stored procedures
  • Widely used HTML for web based design.
  • Involved in Unit testing for various components.
  • Worked on database interaction layer for insertions, updating and retrieval operations of data from oracle database by writing stored procedures.
  • Used Spring Framework for Dependency Injection and integrated with Hibernate.
  • Involved in writing JUnit Test Cases.
  • Used Log4J for any errors in the application

Environment: Java, J2EE, JSP, Servlets, HTML, DHTML, XML, JavaScript, Struts, Eclipse, WebLogic, PL/SQL and Oracle.

We'd love your feedback!