We provide IT Staff Augmentation Services!

Hadoop Developer Resume

0/5 (Submit Your Rating)

Piscataway, NJ

SUMMARY:

  • About 8 years of experience wif emphasis on Big Data Technologies, Development and Design of Java based enterprise applications.
  • Excellent understanding / noledge of Hadoop architecture and various components such as HDFS, JobTracker, TaskTracker, NameNode, DataNode and MapReduce programming paradigm.
  • Experience in installation, configuration, supporting and managing Hadoop Clusters using Apache, Cloudera (CDH3, CDH4) distributions and on amazon web services (AWS).
  • Hands - on experience on major components in Hadoop Ecosystem including Hive, HBase, HBase-Hive Integration, PIG, Sqoop, Flume & noledge of Mapper/Reduce/HDFS Framework.
  • Set up standards and processes for Hadoop based application design and implementation.
  • Worked on NoSQL databases including Hbase, Cassandra and MongoDB.
  • Good experience in analysis using PIG and HIVE and understanding of SQOOP and Puppet.
  • Expertise in database performance tuning & data modeling.
  • Developed automated scripts using Unix Shell for performing RUNSTATS, REORG, REBIND, COPY, LOAD, BACKUP, IMPORT, EXPORT and other related to database activities.
  • Experienced in developing MapReduce programs using Apache Hadoop for working wif Big Data.
  • Good understanding of XML methodologies (XML, XSL, XSD) including Web Services and SOAP.
  • Expertise in working wif different databases likes Oracle, MS-SQL Server, Postgress, and MS Access 2000 along wif exposure to Hibernate for mapping an object-oriented domain model to a traditional relational database.
  • Extensive experience in data analysis using tools like Syncsort and HZ along wif Shell Scripting and UNIX.
  • Involved in log file management where teh logs greater than 7 days old were removed from log folder and loaded into HDFS and stored for 3 months.
  • Experienced in installing, configuring, and administrating Hadoop cluster of major Hadoop distributions.
  • Expertise in development support activities including installation, configuration and successful deployment of changes across all environments.
  • Experience in Creating a design and framework for teh generic Ab initio Code to handle teh multiple and ever expanding list of data files coming from source
  • Good Understanding in using teh Talend components such as tmap, tFileExist, tFileCompare, tELTAggregate, tOracleInput, tOracleOutput.
  • Familiarity and experience wif data warehousing and ETL tools.
  • Good working Knowledge in OOA & OOD using UML and designing use cases.
  • Good understanding of Scrum methodologies, Test Driven Development and continuous integration.
  • Experience in production support and application support by fixing bugs.
  • Used HP Quality Center for logging test cases and defects.
  • Major strengths are familiarity wif multiple software systems, ability to learn quickly new technologies, adapt to new environments, self-motivated, team player, focused adaptive and quick learner wif excellent interpersonal, technical and communication skills.

TECHNICAL SKILLS:

Big Data: Hadoop, Hive, Sqoop, Pig, Puppet, Ambari, HBase, MongoDB, Cassandra, PowerPivot, Datameer, Pentaho, spark, Flume, SolrCloud

Operating Systems: Windows, Ubuntu, Red Hat Linux, Linux, UNIX

Project Management: Plan View, MS-Project

Programming or Scripting Languages: Java, SQL, Unix Shell Scripting, C, Python

Modelling Tools: UML, Rational Rose

IDE/GUI: Eclipse

Framework: Struts, Hibernate

Database: MS-SQL, Oracle, MS-Access

Middleware: Web Sphere, TIBCO

ETL: Informatica, Pentaho, Netezza

Business Intelligence: OBIEE, Business Objects

Testing: Quality Center, Win Runner, Load Runner, QTP

PROFESSIONAL EXPERIENCE:

Confidential, Piscataway, NJ

Hadoop Developer

Responsibilities:

  • Developed data pipeline using Flume, Sqoop, Pig and Java map reduce to ingest customer behavioral data and financial histories into HDFS for analysis.
  • Worked on importing and exporting data from Oracle and DB2 into HDFS and HIVE using Sqoop for analysis, visualization and to generate reports.
  • Developed multiple MapReduce jobs in java for data cleaning.
  • Developed Hive UDF to parse teh staged raw data to get teh Hit Times of teh claims from a specific branch for a particular insurance type code.
  • Schedule these jobs wif workflow engine like Oozie. Actions can be performed both sequentially and parallely using Oozie.
  • Built wrapper shell scripts to hold dis Oozie workflow.
  • Involved in collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
  • Involved in creating Hadoop streaming jobs using Python.
  • Used Ganglia to Monitor and Nagios to send alerts about teh cluster around teh clock
  • Provided ad-hoc queries and data metrics to teh Business Users using Hive, Pig.
  • Developed PIG Latin scripts to extract teh data from teh web server output files to load into HDFS.
  • Used Pig as ETL tool to do transformations, event joins and some pre-aggregations before storing teh data onto HDFS.
  • Worked on MapReduce Joins in querying multiple semi-structured data as per analytic needs.
  • Used Hive to analyze teh partitioned and bucketed data and compute various metrics for reporting.
  • Created many Java UDF and UDAFs in hive for functions dat were not preexisting in Hive like teh rank, Csum, etc.
  • Used Hive and created Hive tables and involved in data loading and writing Hive UDFs.
  • Developed POC for Apache Kafka.
  • Do various performance optimizations like using distributed cache for small datasets, partition and bucketing in hive, doing mapside joins etc..
  • Storing and loading teh data from HDFS to Amazon S3 and backing up teh Namespace data into NFS Filers.
  • Created concurrent access for hive tables wif shared and exclusive locking dat can be enabled in hive wif teh halp of Zookeeper implementation in teh cluster.
  • Wrote teh shell scripts to monitor teh health check of Hadoop daemon services and respond accordingly to any warning or failure conditions.
  • Familiarity wif NoSQL databases including HBase, MongoDB.
  • Wrote shell scripts for rolling day-to-day processes and it is automated.
  • Involved in story-driven agile development methodology and actively participated in daily scrum meetings.

Environment: Hadoop, MapReduce, Hive, HDFS, PIG, Sqoop, Oozie, Cloudera, Flume, HBase, ZooKeeper, CDH3, MongoDB, Cassandra, Oracle, NoSQL and Unix/Linux, Kafka, Amazon web services.

Confidential, Englewood, CO

Hadoop Developer / Admin

Responsibilities:

  • Experience wif Cloudera and Horton works distribution of Hadoop
  • Installed/Configured/Maintained Apache Hadoop clusters for application development and Hadoop tools like Hive, Pig, HBase, Zookeeper and Sqoop.
  • Involved in analyzing system failures, identifying root causes and recommended course of actions.
  • Worked on Hive for exposing data for further analysis and for generating transforming files from different analytical formats to text files.
  • Wrote teh shell scripts to monitor teh health check of Hadoop daemon services and respond accordingly to any warning or failure conditions.
  • Managing and scheduling Jobs on a Hadoop cluster.
  • Deployed Hadoop Cluster in teh following modes.
  • Developed multiple MapReduce jobs in java for data cleaning and accessing.
  • Managed Hadoop clusters: monitor, maintain, setup.
  • Importing and exporting data into HDFS and Hive using Sqoop
  • Experienced in defining job flows.
  • Implemented Name Node backup using NFS. dis was done for High availability.
  • Worked on importing and exporting data from Oracle and DB2 into HDFS using Sqoop.
  • Developed PIG Latin scripts to extract teh data from teh web server output files to load into HDFS.
  • Worked on custom Pig Loaders and Storage classes to work wif a variety of data formats such as JSON, Compressed CSV, etc.
  • Monitored workload, job performance and capacity planning using Cloudera Manager.
  • Created Hive External tables and loaded teh data in to tables and query data using HQL.
  • Wrote shell scripts to automate document indexing to SolrCloud in production.
  • Used Flume to collect, aggregate, and store teh web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
  • Used CGI scripts to access DocDB database.
  • Analyzed teh web log data using teh HiveQL to extract number of unique visitors per day, page-views, visit duration, most purchased product on website.
  • Converting teh Oracle table components to Teradata Table Components in Abinitio Graphs
  • Used Ambari to manage, provision and monitor Hadoop cluster.
  • Implemented Fair schedulers on teh Job tracker to share teh resources of teh Cluster for teh Map Reduce jobs given by teh users.

Environment: Hadoop 1x, Hive, Pig, HBASE, Sqoop, Flume, Zookeeper, Pig, HDFS, Ambari, Oracle, CDH3.

Confidential - Cambridge, MA

Java/Hadoop Developer

Responsibilities:

  • Exported data from DB2 to HDFS using Sqoop.
  • Developed MapReduce jobs using Java API.
  • Installed and configured Pig and also wrote Pig Latin scripts.
  • Wrote MapReduce jobs using Pig Latin.
  • Developed workflow using Oozie for running MapReduce jobs and Hive Queries.
  • Worked on Cluster coordination services through Zookeeper.
  • Worked on loading log data directly into HDFS using Flume.
  • Involved in loading data from LINUX file system to HDFS.
  • Responsible for managing data from multiple sources.
  • Experienced in running Hadoop streaming jobs to process terabytes of xml format data.
  • Responsible to manage data coming from different sources.
  • Assisted in exporting analyzed data to relational databases using Sqoop.
  • Implemented JMS for asynchronous auditing purposes.
  • Created and maintained Technical documentation for launching Cloudera Hadoop Clusters and for executing Hive queries and Pig Scripts
  • Experience wif CDH distribution and Cloudera Manager to manage and monitor Hadoop clusters
  • Experience in defining, designing and developing Java applications, specially using Hadoop Map/Reduce by leveraging frameworks such as Cascading and Hive.
  • Experience in Develop monitoring and performance metrics for Hadoop clusters.
  • Experience in Document designs and procedures for building and managing Hadoop clusters.
  • Strong Experience in troubleshooting teh operating system, maintaining teh cluster issues and also java related bugs.
  • Experienced import/export data into HDFS/Hive from relational database and Teradata using Sqoop.
  • Successfully loaded files to Hive and HDFS from Mongo DB Solar.
  • Experience in Automate deployment, management and self-serve troubleshooting applications.
  • Define and evolve existing architecture to scale wif growth data volume, users and usage.
  • Design and develop JAVA API (Commerce API) which provides functionality to connect to teh Cassandra through Java services.
  • Installed and configured Hive and also written Hive UDFs.
  • Experience in managing teh CVS and migrating into Subversion.
  • Experience in managing development time, bug tracking, project releases, development speed, release forecast, scheduling and many more.

Environment: Hadoop, HDFS, Hive, Flume, Sqoop, HBase, PIG, Eclipse, MySQL and Ubuntu, Zookeeper, Java (JDK 1.6)

Confidential, Chicago, IL

Java Developer

Responsibilities:

  • Gatheird user requirements followed by analysis and design. Evaluated various technologies for teh client.
  • Developed HTML and JSP to present Client side GUI.
  • Involved in development of JavaScript code for client side Validations.
  • Designed teh HTML based web pages for displaying teh reports.
  • Developed teh HTML based web pages for displaying teh reports.
  • Developed java classes and JSP files.
  • Extensively used JSF framework.
  • Extensively used XML documents wif XSLT and CSS to translate teh content into HTML to present to GUI.
  • Developed dynamic content of presentation layer using JSP.
  • Develop user-defined tags using XML.
  • Developed Java Mail for automatic emailing and JNDI to interact wif teh noledge server.
  • Used Struts Framework to implement J2EE design patterns (MVC).
  • Developed, Tested and Debugged teh Java, JSP and EJB components using Eclipse.
  • Developed Enterprise java Beans like Entity Beans, session Beans (both Stateless and State full Session beans) and Message Driven Beans.

Environment: Java, J2EE, EJB 2.1, JSP 2.0, Servlets 2.4, JNDI 1.2, Java Mail 1.2, JDBC 3.0, Struts, HTML, XML, CORBA, XSLT, Java Script, Eclipse3.2, Oracle10g, Weblogic8.1, Windows 2003.

Confidential

Java Developer

Responsibilities:

  • Created teh Database, User, Environment, Activity, and Class diagram for teh project (UML).
  • Implement teh Database using Oracle database engine.
  • Designed and developed a fully functional generic n-tiered J2EE application platform teh environment was Oracle technology driven. Teh entire infrastructure application was developed using Oracle JDeveloper in conjunction wif Oracle ADF-BC and Oracle ADF- RichFaces.
  • Created an entity object (business rules and policy, validation logic, default value logic, security)
  • Created View objects, View Links, Association Objects, Application modules wif data validation rules (Exposing Linked Views in an Application Module), LOV, dropdown, value defaulting, transaction management features.
  • Web application development using J2EE: JSP, Servlets, JDBC, Java Beans, Struts, Ajax, JSF, JSTL, Custom Tags, EJB, JNDI, Hibernate, ANT, JUnit and Apache Log4J, Web Services, Message Queue (MQ).
  • Designing GUI prototype using ADF 11G GUI component before finalizing it for development.
  • Create Reusable Component (ADF Library and ADF Task Flow)
  • Experience using Version controls such as CVS, PVCS, and Rational Clear Case.
  • Creating Modules Using Task Flow wif Bounded and Unbounded
  • Generating WSDL (Web Services) And Create Work Flow Using BPEL
  • Handel teh AJAX functions (partial trigger, partial Submit, auto Submit)
  • Created teh Skin for teh layout.

Environment: Java core, Servlet, JSF, ADF Rich client UI Framework ADF-BC (BC4J) 11g, web services Using Oracle SOA (BPEl), Oracle WebLogic.

We'd love your feedback!