Hadoop Developer Resume
Englewood, CO
SUMMARY:
- About 7+ years of experience wif emphasis on Big Data Technologies, Development and Design of Java based enterprise applications.
- Excellent understanding / knowledge of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
- Experience in installation, configuration, supporting and managing Hadoop Clusters using Apache, Cloudera (CDH3, CDH4) distributions and on Amazon web services (AWS).
- Hands - on experience on major components in Hadoop Ecosystem including Hive, HBase, HBase-Hive Integration, PIG, Sqoop, Flume& knowledge of Mapper/Reduce/HDFS Framework.
- Set up standards and processes for Hadoop based application design and implementation.
- Worked on NoSQL databases including HBase, Cassandra and Monod.
- Good experience in analysis using PIG and HIVE and understanding of SQOOP and Puppet.
- Expertise in database performance tuning data modeling.
- Developed automated scripts using Unix Shell for performing RUNSTATS, REORG, REBIND, COPY, LOAD, BACKUP, IMPORT, EXPORT and other related to database activities.
- Experienced in developing Map Reduce programs using Apache Hadoop for working wif Big Data.
- Good understanding of XML methodologies (XML, XSL, XSD) including Web Services and SOAP.
- Expertise in working wif different databases likes Oracle, MS-SQL Server, Postures, and MS Access 2000 along wif exposure to Hibernate for mapping an object-oriented domain model to a traditional relational database.
- Extensive experience in data analysis using tools like Sync sort and HZ along wif Shell Scripting and UNIX.
- Involved in log file management where teh logs greater than 7 days old were removed from log folder and loaded into HDFS and stored for 3 months.
- Expertise in development support activities including installation, configuration and successful deployment of changes across all environments.
- Familiarity and experience wif data warehousing and ETL tools.
- Good working Knowledge in OOA&OOD using UML and designing use cases.
- Good understanding of Scrum methodologies, Test Driven Development and continuous integration.
- Experience in production support and application support by fixing bugs.
- Used HP Quality Center for logging test cases and defects.
- Major strengths are familiarity wif multiple software systems, ability to learn quickly new technologies, adapt to new environments, self-motivated, team player, focused adaptive and quick learner wif excellent interpersonal, technical and communication skills.
TECHNICAL SKILLS:
Big Data: Hadoop, Hive, Sqoop, Pig, Puppet, Ambary, HBase, Monod, Cassandra, Power Pivot, Defamer, Pentaho, Flume
Operating Systems: Windows, Ubuntu, Red Hat Linux, Linux, UNIX
Project Management: Plan View, MS-Project
Programming or Scripting Languages: Java, SQL, Unix Shell Scripting, C.
Modeling Tools: UML, Rational Rose
IDE/GUI: Eclipse
Database: MS-SQL, Oracle, MS-Access
Middleware: Web Sphere, TIBCO
ETL: Informatica, Pentaho
Business Intelligence: OBIEE, Business Objects
Testing: Quality Center, Win Runner, Load Runner, QTP
PROFESSIONAL EXPERIENCE:
Confidential
Hadoop Developer
Responsibilities:
- Developed data pipeline using Flume, Sqoop, Pig and Java map reduce to ingest customer behavioral data and financial histories into HDFS for analysis.
- Worked on importing and exporting data from Oracle and DB2 into HDFS and HIVE using Sqoop for analysis, visualization and to generate reports.
- Developed multiple Map Reduce jobs in java for data cleaning.
- Developed Hive UDF to parse teh staged raw data to get teh Hit Times of teh claims from a specific branch for a particular insurance type code.
- Schedule these jobs wif workflow engine like Oozie. Actions can be performed both sequentially and parallel using Oozie.
- Built wrapper shell scripts to hold this Oozie workflow.
- Involved in collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
- Involved in creating Hadoop streaming jobs using Python.
- Provided ad-hoc queries and data metrics to teh Business Users using Hive, Pig.
- Developed PIG Latin scripts to extract teh data from teh web server output files to load into HDFS.
- Used Pig as ETL tool to do transformations, event joins and some pre-aggregations before storing teh data onto HDFS.
- Worked on Map reduce Joins in querying multiple semi-structured data as per analytic needs.
- Used Hive to analyze teh partitioned and bucketed data and compute various metrics for reporting.
- Created many Java UDF and UDAFs in hive for functions that were not preexisting in Hive like teh rank, Scum, etc.
- Used Hive and created Hive tables and involved in data loading and writing Hive UDFs.
- Developed POC for Apache Kafka.
- Do various performance optimizations like using distributed cache for small datasets, partition and bucketing in hive, doing map side joins etc..
- Storing and loading teh data from HDFS to Amazon S3 and backing up teh Namespace data into NFS Filers.
- Created concurrent access for hive tables wif shared and exclusive locking that can be enabled in hive wif teh help of Zookeeper implementation in teh cluster.
- Wrote teh shell scripts to monitor teh health check of Hadoop daemon services and respond accordingly to any warning or failure conditions.
- Familiarity wif NoSQL databases including HBase, Monod.
- Wrote shell scripts for rolling day-to-day processes and it is automated.
Environment: Hadoop, Map Reduce, Hive, HDFS, PIG, Sqoop, Oozie, Cloudera, Flume, HBase, Zookeeper, CDH3, Monod, Cassandra, Oracle, NoSQL and Unix/Linux, Kafka, Amazon web services.
Confidential, Englewood, Co
Hadoop Developer / Admin
Responsibilities:
- Installed/Configured/Maintained Apache Hadoop clusters for application development and Hadoop tools like Hive, Pig, HBase, Zookeeper and Sqoop.
- Involved in analyzing system failures, identifying root causes and recommended course of actions.
- Worked on Hive for exposing data for further analysis and for generating transforming files from different analytical formats to text files.
- Wrote teh shell scripts to monitor teh health check of Hadoop daemon services and respond accordingly to any warning or failure conditions.
- Managing and scheduling Jobs on a Hadoop cluster.
- Deployed Hadoop Cluster in teh following modes.
- Developed multiple Map Reduce jobs in java for data cleaning and accessing.
- Managed Hadoop clusters: monitor, maintain, setup.
- Importing and exporting data into HDFS and Hive using Sqoop
- Experienced in defining job flows.
- Implemented Name Node backup using NFS. This was done for High availability.
- Worked on importing and exporting data from Oracle and DB2 into HDFS using Sqoop.
- Developed PIG Latin scripts to extract teh data from teh web server output files to load into HDFS.
- Monitored workload, job performance and capacity planning using Cloudera Manager.
- Created Hive External tables and loaded teh data in to tables and query data using HQL.
- Wrote shell scripts for rolling day-to-day processes and it is automated.
- Used Flume to collect, aggregate, and store teh web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
- Analyzed teh web log data using teh Havel to extract number of unique visitors per day, page-views, visit duration, most purchased product on website.
- Used Ambary to manage, provision and monitor Hadoop cluster.
- Implemented Fair schedulers on teh Job tracker to share teh resources of teh Cluster for teh Map Reduce jobs given by teh users.
Environment: Hadoop 1x, Hive, Pig, HBASE, Sqoop, Flume, Zookeeper, Pig, HDFS, Ambary, Oracle, CDH3.
Confidential - Cambridge, MA
Java/Hadoop Developer
Responsibilities:
- Exported data from DB2 to HDFS using Sqoop.
- Developed Map Reduce jobs using Java API.
- Installed and configured Pig and also wrote Pig Latin scripts.
- Wrote Map Reduce jobs using Pig Latin.
- Developed workflow using Oozie for running Map Reduce jobs and Hive Queries.
- Worked on Cluster coordination services through Zookeeper.
- Worked on loading log data directly into HDFS using Flume.
- Involved in loading data from LINUX file system to HDFS.
- Responsible for managing data from multiple sources.
- Experienced in running Hadoop streaming jobs to process terabytes of xml format data.
- Responsible to manage data coming from different sources.
- Assisted in exporting analyzed data to relational databases using Sqoop.
- Implemented JMS for asynchronous auditing purposes.
- Created and maintained Technical documentation for launching Cloudera Hadoop Clusters and for executing Hive queries and Pig Scripts
- Experience in defining, designing and developing Java applications, specially using Hadoop Map/Reduce by leveraging frameworks such as Cascading and Hive.
- Experience in Develop monitoring and performance metrics for Hadoop clusters.
- Experience in Document designs and procedures for building and managing Hadoop clusters.
- Strong Experience in troubleshooting teh operating system, maintaining teh cluster issues and also java related bugs.
- Successfully loaded files to Hive and HDFS from Mongo DB Solar.
- Experience in Automate deployment, management and self-serve troubleshooting applications.
- Define and evolve existing architecture to scale wif growth data volume, users and usage.
- Design and develop JAVA API (Commerce API) which provides functionality to connect to teh Cassandra through Java services.
- Installed and configured Hive and also written Hive UDFs.
- Experience in managing teh CVS and migrating into Subversion.
- Experience in managing development time, bug tracking, project releases, development speed, release forecast, scheduling and many more.
Environment: Hadoop, HDFS, Hive, Flume, Sqoop, HBase, PIG, Eclipse, MySQL and Ubuntu, Zookeeper, Java (JDK 1.6)
Confidential, Chicago, IL
Java Developer
Responsibilities:
- Gathered user requirements followed by analysis and design. Evaluated various technologies for teh client.
- Developed HTML and JSP to present Client side GUI.
- Involved in development of JavaScript code for client side Validations.
- Designed teh HTML based web pages for displaying teh reports.
- Developed teh HTML based web pages for displaying teh reports.
- Developed java classes and JSP files.
- Extensively used JSF framework.
- Extensively used XML documents wif XSLT and CSS to translate teh content into HTML to present to GUI.
- Developed dynamic content of presentation layer using JSP.
- Develop user-defined tags using XML.
- Developed Java Mail for automatic emailing and JNDI to interact wif teh knowledge server.
- Used Struts Framework to implement J2EE design patterns (MVC).
- Developed, Tested and Debugged teh Java, JSP and EJB components using Eclipse.
- Developed Enterprise java Beans like Entity Beans, session Beans (both Stateless and State full Session beans) and Message Driven Beans.
Environment: Java, J2EE, EJB 2.1, JSP 2.0, Servlets 2.4, JNDI 1.2, Java Mail 1.2, JDBC 3.0, Struts, HTML, XML, CORBA, XSLT, Java Script, Eclipse3.2, Oracle10g, Weblogic8.1, Windows 2003.
Confidential, Winston Salem, NC
JAVA Developer
Responsibilities:
- Created teh Database, User, Environment, Activity, and Class diagram for teh project (UML).
- Implement teh Database using Oracle database engine.
- Designed and developed a fully functional generic n-tiered J2EE application platform teh environment was Oracle technology driven. Teh entire infrastructure application was developed using Oracle JDeveloper in conjunction wif Oracle ADF-BC and Oracle ADF- Rich Faces.
- Created an entity object (business rules and policy, validation logic, default value logic, security)
- Created View objects, View Links, Association Objects, Application modules wif data validation rules (Exposing Linked Views in an Application Module), LOV, dropdown, value defaulting, transaction management features.
- Web application development using J2EE: JSP, Servlets, JDBC, Java Beans, Struts, Ajax, JSF, JSTL, Custom Tags, EJB, JNDI, Hibernate, ANT, JUnit and Apache Log4J, Web Services, Message Queue (MQ).
- Designing GUI prototype using ADF 11G GUI component before finalizing it for development.
- Create Reusable Component (ADF Library and ADF Task Flow)
- Experience using Version controls such as CVS, PVCS, and Rational Clear Case.
- Creating Modules Using Task Flow wif Bounded and Unbounded
- Generating WSDL (Web Services) And Create Work Flow Using BPEL
- Handel teh AJAX functions (partial trigger, partial Submit, auto Submit)
- Created teh Skin for teh layout.
Environment: Java core, Servlet, JSF, ADF Rich client UI Framework ADF-BC (BC4J) 11g, web services Using Oracle SOA (Bell), Oracle WebLogic.
