Hadoop Admin/ Developer Resume
New York, CitY
SUMMARY
- Senior Software Developer, having 8 years of extensive experience in delivering challenging technology solutions, working wif geographically distributed teams.
- 3 years of experience in dealing wif Apache Hadoop components like BIG Data, HDFS,
- MapReduce, Hive, Pig, Sqoop, Oozie, and Big Data Analytics.
- Solid experience in designing, implementing, and improving analytic solutions for Big Data on Apache Hadoop. Experience in Map Reduce programming model and Hadoop Distributed File Systems.
- Deep understanding of data import & export from relational database into Hadoop cluster.
- Good Knowledge in NoSQL databases like MongoDB, Cassandra and HBASE.
- Experience in writing UDFS in java for hive and pig. Experience in using HCatalog for Hive, Pig, and HBase.
- Extensive noledge on Flume and HBase technologies. Experience in working wif flume to load teh log data from multiple sources directly into HDFS.
- Hands on experience in Import/Export of data using Hadoop Data Management tool SQOOP.
- Provided technical assistance for configuration, administration and monitoring of Hadoop clusters.
- Experience in test environment setup and test infrastructure development in both manual and automation.
- Experience wif database testing using various complex set of query and UDF’s.
- Experience in all phases of Software Development Life Cycle for maintaining and supporting teh Java, and J2EE applications.
- Quick learner and self - starter wif TEMPeffective communication, motivation and organizational skills combined wif attention to details and business process improvements.
- Excellent team player as well as an individual wif strong analytical, communication and interpersonal skills.
PROFESSIONAL EXPERIENCE
Confidential, New York City
Hadoop Admin/ Developer
Responsibilities:
- Responsible for architecting Hadoop clusters Translation of functional and technical requirements into detailed architecture and design.
- Installed and configured multi-nodes fully distributed Hadoop cluster of large number of nodes.
- Provided Hadoop, OS, Hardware optimizations.
- Setting up teh machines wif Network Control, Static IP, Disabled Firewalls, Swap memory.
- Installed and configured Cloudera Manager for easy management of existing Hadoop cluster
- Administered and supported distribution of Hortonworks.
- Worked on setting up high availability for major production cluster and designed automatic failover control using zookeeper and quorum journal nodes
- Implemented Fair scheduler on teh job tracker to allocate fair amount of resources to small jobs.
- Performed operating system installation, Hadoop version updates using automation tools.
- Configured Oozie for workflow automation and coordination.
- Implemented rack aware topology on teh Hadoop cluster.
- Importing and exporting structured data from different relational databases into HDFS and Hive using Sqoop
- Configured ZooKeeper to implement node coordination, in clustering support.
- Configured Flume for efficiently collecting, aggregating and moving large amounts of log data from many different sources to HDFS.
- Involved in collecting and aggregating large amounts of streaming data into HDFS using Flume and defined channel selectors to multiplex data into different sinks.
- Worked on developing scripts for performing benchmarking wif Terasort/Teragen.
- Implemented Kerberos Security Autantication protocol for existing cluster.
- Good experience in troubleshoot production level issues in teh cluster and its functionality.
- Backed up data on regular basis to a remote cluster using distcp.
- Regular Commissioning and Decommissioning of nodes depending upon teh amount of data.
- Monitored and configured a test cluster on amazon web services for further testing process and gradual migration
- Installed and maintain puppet-based configuration management system
- Deployed Puppet, Puppet Dashboard, and PuppetDB for configuration management to existing infrastructure.
- Using Puppet configuration management to manage cluster.
- Experience working on API
- Generated reports using teh Tableau report designer
Environment: BigData, Hadoop, MapReduce, Pig, Hive, Sqoop, Oozie, Crunch, Scala, Spark, Strom, kafka, Tableau, Cassandra, Linux, Python, R, RHadoop,Oracle10g,Cloudera manager, Maven, MRUnit, Junit
Confidential, Portland, OR
Hadoop Developer
Responsibilities:
- Participated in Gathering requirements, analyze requirements and design technical documents for business requirements.
- Involved different phases in big data projects like data acquiring, data processing and data serving using dash boards.
- Implemented different aggregations and filtering operations using JavaAPI in Cassandra
- Responsible for Data modeling in Cassandra and deciding teh row key and different column families in Cassandra.
- Import/export data from Oracle data base to/from HDFS using Sqoop, Hue and JDBC.
- Gathered data from different sources like Internet, sensors, user behavior using Flume and Kafka and moved to HDFS and implemented Optimized join base using MapReduce programs.
- Implemented Custom Input formats dat handles wide range of input files received from java applications to process in MapReduce.
- Implemented joins and data aggregation using Apache Crunch.
- Writing MapReduce pipeline programs for testing using Apache Crunch.
- Experienced in unit testing MapReduce programs using MRUnit.
- Divided each data set in to corresponding categories by following MapReduce Binning design pattern.
- Implemented FilterMappers to eliminate un-necessary records and perform data and schema validation.
- Experience in using Pig as an ETL tool for event joins, filters, transformations and pre- aggregations.
- Created partitions, bucketing across state in Hive to handle structured data.
- Implemented Dash boards dat handle HiveQL queries internally like Aggregation functions, basic hive operations, and different kind of join operations.
- Implemented business logic based on state in Hive using Generic UDF's.
- Involved in creating data-models for customer data using Cassandra Query Language.
- Managing and scheduling batch Jobs on a Hadoop Cluster using Oozie.
- Created production jobs using Oozie work flows dat integrated different actions like MapReduce, Sqoop, and Hive by utilizing fork and join operations in Oozie.
- Experience in managing and reviewing Hadoop Log files.
- Experienced wif monitoring Cluster using Cloudera manager and Ambari.
- Involved in managing Hadoop distributions using Hortonworks.
- Worked on building BI reports in Tableau wif Spark using SparkSQL.
- Experience in deploying data from various sources into HDFS and building reports using Tableau.
- Implemented Spark using Scala and SparkSQL for faster testing and processing of data.
- Implemented data ingestion and handling clusters in real time processing using kafka.
- Experience wif Core Distributed computing and Data Mining Library using ApacheSpark.
- Experienced in data modelling in hive implementing hive-indexing.
- Experience in utilizing spark machine learning techniques implemented in scala.
- Experienced in configuring maven builds dat integrated dependencies check styles, test coverage's.
- Responsible for analyzing multi-platform applications using python.
- Developed MapReduce jobs in Python for data cleaning and data processing.
- Designing Test Plans, Test Cases and performed System Testing.
- Involved in daily SCRUM meetings to discuss teh development/progress ofSprints and was active in making scrum meetings more productive.
- Experience in integrating RHadoop for categorization and statistical analysis to generate reports. Environment: BigData, Hadoop, MapReduce, Pig, Hive, Sqoop, Oozie, Crunch, Scala, Spark, Strom, kafka, Tableau, Cassandra, Linux, Python, R, RHadoop,Oracle10g,Cloudera manager, Maven, MRUnit, Junit
Environment: Hadoop, HDFS, Hive, HBase, Map Reduce, Pig, Cassandra Hive, Sqoop, Oozie, HL7, MRUnit, Junit, UNIX, Shell Scripting, MS Visio
Confidential, MO
Senior Hadoop Developer
Responsibilities:
- Evaluated business requirements and prepared detailed specifications dat follow
- Responsible for building scalable distributed data solutions using Hadoop.
- Analyzed large amounts of data sets to determine optimal way to aggregate and
- Developed Simple to complex Map reduce Jobs using Hive and Pig Optimized Map Reduce Jobs to use HDFS efficiently by using various compression mechanisms
- Handled importing of data from various data sources, performed transformations using Hive.
- MapReduce, loaded data into HDFS and Extracted teh data from MySQL into HDFS using Sqoop
- Exported teh analyzed data to teh relational databases using Sqoop for visualization and to generate reports for teh BI team
- Extensively used Pig for data cleansing.
- Gained good experience wif NOSQL database Cassandra.
- Performed importing data from various sources to teh Cassandra cluster using Java APIs or Sqoop.
- Worked wif NoSQL database Hbase to create tables and store data.
- Extract data from teh server and put it into HDFS and Bulk Loaded teh cleaned data into HBase
- Created tables, inserted data and executed variousCassandraQuery Language
- Created partitioned tables in Hive.
- Managed and reviewed Hadoop log files.
- Communicate wif teh clients on modules, requirements and change requests for any project guidelines required to develop written programs.
Environment: Hadoop, HDFS, Hive, Sqoop, Flume, HBase, Linux, JDBC, SQL, PL/SQL, Servlets, JSP
Confidential, Austin, TX
J2EE Developer
Responsibilities:
- Developed teh application using Struts Frameworkdat leverages classical Model View Layer (MVC) architecture UML diagrams like use cases, class diagrams, interaction diagrams, and activity diagrams were used
- Participated in requirement gathering and converting teh requirements into technical specifications
- Created Business Logic using Servlets, Sessionbeans and deployed them on Web logic server
- Wrote complex SQL queries and stored procedures
- Developed teh XMLSchema and Web services for teh data maintenance and structures
- Implemented teh Web Service client for teh login autantication, credit reports and applicant information using Apache Axis 2 Web Service
- Responsible to manage data coming from different sources.
- Designed teh logical and physical data model, generated DDL scripts, and wrote DML scripts for Oracle 9i database
- Used Hibernate ORM framework wif Spring framework for data persistence and transaction management
- Used struts validation framework for form level validation
- Wrote test cases in JUnitfor unit testing of classes.
- Provided Technical support for production environments resolving teh issues, analysing teh defects, providing and implementing teh solution defects
- Built and deployed Java applications into multiple Unix based environments and produced both unit and functional test results along wif release notes.
Environment: J2EE, Struts, JSP, Servlets, Web Sphere, HTML, XML, ANT, Oracle, JavaScript.
Confidential
Java Developer
Responsibilities:
- Involved in teh complete SDLC software development life cycle of teh application from requirement analysis to testing.
- Developed teh modules based on struts MVC Architecture.
- Developed Teh UI using JavaScript, JSP, HTML, and CSS for interactive cross browser functionality and complex user interface.
- Created Business Logic using Servlets, Session beans and deployed them on Weblogic server.
- Used MVC struts framework for application design.
- Created complex SQL Queries, PL/SQL Stored procedures, Functions for back end.
- Prepared teh Functional, Design and Test case specifications.
- Involved in writing Stored Procedures in Oracle to do some database side validations.
- Performed unit testing, system testing and integration testing
- Developed Unit Test Cases. Used JUNIT for unit testing of teh application.
- Provided Technical support for production environments resolving teh issues, analyzing teh defects, providing and implementing teh solution defects. Resolved more priority defects as per teh schedule.
Environment: JavaScript, JSP, Junit, HTML, CSS, PL/SQL, Oracle
