We provide IT Staff Augmentation Services!

Hadoop Administrator Resume

4.00/5 (Submit Your Rating)

Bellevue, SeattlE

SUMMARY

  • Seven years extensive IT experience that includes Data warehousing, Systems Administration, monitoring and troubleshooting experience on UNIX, Hadoop environments.
  • A qualified Technocrat and a seasoned professional in IT with 3 years of experience in Hadoop Administration, Big Data Ecosystem and 4.5 years of experience in Linux Administration.
  • Installation, configuration and administration experience inBig Dataplatforms Cloudera CDH, Hortonworks Ambari, Apache Hadoop on Redhat, and Centos asadatastorage, retrieval, and processing systems.
  • Hands on experience in installing, configuring, and using Hadoop ecosystem components like Map Reduce, HDFS, HBase, Oozie, Hive, Sqoop, Pig, impala and Zoo keeper.
  • As a Hadoop Administration responsibilities include software installation, configuration, software upgrades, backup and recovery, commissioning and decommissioning data nodes, cluster setup, cluster performance and monitoring on daily basis, maintaining cluster on healthy on different Hadoop distributions (Hortonworks & Cloudera).
  • Hadoop environment build and support including design, capacity planning, configuration management, monitoring, debugging, and performance tuning.
  • Hands on experience in all aspects of software development life cycle(SDLC) in Waterfall & Agile environments.
  • Experience in Implementing High Availability of Name Node and Hadoop Cluster capacity planning.
  • Experience in large - scale data processing, on an Amazon EC2 cluster.
  • Experience in analyzing data using HiveQL, Pig Latin, HBase and custom Map Reduce programs in Java.
  • Experience in support analysts by administering and configuring Hive.
  • Extending Hive and Pig core functionality by writing custom UDFs.
  • Extensive experience in developing PIG Latin Scripts and using Hive Query Language for data analytics.
  • Experience in troubleshooting errors in HBase Shell/API, Pig, Hive and MapReduce.
  • Experience in managing and reviewing Hadoop log files.
  • Experience in upgrading Hadoop cluster from current version to minor version upgrade as well as to major versions.
  • Experience in Unix Shell Scripting, SQL, Reporting and validating complex Stored Procedures, Triggers
  • Monitored the cluster resources & configured the alerts using Cloudera Manager for the Hadoop cluster.
  • Supported technical team members for automation, installation and configuration tasks.
  • Performed Importing and exporting data into HDFS and Hive using Sqoop.
  • Manage nodes on Hadoop cluster.
  • Expertise in monitoring tools like Oozie and Zookeeper.
  • Successfully uploaded files into Hive and HDFS from MongoDB, Hbase.
  • Good Knowledge on NoSQL databases MongoDB & Hbase.
  • Hadoop cluster connectivity check.
  • Implement new Hadoop hardware infrastructure.
  • HDFS support and maintenance.
  • Adding/Removing a Node, Data Rebalancing.
  • Tuning map reduce jobs and Maintaining backups for name node.
  • Determined, committed, hardworking with strong communication, interpersonal and organizational skills.
  • Experienced with handling different phases in big data environments.
  • Excellent communication, verbal skills to interact with client, onsite-offshore coordination.
  • Ability to adapt to evolving technology, strong sense of responsibility and accomplishment.

TECHNICAL SKILLS

Big Data: Cloudera Manager(CDH5), Hortonworks HDP 2.x

Big Data Ecosystem: HDFS, Map Reduce, Sqoop, Zoo keeper, Oozie, Hive, Pig

Operating Systems: Linux, Cent OS, Red Hat Linux, MS Windows family.

Scripting Languages: Java, C, C++, HTML/XHTML, SQL, PL/SQL, Python, Linux Shell s

Database: DB2, Oracle

ETL Tools: Informatica, IBM Infosphere and Qlikview.

Web Servers: Web Logic, Web Sphere, Apache Tomcat.

NoSQL Databases: HBase, Cassandra, MongoDB and CouchDB.

Methodologies: Design Patterns, Agile, V-model, Water Fall and UML

Cloud Services: Amazon Web Services, Microsoft Azure and Red-hat Openstack.

IDE’s: Eclipse, Dreamweaver, Net Beans

PROFESSIONAL EXPERIENCE

Confidential, Bellevue, Seattle

Hadoop Administrator

Responsibilities:

  • Experience in setup, configuration and management of security for Cloudera Hadoop clusters.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Responsible for day-to-day activities which includes HDFS support and maintenance, Cluster maintenance, creation/removal of nodes, Cluster Monitoring/ Troubleshooting,
  • Involved in manage and review the Hadoop log files, Backup and restoring, capacity planning
  • Worked with Hadoop developers and operating system admins in designing scalable supportable infrastructure for Hadoop
  • Worked in the administration activities in providing installation, upgrades, patching and configuration for all hardware and software Hadoop components.
  • Responsible for cluster availability and available 24x7 on call support.
  • Managing and reviewing data backups and Hadoop log files.
  • Responsible for deciding the hardware configurations for the cluster along with other teams.
  • Implemented the Cluster High Availability in case of crash or planned maintenance.
  • Responsible for scheduling jobs in Hadoop using Fair scheduler.
  • Involved in configuring Oozie workflow engine to run multiple Hive jobs.
  • Configured Sqoop and developed scripts to extract data from DB2 into HDFS.
  • Worked extensively with Sqoop for importing metadata from DB2.
  • Involved in administration, configuration management, monitoring, debugging and performance tuning of Hadoop environments.
  • Experience in configuring, installing, managing and administrating HBase clusters.
  • Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
  • Configured the Alerts using Cloudera Manager for the Hadoop cluster.
  • Installed and configured Flume, Hive, Sqoop, Zookeeper and Oozie on the Hadoop cluster.
  • Maintain extensive documentation on Hadoop cluster, policies and configurations.
  • Diligently teaming with the infrastructure, network, database, application and business intelligence teams to guarantee high data quality and availability.

Environment: CDH 5.4, Cloudera Manager, Hive, Sqoop, Zookeeper, Oozie, CENT OS, Unix scripts, YARN, Capacity Scheduler, Kerberos, Oracle, MySQL, Ganglia, DB2.

Confidential, San Francisco, CA

Hadoop Developer/Admin

Responsibilities:

  • Primary responsibilities include building scalable distributed data solutions using Hadoop ecosystem
  • Installed and configured Hive on the Hadoop cluster
  • Worked alongside with the accountants, financial analysts, data analysts, data scientists, statisticians, compliance, sales, marketing, pricing strategists, product development, and business analysts to create solutions for their issues.
  • Highly experienced in Developing complex MapReduce streaming jobs using Java language that are implemented Using Hive and Pig.
  • Optimized MapReduce Jobs to use HDFS efficiently by using various compression mechanisms.
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and Extracted the data from MySQL into HDFS using Sqoop
  • Analyzed the data by performing Hive queries (HiveQL) and running Pig scripts (Pig Latin) to study customer behavior.
  • Tested Apache(TM) Tez, an extensible framework for building high performance batch and interactive data processing applications, on Pig and Hive jobs.
  • Used Impala to query the Hadoop data stored in HDFS.
  • Working as aHadoopconsultant for converting the Oracle Stored Procedures based DataWarehouse Solution to Hadoop based Solution.
  • Filtered, transformed and combined data from multiple providers based on payer filter criteria using custom Pig UDFs.
  • Used the RegEx, JSON and Avro SerDe's for serialization and de-serialization packaged with Hive to parse the contents of streamed log data and implemented Hive custom UDF's.
  • Extensively used Informatica Power Center in end-to-end of Data warehousing ETL routines, which includes writing custom scripts, data mining and data quality process.
  • Continuous monitoring and managing the Hadoop cluster using Cloudera Manager
  • Experience in using Sqoop to migrate data to and fro from HDFS and My SQL or Oracle and deployed Hive and HBase integration to perform OLAP operations on HBase data.
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required
  • Exported the analyzed data to the relational databases using HIVE for visualization and to generate reports for the BI team
  • Perform data analysis on large datasets and present results to risk, finance, accounting, pricing, sales, marketing, and compliance teams.
  • Experienced on loading and transforming of large sets of structured, semi structured and unstructured data.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Written multiple MapReduce programs in Java for data extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV and other compressed file formats.

Environment: Hadoop 2.4.0 - PIG, Hive, Java, Cloudera manager, 30 Node cluster with Linux-Ubuntu, Linux, MapReduce, HDFS, Sqoop, Shell Scripting, Java (JDK1.6), Java 6, Eclipse, Oracle 10g, PL/SQL, SQL*PLUS, Toad 9.6, Linux, JIRA 5.1, CVS, JIRA 5.2.

Confidential, Dulles, VA

MongoDB Developer/Admin

Responsibilities:

  • Gathered the business requirements from the Business Partners and Subject Matter experts.
  • Involved in installing Hadoop ecosystem components.
  • Responsible to manage data coming from different sources.
  • Monitoring the jobs to analyze performance statistics.
  • Performing Unit Testing of completed jobs.
  • Installed and configured MongoDB for dev and test environments.
  • Involved in user location based advertise service using mongodB’s geospatial indexing.
  • Tuned mongodb queries and indexes so that response time was under five milliseconds and processing 12,000 transactions per second, or several billion each month.
  • Involved in setting up MongoDB connector for Hadoop as plugin in Hadoop so that user information is staged to HDFS for further AI algorithms.
  • Involved in design of efficient shard key and setting up of shard and replica set.
  • Managing Mongo databases using MMS monitoring tool.
  • Responsible to manage data coming from different sources.
  • Monitoring the jobs to analyze performance statistics.
  • Performing Unit Testing of completed jobs.
  • Applying optimization techniques at both Hadoop and Database level.
  • Involved in running Hadoop jobs for processing millions of records of text data.
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
  • Experienced in defining job flows.
  • Experienced in managing and reviewing Hadoop log files.
  • Extracted files from MongoDB through Sqoop and placed in HDFS and processed.
  • Experienced in running Hadoop streaming jobs to process terabytes of XML format data.
  • Load and transform large sets of structured, semi structured and unstructured data.
  • Plan, design, and implement processing massive amounts of marketing information, complete with information enrichment, text analytics, and natural language processing.
  • Prepare multi-cluster test harness to exercise the system for performance and failover.

Environment: Hadoop, HDFS, Map Reduce, Hive, Pig, Sqoop, Oozie, HBase, Linux, Java, Xml, MongoDB.

Confidential

Linux Administrator

Responsibilities:

  • Installation and configuration of Apache and supporting them on Linux production servers.
  • Administration of RHEL4.x, 5.x which includes installation, testing, tuning, upgrading and loading patches, troubleshooting both physical and virtual server issues.
  • Installing RedHat Linux using kick start and applying security polices for hardening the server based on company’s policies.
  • Installed and verified that all AIX/Linux patches are applied to the servers.
  • Maintenance and installation of RPM and YUM package installations and other server management.
  • Managing and scheduling cron jobs such as enabling system logging, network logging of servers for maintenance, performance tuning and testing.
  • Performed data-center operations including rack mounting and cabling.
  • Set up user and group login ID, printing parameters, network configuration, password, resolving permissions issues, user and group quota.
  • Worked on various applications and improving their performance by performance tuning and analysis.
  • Configuring multipath, adding SAN and creating physical volumes, volume groups and logical volumes.
  • Performed various configurations which include networking and IPTables, resolving hostnames, SSH key less login.
  • Troubleshooting Linux network and capturing packets using tools such as IPtables, firewall, TCP wrappers and NMAP.

Environment: LINUX, TCP/IP, TELENT, UBUNT, RHEL.

Confidential

Linux Administrator

Responsibilities:

  • Build Linux servers. Upgrade and patch existing servers. Compile, built and upgrade Linux kernel.
  • Setup Solaris Custom Jumpstart server and clients and implement Jumpstart installation.
  • Worked with Telnet, rlogin, used to inter-operate hosts.
  • Contact various systems administration works under CentOS, Red Hat Linux environments.
  • Performed regular day-to-day system administrative tasks including User Management, Backup, Network Management, and Software Management including Documentation etc.
  • Recommend system configurations for clients based on estimated requirements.
  • Performed reorganization of disk partitions, file systems, hard disk addition, and memory upgrade.
  • Monitored system activities, log maintenance, and disk space management.
  • Encapsulated root file systems, and mirrored the file systems were mirrored to ensure systems had redundant boot disks.
  • Administer Apache Servers. Published client’s web site in our Apache server.
  • Fix all the system problems, based on system email information and users’ complaints.
  • Upgrade software, add patches, and add new hardware in UNIX machines.

Environment: UNIX, FTP, TCP/IP, Red Hat Linux.

Confidential

Java/J2EE Developer

Responsibilities:

  • Involved in various phases of Software Development Life Cycle (SDLC) as design development and unit testing.
  • Developed and deployed UI layer logics of sites using JSP, XML, JavaScript, HTML/DHTML, and Ajax.
  • CSS and JavaScript were used to build rich internet pages.
  • Agile Scrum Methodology been followed for the development process.
  • Designed different design specifications for application development that includes front-end, back-end using design patterns.
  • Developed proto-type test screens in HTML and JavaScript.
  • Involved in developing JSP for client data presentation and, data validation on the client side with in the forms.
  • Developed the application by using the Spring MVC framework.
  • Collection framework used to transfer objects between the different layers of the application.
  • Developed data mapping to create a communication bridge between various application interfaces using XML, and XSL.
  • Spring IOC being used to inject the parameter values for the Dynamic parameters.
  • Developed JUnit testing framework for Unit level testing.
  • Actively involved in code review and bug fixing for improving the performance.
  • Documented application for its functionality and its enhanced features.
  • Created connection through JDBC and used JDBC statements to call stored procedures.

Environment: Spring MVC, Oracle 11g J2EE, Java, JDBC, Servlets, JSP, XML, Design Patterns, CSS, HTML, JavaScript 1.2, JUnit, Apache Tomcat, My SQL Server 2008.

We'd love your feedback!