Hadoop Administrator Resume
Hartford, CT
SUMMARY
- 18 Years of diverse experience in Software Engineering and Administration.
- 5 years of Experience in HADOOP Administration and AWS EMR.
- 3 Year of Automation Experience in Perl, Shell,Unix.
- 2 Years of Experience as Linux Support Engineer, QA Analyst.
- 8 Years of Software development experience in Java,C,C++ and Linux
- Experience in installation, configuration and management of Hadoop Clusters
- Experience with EMR Hortonworks HDP 2.3.4, HDP 2.6.1, HDP 2.6.3 and Cloudera CDH4, CDH5 distributions
- Experience in using Ambari, Cloudera Manager for tracking cluster utilization and Cloudera navigator for defining data lifecycle rules
- Extensive experience on configuration of cluster using Ambari Server for HDP
- Good Experiencing of deploying Hadoop2 cluster on EC2 cloud service by AWS.
- Worked on Migration of On - Prem to AWS EMR Cluster.
- Worked on various services like SPARK, HIVE, HBASE,HDFS, YARN,RANGER,AIRFLOW,KERBEROS
- EMR Automation using Ansible/Cloud Formation
- Experienced with DevOps tools like Chef, Puppet, Ansible, Jenkins, Jira and Docker.
- In depth knowledge on functionalities of every Hadoop daemon, interaction between them, resource utilizations and dynamic tuning to make cluster available and efficient
- Experience in providing security for Hadoop Cluster with Kerberos, Ranger
- Experience in creating job pools, assigning users to pools and restricting production job submissions based on pool
- Experience in setting up the monitoring tools such as Pepperdata Nagios and Ganglia to monitor and analyze the functioning of cluster.
- Good working knowledge of NoSQL databases such as Hbase, and knowledge on Cassandra and MongoDB
- Experience in analyzing data on HDFS through MapReduce, Hive
- Experience in setting up workflows and scheduling the workflows using Oozie
- Experience on UNIX commands and Shell Scripting and Python
- Excellent interpersonal, communication, documentation and presentation skills
- Strong experience in interacting with business analysts and developers to analyze the user requirements, functional specifications and system specifications.
- Working Knowledge on Configuration tools such as Ansible and Puppet.
- Good Understanding of Kafka Cluster.
TECHNICAL SKILLS
Hadoop/Big Data platform: HDFS, MapReduce, Hbase, Cassandra, Hive, Pig, Oozie, Zookeeper, Flume, Sqoop, Spark
Hadoop distribution: Horton Works,Cloudera, AWS EMR
Admin operations: Access control, Cluster maintenance, Performance tuning, Storage capacity management
Programming Languages: C,C++, Java, Python
Web Development Tools: VB Script
Operating Systems: Windows Series, HP Unix, Linux
Databases: MYSQL, Hbase, Cassandra
Scripting Languages: Perl, Shell, Python
PROFESSIONAL EXPERIENCE
Confidential, Hartford, CT
Hadoop Administrator
Responsibilities:
- Responsible for building scalable distributed data solutions using HDP Hortonworks
- Currently working on the migration of On-Prem cluster to AWS EMR cluster.
- Building Production and Non Production Clusters of AWS EMR .
- Working on Transient Clusters in AWS EMR.
- Migrating data from on-prem to S3 using wandisco tool
- Maintenance of the HDP Hortonworks Cluster on Teradata servers
- Configured the Clusters for various environments like PROD, DEV,TEST and Backup
- Creation of HDFS user directories, Hive databases and provide authorization using Ranger and IAM Policies for the new users.
- Deployment of code using Jenkins, SVN and git
- Interacting with Dev teams for day to day support.
- POC on Microsoft Azure for the migration of HDP cluster from on prem to HD Insights
- Monitoring the AWS EMR cluster logs using Sumologic
- Working with Unix/Linux Admin team in the administering of hardware and OS requirements for the Hortonworks Cluster.
- Working on HDP 2.6.3 HDP env in all the clusters.
- Enabled HA for Namenode, Resource Manager,Yarn Configuration, Hive Metastore and HBase
- Worked on Kafka for data streaming.
- Hands on Experience on Kerberos .Enabled Kerberos for authentication with Active Directory.
- Working on DevOps tools like Chef, Puppet, Ansible, Jenkins, Jira and Docker.
- Managing Amazon Web Services (AWS) infrastructure with automation and configuration management tools such as Chef, Ansible and Puppet.
- Experience on Ranger Configuration with AD for Authorization of cluster services
- Created Nifi setup in EMR and on-prem environments.
- Configuration of Flume and Sqoop in the environments
- Performed Ambari server upgrades
- Adding and Decommissioning of nodes as part of maintenance of the cluster
- Good Knowledge on HBase and Phoenix.
- Worked on taking snapshots of the Hbase tables and exporting to the backup cluster.
- Worked on Spark issues and Configuration and tuning of spark jobs.
- Monitored Hadoop cluster job performance and capacity planning.
- HDFS, Hive and HBase Performance Tuning
- Worked on implementation of SSL /TLS implementation.
- Taking blueprints and the snapshots of the clusters during any major changes to the cluster.
- Experience on Backup and restoration of Hive, Hue, Ranger, Ambari and Namenode Metadata backups .
- Fixing Issues related to the cluster configuration and NameNodes
- Monitoring and troubleshooting, and review Hadoop log files.
- Working using Curl Commands to fix any issues
- Configuration of SSL and trouble shooting in Hue.
- Optimized Map/Reduce Jobs to use HDFS efficiently by using various compression mechanisms
- Configured Journal nodes and Zookeeper Services for the cluster.
- Monitored Hadoop cluster job performance and capacity planning.
- Monitored and reviewed Hadoop log files.
Environment: AWS EMR, MapReduce, HDFS, Hive, SQL, Oozie, Sqoop, UNIX Shell Scripting, Yarn, Ranger
Confidential, Raleigh, NC
Hadoop Administrator
Responsibilities:
- Responsible for building scalable distributed data solutions using HDP Hortonworks
- Maintenance of the HDP Hortonworks Cluster of 280 Nodes.
- Configured the Clusters for various environments like PROD, SIT, CAT and DEV
- Interacting with Dev team for day to day support.
- Working with Unix/Linux Admin team in the administering of hardware and OS requirements for the Hortonworks Cluster.
- Performed HDP 2.6.1 upgrades in PROD, SIT, CAT and DEV environments.
- Enabled HA for Namenode, Resource Manager,Yarn Configuration, Hive Metastore and HBase
- Worked on Kafka for data streaming.
- Deployment of code using Jenkins, SVN and GIT.
- Hands on Experience on Kerberos .Enabled Kerberos for authentication with Active Directory.
- Experience on Ranger, Knox Configuration with AD for Authorization of cluster services
- Hands on Experience on Puppet in pushing/deploying the configurations to the cluster.
- Monitoring the Kafka data synchronization using Commands and Zookeeper administration
- Configuration of Flume and Sqoop in the environments
- Configuring and Moving of Journal Nodes during the expansion of the cluster.
- Performed Ambari server upgrades
- Adding and Decommissioning of nodes as part of maintenance of the cluster
- Good Knowledge on Cassandra
- Worked on Spark issues and Configuration
- Monitored Hadoop cluster job performance and capacity planning.
- HDFS, Hive and HBase Performance Tuning
- Performed Cluster to Cluster Copy using distcp
- Worked on implementation of SSL /TLS implementation.
- Taking blueprints and the snapshots of the clusters during any major changes to the cluster.
- Taking backup of Critical data, Hive data and creating snapshots.
- Fixing Issues related to the cluster configuration and NameNodes
- Monitoring and troubleshooting, and review Hadoop log files.
- Working using Curl Commands to fix any issues
- Automating the process for kafka and the other scripts for running commands using the shell scripting.
- Configuration of SSL and trouble shooting in Hue.
- Optimized Map/Reduce Jobs to use HDFS efficiently by using various compression mechanisms
- Enabled HA for Namenode, Resource Manager, Yarn Configuration and Hive Metastore.
- Configured Journal nodes and Zookeeper Services for the cluster using Cloudera.
- Monitored Hadoop cluster job performance and capacity planning.
- Monitored and reviewed Hadoop log files.
- Responsible for building scalable distributed data solutions using Hadoop.
- Responsible for cluster maintenance, adding and removing cluster nodes, cluster
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce and loaded data into HDFS. Extraction data using Flume. Import/Export to HDFS/RDMS using Sqoop
- Analyzed the data by performing Hive queries and running Pig scripts to know user behavior.
- Good Knowledge of NoSQL database like HBase
- Performance tuning of Impala jobs and resource management in cluster.
Environment: MapReduce, HDFS, Hive, SQL, Oozie, Sqoop, UNIX Shell Scripting, Yarn.
Confidential, Sugarland, TX
Linux Support /QA Analyst
Responsibilities:
- Understanding the existing system and working in Ajile environment
- Updating the Product Owner on daily activities in Standup meetings
- Created Test scenario and Test cases based on the User Stories.
- Involvement in Sprint Planning.
- Executed test cases using ALM and updates the Testcases.
- Validated Test results and raised defects using ALM.
- Writing Complex SQL queries for Validations.
- Analyzed Business and System Requirement documents and Created Test Plan Document.
- Created Test data for Positive and negative scenarios and worked with Development team to get the accurate data for each test cases.
- Review of Testcases with the development team.
- Participate and active involvement in the Sprint Demos.
- Discussions with Product owner regarding the Testing activities for the user stories that is being developed.
- Creating Defects and follow up with the developers to fix the issues and retesting.
Environment: MapReduce, HDFS, Hive, SQL, Oozie, Sqoop, UNIX Shell Scripting, Yarn.
Confidential, IL
Linux Support Engineer/Automation
Responsibilities:
- As a Linux Support Engineer, understand the various business processes implemented via TMS6 - loading processes, End of Day and End of Month activities, Testing and Deployment process.
- Involved in setting up the lab for the test environment.
- Setting up of Virtual Machines using VMWare
- Installation of Linux and configuring the system.
- Configured Field Devices (Acculoads, PLC) to TAS Servers
- Install and maintain all server hardware and software systems and administer all server performance and ensure availability for same.
- Monitor everyday systems and evaluate availability of all server resources and perform all activities for Linux servers.
- Maintain and monitor all system frameworks and provide after call support to all systems and maintain optimal Linux knowledge.
- Perform tests on all new software and maintain patches for management services and perform audit on all security processes.
- Gathered test data requirements for data conditioning from Business Units to test total application functionality.
- Developed Automation Scripts using shell scripts to check the log files size, and report the application.
- Responsible in writing the cron jobs to start the processes at regular intervals.
- Involved in Database Testing Using SQL to pull data from database and check whether it matches with GUI.
- Responsible in supporting the Deployment activities, tracking the data flow from TMS6 application to Tophat application and other external systems and Bubble support activities.
Environment: RedHat Linux 6.3, MySQL, VMware, Shell, Perl
Confidential
Systems Analyst
Responsibilities:
- Involved in developing detailed test strategy, test plan, test cases and test scripts for Automation Testing.
- Set up the test environment, defining detailed Test Requirements, converting them into Test Cases and collected Test Metrics for analyzing the Testing Effort.
- Develop, maintain and conduct smoke test cases for QA environments.
- Development of library functions for the Automation Test cases.
- Coordinate and work along with the development and business teams. Controlled testing projects at every step of the quality cycle from test planning through execution of defect management.
- Involved as part of automation team for the development of Perl and Shell Scripts to automate the DPG Product.
- DPG product involves snap mirroring among volumes and aggregates, taking backup from disk to tape and other operations on filers.
- Debugging the logs when the problem occurs during execution.
- Unit testing and System testing of the scripts.
Environment: QTP, Linux, Filers, Perl
Confidential, San Diego, CA
Automation Engineer
Responsibilities:
- Coordinate and work along with the onsite coordinator
- Automation of scripts in Perl.
- Debugging the logs when the problem occurs during execution.
- Unit testing and System testing and UAT of the scripts.
- Performed regression, functional, system, UAT testing on main application
- Developing and maintaining test scripts, analyzing bugs and interacting with development team members in fixing the defects.
- Responsible for writing simple to complex SQL queries to verify the data in database.
- Responsible for analysis, reports and defect tracking.
Environment: Linux, Perl, Java
Confidential
Sr Software Engineer/Module Lead
Responsibilities:
- Solving the bugs raised by customers as problems reports,
- Helpdesk requests and NOKs: ensuring precise and speedy resolution to global clients of Confidential .
- Software maintenance and Enhancements: Fixing bugs, releasing Change/technical notes.
- Responsible to lead the module related activities
- Developed Perl scripts to verify the output of the Network Elements.
- Query the data using SQL*Plus
- Responsible for extracting and loading data into database for report generation and other functionalities.
Environment: GUI based application, C++, HP UNIX, Perl, Shell, and Oracle 7.3
Confidential
Sr. Software Engineer
Responsibilities:
- Solving the bugs raised by customers as problems reports.
- Helpdesk requests and NOKs: ensuring precise and speedy resolution to global clients of Confidential .
- Software maintenance and Enhancements: Fixing bugs, releasing Change/technical notes.
- Responsible to lead the module related activities
- Developed Perl scripts to verify the output of the Network Elements.
- Query the data using SQL*Plus
- Responsible for extracting and loading data into database for report generation and other functionalities.
- Worked on fixing critical bugs found in customer labs. Analyzed bugs prepared Implementation analysis reports and then once approval was got from product management, implemented the fix.
Environment: GUI based application, C++, HP UNIX, Perl, Shell, and Oracle 7.3
Confidential
Software Engineer
Responsibilities:
- Handled Collection and distribution modules in C++ and Linux
- Requirements gathering for the project.
- Involved in development of Collection and distribution modules in C++ contains CMIP protocol.
- Processing of CDR (Call Detail Record) information in the Raw CDR files.
- Unit testing of the developed modules.
- Involved in the module of putting the processed information into the database.
- Responsible for unit Testing and UAT testing of the product.
- Collection of Data from the switches using RS232.
Environment: C++, Linux, MySQL
Confidential
Software Engineer
Responsibilities:
- Responsible for attending release progress meetings
- Normalization of the Raw CDR files
- Requirements Gathering for the project
- Processing of Raw CDRs files of Alcatel OCB 283 switch
- Processing of Raw CDR files of EWSD Confidential switch
- Responsible for the unit testing and UAT testing of the project.
- Collection of data from the switches using RS232.
Environment: C++, Linux, MySQL
