We provide IT Staff Augmentation Services!

Senior Hadoop Admin Resume

5.00/5 (Submit Your Rating)

Irving, TX

SUMMARY

  • Overall 8 years of experience in IT industry including 3+ years of experience in BIG DATA Hadoop Administration and 4 years of experience in Linux ecosystems.
  • Responsible for Cluster maintenance, Adding and removing cluster nodes, Cluster Monitoring and Troubleshooting, Manage and review data backups, Manage and review Hadoop log files on Horton works, MapR and Cloudera clusters.
  • Experience in installation, configuration, supporting and managing using Apache, Cloudera (CDH3, CDH4) distributions and on amazon web services (AWS).
  • Experience with complete Software Design Lifecycle including design, development, testing and implementation of moderate to advanced complex systems.
  • Worked on NoSQL databases including Hbase, Cassandra and MongoDB.
  • Good experience in analysis using PIG and HIVE and understanding of SQOOP and Puppet.
  • Excellent understanding / knowledge of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
  • Developed automated scripts using Unix Shell for performing RUNSTATS, REORG, REBIND, COPY, LOAD, BACKUP, IMPORT, EXPORT and other related to database activities.
  • Experienced in developing Map Reduce programs using Apache Hadoop for working with Big Data.
  • Good understanding of XML methodologies (XML, XSL, XSD) including Web Services and SOAP.
  • Experience in deploying Hadoop 2.0(YARN).
  • Good Experience in Planning, Installing and Configuring Hadoop Clusters in Cloudera and Hortonworks Distributions.
  • Expertise in implementing enterprise level security using AD/LDAP, Kerberos, Knox, Sentry and Ranger.
  • Extensive experience in developing the SOA middleware based out of Fuse ESB and Mule ESB.And Configured, Elastic Search Log Stash, Kibana to monitor spring batch jobs.
  • Extensive experience in data analysis using tools like Sync sort and HZ along with Shell Scripting and UNIX.
  • Excellent command in creating Backups & Recovery and Disaster recovery procedures and Implementing BACKUP and RECOVERY strategies for off - line and on-line Backups.
  • Good understanding of Scrum methodologies, Test Driven Development and continuous integration.
  • Expertise in working with different databases likes Oracle, MS-SQL Server, Postgress, and MS Access 2000 along with exposure to Hibernate for mapping an object-oriented domain model to a traditional relational database.
  • Experienced with Devops tools like Chef, Puppet, Ansible, Jenkins, Jira, Docker and Splunk.
  • Prepared, arranged and tested Splunk search strings and operational strings.
  • Experience in importing and exporting the data using Sqoop from HDFS to Relational Database systems/mainframe and vice-versa.
  • Familiar with writing Oozie workflows and Job Controllers for job automation.

TECHNICAL SKILLS

Big Data Technologies: Hadoop, HDFS, Hive, Map Reduce, Cassandra, Pig, Scoop, Falcon, Flume, Zookeeper, YarnMahout, Oozie, Avro, HBase, MapReduce, HDFS, Storm, CDH 5.3, CDH 5.4.

Operating Systems: Linux RHEL/Ubuntu/CentOS, Windows (XP/7/8/10)

Database& NoSql: Database Systems Oracle 11g/10g, DB2, SQL, My SQL, HBASE, Mongo DB, Cassandra

Scripting & security: Shell Scripting, HTML Scripting, Python, Kerberos, Dockors

Monitoring Tools: Cloudera Manager, Ambari, Nagios, Ganglia

Java Technologies: Java, J2EE, JSP, Servlets, Struts, Hibernate, Spring

Testing: Capybara, Web Driver Testing Frameworks RSpec, Cucumber, Junit, SVN

Server: WEBrick, Thin, Unicorn, Apache, AWS

Security: Kerberos

Other tools: Angular.js, knockout.js, backbone.js, ember.js, react.js, node.js, bootstrap, Redmine, Bugzilla, JIRAAgile SCRUM,SDLC Waterfall.

PROFESSIONAL EXPERIENCE

Confidential, Irving, TX

Senior Hadoop Admin

RESPONSIBILITIES:

  • Worked on Distributed/Cloud Computing (Map Reduce/ Hadoop, Hive, Pig, Hbase, Sqoop, Flume, Spark AVRO, Zookeeper, Tableau, etc.), Horton works (HDP 2.2.4.2), for 4 clusters ranges from POC to PROD contains nearly 100 nodes.
  • Performed a Major upgrade in production environment from HDP 1.3 to HDP 2.2. As an admin followed standard Back up policies to make sure the high availability of cluster.
  • Implement Flume, Spark, Spark Stream framework for real time data processing. Developed analytical components using Scala, Spark and Spark Stream. Implemented Proofs of Concept on Hadoop and Spark stack and different big data analytic tools, using Spark SQL as an alternative to Impala.
  • Involved in Analyzing system failures, identifying root causes, and recommended course of actions.
  • Imported logs from web servers with Flume to ingest the data into HDFS.
  • Involved in implementing security on Horton works Hadoop Clusters using with Kerberos by working along with operations team to move non secured cluster to secured cluster.
  • Responsible for upgrading Horton works Hadoop HDP2.2.0 and Map reduce 2.0 with YARN in Multi Clustered Node environment.Handled importing of data from various data sources, performed transformations using Hive, Map Reduce, Spark and loaded data into HDFS.
  • Migrated services from a managed hosting environment to AWS including: service design, network layout, data migration, automation, monitoring, deployments and cutover, documentation, overall plan, cost analysis, and timeline.
  • Performance tuning of Hadoop clusters and Hadoop MapReduce routines.
  • Working on this project spanning from Architecting, Installation, Configuration and Management of Hadoop Clusters.
  • Managing Amazon Web Services (AWS) infrastructure with automation and configuration management tools such as Chef, Ansible, Puppet, or custom-built designing cloud-hosted solutions, specific AWS product suite experience.
  • Development projects Extensively on Hive, Spark, Pig, Sqoop and Gem fire XD throughout the development Lifecycle until the projects went into Production. Created reporting views in Impala using Sentry Policy files.
  • Responsible for Handler configuration and handler ESB mappings. Also Involved in Integrating Hive with Mulesoft ESB to land data into applications running on Sales force and vice versa.
  • Implemented dual data center set up for all Cassandra cluster.Performed many complex system analysis in order to improve ETL performance, identified high critical batch jobs to prioritize
  • Implemented Spark solution to enable real time reports from Cassandra data.Was also actively involved in designing column families for various Cassandra Clusters.

ENVIRONMENT: RHEL, Ubuntu, Cloudera Manager, Cloudera Search, CDH4, HDFS, Hbase, Hive, Pig, ZooKeeper, Monitoring Cluster with automated scripts, Map Reduce2 (YARN), PostgreSQL, MySQL, QAS and Ganglia.

Confidential, Torrance, CA

Hadoop Admin

RESPONSIBILITIES:

  • Responsible for implementation and ongoing administration of Hadoop infrastructure.
  • Utilize big-data technologies such as Elastic Search, Riak, RabbitMQ, Couchbase, Redis, Docker, Mesos/Marathon, Jenkins, Puppet/Chef, Github, and much more.
  • Implemented a distributed messaging queue to integrate with Cassandra using Apache Kafka and ZooKeeper. Involved in a POC to implement a failsafe distributed data storage and computation system using Apache YARN.
  • Created Talend mappings for initial load and daily updates and also involved in the ETL migration jobs from Infromatica to Talend.
  • Adding/installation of new components and removal of them through Clouderav5 Manager.
  • Monitored workload, job performance and capacity planning using Cloudera Manager.
  • Integrated Impala to use the same file and data formats, metadata, security and resource management frameworks.
  • Written scripts to automate application deployments and configurations. Hadoop cluster performance tuning and monitoring. Troubleshoot and resolve Hadoop clusters related system problems.
  • Implemented test scripts to support test driven development and continuous integration.
  • Extensively worked in Hadoop, spark cluster and streams processing using Spark Streaming. Experience in Spark, python interfaces to Spark. Write Scoop, Spark and Map Reduce scripts and workflows.
  • Worked on creating reports, dashboards and alerts in Splunk for proactive monitoring
  • Working with data delivery teams to setup new Hadoop users. This job includes setting up Linux users, setting up Kerberos principals and testing HDFS, Hive.
  • Involved in generating and applying rules to profile data for flat files and relational data by creating rules to case cleanse, parse, standardize data through mappings in IDQ and generated as Mapplets in PC. Working knowledge with Talend ETL tool to filter data based on end requirements.

ENVIRONMENT: Hadoop, HDFS, Map Reduce, Shell Scripting,spark, Splunk, solr, Pig, Hive, HBase, Sqoop, FlumeOozie, Zoo keeper, Base, cluster health, monitoring security, Redhat Linux, impala, Cloudera Manager,Hortonworks.

Confidential, Los Angeles, CA

Linux/ Hadoop Admin

RESPONSIBILITIES:

  • Working on multiple projects spanning from Architecting, Installation, Configuration and Management of Hadoop Clusters.
  • Implemented authentication and authorization service using Kerberos authentication protocol.
  • Integrated Oozie with the rest of the Hadoop stack supporting several types of Hadoop jobs out of the box (like Map Reduce, Pig, Hive, Sqoop) as well as system specific jobs.
  • Developing data pipeline using Flume, Sqoop, Pig and Java map reduce to ingest customer behavioral data and financial histories into HDFS for analysis.
  • Migrated data from SQL Server to HBase using Sqoop.
  • Log data Stored in HBase DB is processed and analyzed and then imported into Hive warehouse, which enabled end business analysts to write HQL queries.
  • Integrated Kafka with Flume in sand box Environment using Kafka source and Kafka sink.
  • Analyzed the alternatives for NOSQL Data stores and intensive documentation for HBASE vs. Accumulo data stores.
  • Involved in setup, installation, configuration of OBIEE 11g in Linux operating system also integrating with the existing environment. Involved in trouble shooting of errors encountered. And worked with Oracle support to analyze the issue.
  • Maintained multiple application servers with latest version of Confidential EnterpriseLinux4.x-7.x.
  • Involved in migrating applications from Solaris toLinux( Confidential EnterpriseLinux).
  • Developed various workflows using custom MapReduce, Pig, Hive and scheduled them using Oozie.
  • Responsible for Installing, setup and Configuring Apache Kafka and Apache Zookeeper.
  • Extensive knowledge in troubleshooting code related issues.
  • Developed suit of Unit Test Cases for Mapper, Reducer and Driver classes using MR Testing library.

ENVIRONMENT: Hadoop, HDFS, Map Reduce, Shell Scripting, Spark, Splunk, Solr, Pig, Hive, HBase, Sqoop, Flume, Oozie, Zoo keeper, cluster health, monitoring security, RedHat Linux, Cloudera Manager.

Confidential

Linux/Unix Systems Administrator

RESPONSIBILITIES:

  • Installed, Configured and Maintained Debian/RedHat Servers at multiple Data Centers.
  • Configured Confidential Kick start server for installing multiple production servers.
  • Configuration and administration of DNS, LDAP, NFS, NIS, NIS+ and Send mail on RedHat Linux/Debian Servers.
  • Hands on experience working with production servers at multiple data centers.
  • Involved in writing scripts to migrate consumer data from one production server to another production server over the network with the help of Bash and Perl scripting.
  • Installed and configured monitoring tools Munin and NagiOS for monitoring the network bandwidth and the hard drives status.
  • Implemented the Clustering Topology that meets High Availability and Failover requirement for performance and functionality.
  • Configured, managed ESX VM's with virtual center and VI client.
  • Day - to-day administration on Sun Solaris, RHEL 4/5 which includes Installation, upgrade & loading patch management & packages
  • Assist with overall technology strategy and operational standards for the UNIX domains.
  • Performed day-to-day administration tasks like User Management, Space Monitoring, Performance Monitoring and Tuning, alert log monitoring and backup monitoring.
  • Provides accurate root cause analysis and comprehensive action plans.
  • Manage daily system administration cases using BMC Remedy Help Desk.
  • Provided 24/7 on call support on Linux Production Servers. Responsible for maintaining security on Confidential Linux.
  • Configured Global File System (GFS) and Zetta byte File System (ZFS).
  • Troubleshooting production servers with IPMI tool to connect over SOL.
  • Configured system imaging tools Clonezilla and System Imager for data center migration.

ENVIRONMENT: RHEL 5.x/4.x, Solaris 8/9/10, Sun Fire, IBM blade servers, Web sphere 5.x/6.x, Apache 1.2/1.3/2.x, iPlanet, Oracle 11g/10g/9i, Logical Volume Manager, Veritas net backup 5.x/6.0, SAN Multipathing (MPIO, HDLM, Power path), VM ESX 3.x/2.x.

We'd love your feedback!