We provide IT Staff Augmentation Services!

Hadoop Administrator Resume

3.00/5 (Submit Your Rating)

Houston, TX

SUMMARY

  • 7+ years of IT experience and with experience in Operations, developing, maintaining, monitoring and upgradingHadoopClusters in the cloud as well as in - house. (ApacheHadoop, Hortonworks and Cloudera distributions).
  • Built large-scale data processing pipelines and data storage platforms using open-source big data technologies.
  • Involved in all phases of the Software Development Life Cycle (SDLC) and Worked on all activities related to the Operations, implementation, administration, and support of ETL processes for large-scale Data Warehouses.
  • Experience in installation, configuration, management and deployment of Big Data solutions and the underlying infrastructure of theHadoopClusters using Hortonworks Ambari, and Apache Hadoop tarball on Ubuntu, Red Hat, and Centos.
  • Experience in installing, configuring Hive, its services, and Metastore. Exposure to Hive Querying Language, knowledge about tables like importing data, altering and dropping tables.
  • Experience in installing and running Pig, its execution types, Grunt, Pig Latin Editors. Good knowledge about how to load, store, filter data and also combining and splitting data.
  • In-depth knowledge about database imports, worked with imported data to populate tables in Hive. Exposure in how to export data from relational databases toHadoopDistributed File System.
  • Experience in setting up the High-AvailabilityHadoopClusters.
  • Well experienced in helping clients with Big Data Strategy, Capacity Planning and designing Hadoop and Kafka cluster on-premise or cloud or hybrid models.
  • Good knowledge about planning aHadoopcluster like choosing the distribution, hardware selection for both master as well as slave nodes and cluster sizing.
  • Experience in working with Microsoft AZURE .
  • Exposure about how to deploy, manage bothHadoopand Cloudera clusters in the cloud with the help of Microsoft Azure.
  • Experience in developing Shell Scripts for system management.
  • Experience inHadoopadministration with good knowledge aboutHadoopfeatures like safe mode, auditing. Maintenance like both minor as well as major upgrades, rolling back. In-depth knowledge about how Distcp is used for Data Backup as well as recovery.
  • Experience with Software Development Processes & Models: Agile, Waterfall & Scrum Model.
  • Have a knowledge of sprint planning tools like hip chat, Jira and GitHub version control tools as well.
  • Team Player and a fast learner with good analytical and problem-solving skills.
  • Self-Starter and Ability to work independently as well as a Team.
  • Expertise in applying Kerberos security for authentication using local KDC and integrated with Active Directory.
  • Proficiency in integrating Ambari, Ranger and Knox with an LDAP and design & implementing authorization policies using Ranger.
  • Experiencein Benchmarking & performance tuning of the distributed cluster environments.
  • Involved in identifying the bottlenecks in the system and fine-tune them with appropriate directions and performance improvements.

TECHNICAL SKILLS

Languages: Java, C, C++

BigData Technologies: Hadoop, Spark, Scala, Kafka, NIFI, Impala, Hive, Pig, SQL, Storm, Sqoop, Flume, Oozie, Zookeeper, HDF, Cloudera CDH, Hortonworks Ambari, MRUnit,Nifi Registry, Kafka SMM/SAM, CDSW.

Security: Kerberos, Active Directory, LDAP, Apache Ranger, Knox and Sentry

RDBMS: Oracle, MySQL.

No SQL: HBase

Scripting & Query Languages: Python, UNIX Shell, Perl, SQL & PL/SQL.

Cloud Platforms: Azure, AWS, EC2, BLOB and S3

PROFESSIONAL EXPERIENCE

Hadoop Administrator

Confidential, Houston, TX

Responsibilities:

  • Administrate and orchestrated all the Hadoop clusters available in Confidential on-site, 2 research, 2 development, 2 UAT/QA, 2 productions, 1 DR clusters.
  • Installed and configured Apache Ranger with HDFS and Hive plugins for user’s data sets authorizations.
  • Installed and configured Kerberos and Integrated it with AD Server for user’s authorization.
  • Installed and configured KNOX gateway and integrated with AD for users authentication.
  • Buildand evaluatethe applicabilityof a document processing pipeline thatgoes from copying and processing files from different locations to enablesearch in the documents, Ingest files into Hadoopand NLP process files and Indexing the result data into Solrfor searching.
  • Buildand evaluated an 18 node HDF NiFi/Kafka cluster in Azure for a specific use case requirement to ingest and process real time Drilling data into NiFi and write to Kafa/ Azure Datalake.
  • Ingest file system databases with file paths precomputed, augment data through Hive processing, and make a final optimizedfile system database availablethrough Hive.
  • Ingestion of workflows for multiple use cases through Nifi and providing the processed data to meet use case needs.
  • Leveraging Nifi to ingest, enrich and route system and application data for data logistics by configuration the priority, throughput, and backpressure.
  • Integratingthe data throughimplementing the data life cycle management and providing the neededsupport byprovisionstorage,compute and engine for data science workloads.
  • Configured HBase Replication across PROD to DR clusters.
  • Configured and Demonstrated HDP Transparent Data Encryption (TDE) on The HDP 2.xcluster.
  • Automate the Hadoop Cluster, Kafka cluster and Cloud Platform Configuration in Ansible.
  • Enabled the High Availability for HiveServer2, Oozie servers and HBase Thrift Servers.
  • Installed and configuredHadoopsecurity and access controls using Kerberos & Ranger.
  • Supporting multiple pipelines which involve continuous data ingestion toHadoop
  • Provided ad-hoc queries and data metrics to the Business Users using Hive, Pig.
  • Debugging and troubleshooting the issues in the development and Test environments.
  • Experience in upgrading the cluster by providing downtime to theHadoopusers as and when required.
  • Supporting multiple teams who run theirHadoopjobs using Hive and Spark.
  • Conducting root cause analysis and resolve production problems and data issues.
  • Proactively involved in ongoing maintenance, support, and improvements in theHadoopcluster.
  • Used Sqoop to import the verifiable data into hive tables in text and Avro file formats.
  • Upgraded the cluster to both minor/major versions and resolved all the issues with minimum help.
  • Implemented Partitioning, Bucketing in Hive for the better association of the data.
  • Implemented Hbase Bucketing cache to improve the performance and reducing evictions to zero.
  • Implemented a script to run compactions off business hours in Hbase to improve performance during business hours.
  • Implemented a script to perform rolling restarts for cluster components to reduce downtime for the service and maintain the expected SLA.

Environment: Hadoop, Hdfs, Spark, MapReduce, Yarn, Hive, Sqoop, Oozie, Airflow, Nifi, Kafka, Kerberos, Shell Scripting, Unravel.

Big Data Administrator

Confidential, Milwaukee, WI

Responsibilities:

  • Configure and Adminstrate the Hadoop cluster in The enterprise Data lake.
  • Installed and Configured multi node PROD, DR and DEV HDP Hadoop clusters in-house and Cloud environments.
  • Installed and configured the plugins for Apache Ranger and integrated the platform with Kerberos.
  • Integrated the HDP platform with enterprise AD.
  • Configured the encrypted zones in the cluster for each department.
  • Configured the capacity scheduler and defined all YARN queues based on the allocation requirements.
  • Supported the jobs related to data ingestion toHDFS.
  • Provided ad-hoc queries and data metrics to the Business Users using Hive, Pig.
  • Debugging and troubleshooting the issues in the development and Test environments.
  • Experience in upgrading the cluster by providing downtime to theHadoopusers as and when required.
  • Supporting multiple teams who run theirHadoopjobs using Hive Oozie and Spark.
  • Conducting root cause analysis and resolve production problems and data issues.
  • Proactively involved in ongoing maintenance, support, and improvements in theHadoopcluster.
  • Experience in managing and reviewingHadooplog files.
  • Involved in the implementation of theHadoopcluster on AZURE as a part of POC.
  • Worked using various Bigdata technologies on theHadoopEcosysteminHortonworks distribution.
  • Analyzed large amounts of data sets to determine the optimal way to aggregate and report on it.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and generated financial reports using Tableau.
  • Supported in setting up a QA environment and updating configurations for implementing scripts with Pig and Sqoop.
  • Working with Jenkins for any automation builds which are integrated with GIT, Docker as part of infra automation under continuous integration.

Environment: Hortonworks HDP 2.6.1, HDP2.5.3, HDP, MapReduce, Hive, TEZ, Pig, Sqoop, Oozie, Spark, Kafka, NIFI, Scala, Python, S3, Amazon EC2, EMR, RDS and, Azure.

Big Data Admin

Confidential, Chicago, IL

Responsibilities:

  • Installed the cluster in AWS and moved to In House Cluster after completion of POC Phase.
  • Adjusted the configuration parameters related to various services and enabled HA for NN & RM.
  • Configured Logstash to collect log files from many servers and collected log files events as streaming reads and customized Logstash Kafka plugin to push logs in to Kafka
  • Developed Impala/Hive Scripts for implementing dynamic partitions.
  • Developed Sqoop commands/jobs to pull data from Oracle, MYSQL and push to HDFS vice versa using Sqoop shared meta store.
  • Streamlined Hadoop jobs and workflow operations using Oozie workflow
  • Generate reporting data using PowerPivot in excel and SSIS packages for testing by connecting to the corresponding Hive / Impala tables using the ODBC connector.
  • Research and do feasibility studies on the latest tools and developments in the Hadoop ecosystem and recommend them to management.

Environment: Cloudera CDH4, Hadoop, Map Reduce Storm, Kafka, Logstash, Java, Python, Hive, Pig, Sqoop, Impala, Oozie, HBase, Oracle, MySQL, Amazon AWS EC2, and GitHub.

Systems Analyst

Confidential 

Responsibilities:

  • Have been responsible for administering large, multi-site UNIX/LINUX server environments and operating systems, software installation, upgrades, system integrity, security, disaster recovery and performance.
  • Implemented ClouderaHadoopclusters for three different environments on HPE ProLiant servers.
  • Configured/Managed Cisco and Brocade Fabric environment for soft/hard Zoning.
  • Worked on MySQL and successfully launched queries to provide required data to the department.
  • Created users, manage user permissions, maintain User & File System quota on Redhat Linux.
  • Installation & Configuration of Logical Volume Manager - LVM and RAID.
  • Automated administration tasks through use of scripting and Job Scheduling using CRON.
  • Wrote shell scripts for taking data backups, cleaning junk content and updating software regularly.
  • Experience in using protocols/services like Http, Https, TLI/SSL, DHCP, DNS, SSH, SFTP, TCP/IP, FTP/SFTP, SMTP
  • Provided Linux SystemAdministration, Linux System Security, Project Management and Risk Management in InformationSystems.
  • Day to day provisioning of storage including (Storage device/LUN/Volume selection & creation, Fabric Zoning, LUN Masking & Mapping).
  • Administration of environment running VMware ESXi Hosts and Virtual Machines.
  • Worked successfully towards improving and maintaining the Backup Success rate to > 98%.
  • Worked with server teams to insure the configuration and installation of the proper drivers, firmware, and multipath drivers to support the SAN environment.
  • Provide performance tuning and regular maintenance in order to minimize downtime and maximize performance.
  • Take care about Data Center (DC) by ordering and upgrading necessary hardware, supporting RAIDs, maintaining servers and installing new ones.

Environment: Linux, VMware, Storage

We'd love your feedback!