We provide IT Staff Augmentation Services!

Hadoop Administrator Resume

2.00/5 (Submit Your Rating)

Dublin, OH

SUMMARY

  • Over 6 years of experience as Informatica administrator wif around 2 years experience in Big Data Hadoop and NoSQL technologies.
  • Experience in all the phases of Data warehouse life cycle involving Requirement Analysis, Design, Coding, Testing, and Deployment.
  • Experience in working wif business analysts to identify study and understand requirements and translated them into ETL code in Requirement Analysis phase.
  • Experience in installation, configuration and management of Hadoop Clusters
  • Good Knowledge on Hadoop Distribution Cloudera
  • Extensive experience in architect the Hadoop cluster.
  • Practical knowledge on functionalities of every Hadoop daemon, interaction between them, resource utilizations and dynamic tuning to make cluster available and efficient.
  • Experience in understanding and managing Hadoop Log Files
  • Experience in managing the Hadoop infrastructure wif Cloudera Manager.
  • Experience in writing custom scripts to check the impact of any changes to cluster.
  • Experience in providing security for Hadoop Cluster wif Kerberos.
  • Experience in setting up the monitoring tools such as Nagios and Gangalia to monitor and analyze the functioning of cluster.
  • Experience in setting up data gathering tools such as Flume and Sqoop
  • Experience in setting up and managing the batch scheduler Oozie.
  • Good understanding of NoSQL databases such as Hbase and Cassandra and Querying Language called CQL for Apache Cassandra.
  • Design, implement and review features and enhancements to Cassandra.
  • Deployed a Cassandra cluster in cloud environment as per the requirements.
  • Experience in analyzing data in HDFS through MapReduce, Hive and Pig
  • Experience on UNIX commands and Shell Scripting.
  • Extensively worked on the ETL mappings, analysis and documentation of OLAP reports requirements. Solid understanding of OLAP concepts and challenges, especially wif large data sets.
  • Proficient in Oracle 9i/10g/11g, SQL and PL/SQL.
  • Experience in integration of various data sources like Oracle, DB2, Sybase, SQL server and MS access and non - relational sources like flat files into staging area.
  • Proficient in creating mappings wif Transformations like Source qualifier, Joiner, Sorter, Aggregator, Expression, Lookup, Router, Filter, Update Strategy, Sequence Generator, Normalizer and Rank Informatica Designer and processing tasks using Workflow Manager to move data from multiple sources into targets.
  • Experience in Data Analysis, Data Cleansing (Scrubbing), Data Validation and Verification, Data Conversion, Data Migrations and Data Mining.
  • Excellent interpersonal, communication, documentation and presentation skills.

TECHNICAL SKILLS

Hadoop/Big Data platform: HDFS, MapReduce, Hbase, Cassandra, Hive, Pig, Oozie, Zookeeper, Flume, Sqoop

Hadoop distribution: Cloudera

Programming languages: C/C++, Java, Unix shell scripts, Perl, Pig Latin, PL/SQL, CQL

Databases: Oracle10g/11g, DB2, MySQL, Cassandra, Hbase, Teradata

ETL tools: Informatica 9.x/8.x, Informatica Power Exchange 9.x/8.x

Oracle Tools: Oracle Enterprise Manager, Quest TOAD, SQL*PLUS, SQL*Loader, SQL*Net, SQL Navigator Export/Import

Admin operations: Performance tuning, Storage capacity management, System dump analysis

PROFESSIONAL EXPERIENCE

Confidential, Dublin, OH

Hadoop administrator

Responsibilities:

  • Handle the installation and configuration of a Hadoop cluster.
  • Build and maintain scalable data pipelines using the Hadoop ecosystem and other open source components like Hive, and HBase.
  • Handle the data exchange between HDFS and different web sources using Flume and Sqoop
  • Monitor the data streaming between web sources and HDFS.
  • Monitor the Hadoop cluster functioning through monitoring tools.
  • Close monitoring and analysis of the MapReduce job executions on cluster at task level.
  • Inputs to development regarding the efficient utilization of resources like memory and CPU utilization based on the running statistics of Map and Reduce tasks
  • Changes to the configuration properties of the cluster based on volume of the data being processed and performance of the cluster.
  • Handle the upgrades and Patch updates.
  • Set up automated processes to analyze the System and Hadoop log files for predefined errors and send alerts to appropriate groups.
  • Commission or decommission the datanodes from cluster in case of problems.
  • Set up automated processes to archive/clean the unwanted data on the cluster, in particular on Namenode and Secondary namenode.
  • Set up and manage HA namenode and Namenode federation using Apache 2.0 to avoid single point of failures in large clusters.
  • Set up the checkpoints to gathering the system statistics for critical set ups.
  • Discussions wif other technical teams on regular basis regarding upgrades, Process changes, any Special processing and feedback.

Confidential, OH

Hadoop Engineer

Responsibilities:

  • Participated in design and development of scalable and custom Hadoop solutions as per dynamic data needs.
  • Coordinated wif technical team for production deployment of software applications for maintenance.
  • Provided operational support services relating to Hadoop infrastructure and application installation.
  • Handled the imports and exports of data onto HDFS using Flume and Sqoop.
  • Supported technical team members in management and review of Hadoop log files and data backups.
  • Participated in development and execution of system and disaster recovery processes.
  • Formulated procedures for installation of Hadoop patches, updates and version upgrades.
  • Automated processes for troubleshooting, resolution and tuning of Hadoop clusters.
  • Set up automated processes to send alerts in case of predefined system and application level issues.
  • Set up automated processes to send notifications in case of any deviations from the predefined resource utilization.

Confidential, Newark NJ

Informatica Administrator/Developer

Environment: Informatica Power Center 8.6.1, and Oracle 10G

Responsibilities:

  • Installed and configured PowerCenter 8.6.1 on UNIX platform.
  • Installation and configuration of PowerCenter 8.6.1 hot fixes, patches and version upgrade from PowerCenter 8.5.1 to PowerCenter 8.6.1.
  • Creation and maintenance of Informatica users and privileges.
  • Migration of Informatica Mappings/Sessions/Workflows from Dev, QA to Prod environments.
  • Deployment of Informatica Mappings/Sessions/Workflows into Production environment. Testing the Informatica Objects in QA before moving them into Production.
  • Designed and documented ETL standards document.
  • Worked on shell scripts to sftp (Secured FTP) the file to and from TumbleWeed Server Directory.
  • Worked on SQL queries to query the Repository DB to find the deviations from Company’s ETL Standards for the objects created by users such as Sources,Targets, Transformations, Log Files, Mappings, Sessions and Workflows.
  • Worked on Object Queries, resolving conflicts delete and purge objects.
  • Worked on SQL queries to query the Repository DB to find the deviations from Company’s ETL Standards for the objects created by users such as Sources,re that all support requests are properly approved, documented, and communicated using the Remedy tool. Documenting common issues and resolution procedures.

Confidential, Framingham, MA

ETL Developer

Responsibilities:

  • Creating Informatica Mappings
  • Creating Business Workflows using Informatica
  • Used various transformation such as Filter, Expression, Aggregate, Look-up, sequence, Joiner, Router, Update Strategy, and Stored procedure transformation for better data cleansing and to migrate clean and consistent data.
  • Used debugger to test the mapping and fixed the bugs
  • Developed Mapplets and reusable Transformations
  • Managing, Migrating and Maintaining Informatica Repository in Test and Production environments
  • Testing the Informatica mappings and Workflows
  • Installing and Troubleshooting Informatica Server and Client
  • Preparing & Conducting User Acceptance Testing, Integration & Performance Testing
  • Issue Tracking & Resolution,
  • Change request Tracking & resolution and Documentation

We'd love your feedback!