We provide IT Staff Augmentation Services!

Hadoop Admin/developer Resume

5.00/5 (Submit Your Rating)

Costa Mesa, CA

PROFESSIONAL SUMMARY:

  • 6 years of overall IT experience that includes 3 years of Big Data experience in ingestion, storage, querying, processing and analysis
  • Have knowledge of loading logs from multiple sources directly into HDFS using tools like Flume
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems and vice - versa.
  • Solid understanding of Hadoop, especially in HDFS, Hive, Map Reduce, Pig, Hbase, and Sqoop
  • Excellent understanding of HDFS, Map Reduce, YARN, and tools including Pig and Hive for data analysis, Sqoop for data migration, Flume for data ingestion, Oozie for scheduling and Zookeeper for coordinating cluster resources.
  • Sound knowledge on EMC storage DXP tools and SFDC.
  • Knowledge on EMC VNX,VMAX,isilon,VMware
  • Built data transform framework using Map Reduce and Pig.
  • Worked on disaster management with Hadoop cluster.
  • Worked with application team via scrum to provide operational support, install Hadoop updates, patches and version upgrades as required.
  • Have the motivation to take independent responsibility as well as ability to contribute and be a productive team member.
  • Experience in building Pig scripts to extract, transform and load data onto HDFS for processing.
  • Knowledge and understanding on industry latest Hadoop ecosystems like Apache Spark integration with Hadoop.
  • Knowledge in job work-flow scheduling and monitoring tools like Oozie
  • Worked with business users to extract clear requirements to create business value.
  • Designed, delivered and helped manage a device data analytics at a very large storage vendor.
  • Having Good knowledge on Single node and Multinode Cluster Configurations.
  • Very Good understanding of SQL, ETL and Data Warehousing Technologies
  • Strong communication, collaboration & team building skills with proficiency at grasping new Technical concepts quickly and utilizing them in a productive manner.
  • Storage allocation on Windows, LINUX, HP-ux
  • Extremely motivated with good interpersonal skills; have ability to work in strict deadlines
  • Experienced with Hadoop internals (MapReduce (YARN), HDFS),
  • Extended Hive and Pig core functionality
  • Good understanding of OSI Model, TCP/IP protocol suite (IP, ARP, ICMP, TCP, UDP, SMTP, FTP, TFTP)
  • Experienced in creating a detailed design document to translate the requirements into technical design.
  • Developed Pig Latin scripts using operators such as LOAD, STORE, DUMP, FILTER, DISTINCT, FOREACH, GENERATE, GROUP, COGROUP, ORDER, LIMIT, UNION, SPLIT to extract data from data files to load into HDFS
  • Experience in developing solutions to analyze large data sets efficiently
  • Installed Oozie workflow engine to run multiple Hive and Pig jobs that run independently with time and data availability.
  • Experience to work under tight deadlines and rapidly changing priorities with proactive, creative & focused approach to business needs with high analytical, inters personal, and team playing skills.

TECHNICAL SKILLS:

Operating System: Mac OS, LINUX,Windows,UNIX

Domain: Storage, Cloud, Big Data, Java, HDFS, Hive, Pig, HBase

Big Data Ecosystem: Hadoop, MapReduce, HDFS, HBase, Ambari, Spark, Hive, Pig, Sqoop, Oozie, and Flume

Tools: Eclipse, Dropbox, SFDC, Ubuntu

Programming Languages: C, Java, SQL, JavaScript Python, PL/SQL

WORK EXPERIENCE:

Confidential, Costa Mesa, CA

Hadoop Admin/Developer

Responsibilities:

  • Monitored health of all the Processes related to Name Node HA, HDFS, Yarn, Pig, Hive, and SPARK using Cloudera Manager.
  • Monitored disk, Memory, Heap, CPU utilization on all Master and Slave machines using Cloudera Manager and took necessary measures to keep the cluster up and running on 24/7 basis.
  • Monitored all MapReduce Write Jobs running on the cluster using Cloudera Manager and ensured that they were able to write the data to HDFS without any issues and Data getting evenly distributed over the cluster.
  • Monitored all MapReduce Read Jobs running on the cluster using Cloudera Manager and ensured that they were able to read the data to HDFS without any issues.
  • Monitored workload, job performance and capacity planning using Cloudera Manager.
  • Involved in adding new node to a cluster and decommissioning of the effective nodes from the cluster.
  • Provided Statistics of all successfully completed jobs in detail report format.
  • Provided Statistics of all failed jobs in detail report format and worked on finding the root cause and resolution. Jobs failure due to disc errors, node issues etc.
  • Viewed the performance of the Map and Reduced task that make up the job using Cloudera Manager.
  • Involved in Analyzing system failures, identifying root causes, and recommended course of actions.
  • Fine Tuned JobTracker by changing few properties mapred-site.xml.
  • Fine Tuned Hadoop cluster by setting proper number of map and reduced slots for the TaskTrackers.
  • Migrated Hadoop Cluster from CDH 3.X.X to CDH 4.X.X
  • Integrated Kerberos into Hadoop to make cluster more strong and secure from unauthorized users.
  • Configured user authentication for accessing web UI
  • Involved in Installing Cloudera Manager, Hadoop, Zookeeper, HBASE, HIVE, PIG etc.
  • Experience working on processing unstructured data using Pig and Hive.
  • Developed Pig Scripts, Pig UDFs and Hive Scripts, Hive UDFs to analyze HDFS data
  • Wrote pig jobs to transform data and dump into Hbase.

Environment: Apache Hadoop, Java, MySql, Windows, UNIX, Sqoop, Hive, Oozie

Confidential, Beverly Hills, CA

Hadoop Admin/Developer

Responsibilities:

  • Handled importing of data from various data sources, performed transformations using Hive, Map Reduce, loaded data into HDFS and Extracted the data from MySQL into HDFS using Sqoop.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Analyzed the data by performing Hive queries and running Pig scripts to know user behavior.
  • Created partitioned tables in Hive.
  • Worked on Installed and configured Hadoop 0.22.0 Map Reduce, HDFS, developed multiple Map Reduce jobs in java for data cleaning and preprocessing.
  • Importing and exporting data into HDFS and HIVE using Sqoop.
  • Responsible for manage data coming from different sources
  • Monitoring the running Map Reduce programs on the cluster.
  • Responsible for loading data from UNIX file systems to HDFS.
  • Installed and configured Hive and also wrote Hive UDFs.
  • Involved in creating Hive Tables.
  • Implemented the workflows using Apache Oozie framework to automate tasks.
  • Involved in Installing, Configuring Hadoop ecosystem, and Cloudera Manager using CDH3 Distribution.
  • Involved in requirements gathering and analysis; designed the architecture.
  • Responsible for developing data pipeline using Flume, Sqoop and Pig to extract the data from Facebook feeds and tweets and loaded to HDFS.
  • Involved in processing the ingested raw data using Pig extensively for data cleansing, in order to retrieve the bank related comment.
  • Developed MapReduce programs to processed the cleansed data for analyzing the polarity and emotion.
  • Loaded large sets of semi structured result in HBase and displayed graphical representation to the user.
  • Preparing weekly status and monthly status report
  • Attending Defect calls to provide latest status to client

Environment:: Java 1.7, Spring, Hadoop, YARN, Hive, noSQL, UNIX, Git Cloudera CDH3, Flume, Sqoop, Pig, HDFS, MapReduce, HBase.

Confidential, Bentonville, Arkansas

Hadoop Developer/Admin

Responsibilities:

  • Requirements gathering & analysis
  • Involved in the Software Development Life Cycle (SDLC) of the project from Analysis, Design, Implementation and Testing.
  • Designing fact and dimensions schema models on hadoop.
  • Creating and scheduling end-to-end workflows using Oozie.
  • Designing hive and Impala tables in parquet format for faster querying and reporting
  • Monitoring and keeping Hadoop cluster healthy.
  • Developed Hive DDLs and DMLs to populate data.
  • Developed shell script to pull data from RDMS and apply the incremental and full load to the Hive tables.
  • Technical design of Hadoop end to end job flows and integration components
  • Writing sqoop scripts to extract data from existing RDBMS source Oracle and store it in HDFS
  • Writing Pig scripts to perform various data transformations (ETL tasks) on input data and store output in hive tables.
  • Design and execution of Java Mapreduce programs to perform complex business computations related to revenue management systems which deals with customer sensitive information
  • Worked on Java MapReduce jobs to process in coming data and load it into HBase and Hive tables.
  • Designing, creating, loading Hive tables (for structured data)
  • Design and Develop Spark, Scala applications to replace long running Mapreduce jobs
  • Including and Executing Spark applications using Oozie flow.
  • Development of Pig/Hive UDFs for complex computations
  • Extensively used Hue and Cloudera Manager in CDH
  • Designing and scheduling end-to-end Oozie workflows for various subject areas
  • Demonstration of End to End workflows for target Audience
  • Optimizing the code and ensuring production stability

Environment: Cloudera (CDH4), HDFS, Map Reduce, Java, Pig, Hive, HBase, Sqoop, Flume, Spark, Scala, Oozie, Hue, SVN, D3JS and R

Confidential, CA

Linux System Admin

Responsibilities:

  • Involved in preparation of technical design specifications, installation and configuration of Red Hat Enterprise Linux.
  • Worked with Database administrators to tune kernel for Oracle installations on Linux.
  • Involved in troubleshooting the performance issues related to web application running on virtual machines. Publishing the application on App Store and then pushing updates after every sprint release.
  • Performed routine checks on nodes by monitoring sys logs and error logs for system and hardware errors.
  • Worked with Firewall group, Web group to maintain their software on these systems.
  • Up gradation of Operating System Environment/ kernel patches, packages, web server, and applications.
  • Designed Shell scripts for process automation of databases, applications, backup and scheduling.
  • Configuring IP connectivity, routing, checkpoint firewall and network interfaces. Also maintaining network connectivity of servers.
  • Implemented software RAID at install-time and run-time on Linux.
  • Perform system administration tasks on Linux servers such as configuring logs and applications, setting up new drives, and user management
  • Analyze logs and use Linux and monitoring tools to troubleshoot and debug problems related to applications, Chef, Ubuntu, and TCP/IP
  • Involved in software development life cycle (SDLC) development.
  • Managed existing documentation for systems and created new procedures to support new products. Created documentation for disaster recovery project
  • Automating many day to day tasks through Bash scripting.
  • Worked closely with DBA Team in order to adjust kernel parameters as per requirements
  • Troubleshooting network administration, IIS configuration, DNS setup and modifications, firewall rule sets, local and distributed director, connectivity, and supporting applications
  • Remote Desktop Monitoring using Microsoft Terminal Services/Client.

Environment: Red hat Enterprise Linux 4.x/5.x/6.1, AIX 6.x, windows 2008 R2 & 2008 servers, Windows 2003, IIS 7.0 & 7.5

We'd love your feedback!