We provide IT Staff Augmentation Services!

Bigdata Architect Resume

2.00/5 (Submit Your Rating)

PROFESSIONAL CAREER:

  • 13 years of IT experience currently working as Hadoop Administration & Big Data Technologies.
  • Hands on experience in installation, configuration, supporting and managing Hadoop Clusters using Horton works, Cloudera.Hadoop Cluster capacity planning, performance tuning, cluster Monitoring, Troubleshooting.Backup configuration and Recovery from a NameNode failure.
  • Excellent command in creating Backups & Recovery and Disaster recovery procedures and Implementing BACKUP and RECOVERY strategies for off - line and on-line Backups.Involved in bench marking Hadoop cluster file systems with various batch jobs and workloads.Making Hadoop cluster ready for development team working on POCs.
  • Experience in minor and major upgrades of Hadoop and Hadoop eco system
  • Experience monitoring and troubleshooting issues with Linux memory, CPU, OS, storage and networkHands on experience in analyzing Log files for Hadoop and eco system services and finding root cause.Experience on Commissioning, Decommissioning, Balancing, and Managing Nodes and tuning server for optimal performance of the cluster.As an admin involved in Cluster maintenance, trouble shooting, Monitoring and followed proper backup& Recovery strategies.
  • Good Experience in setting up the Linux environments, Password less SSH, Creating file systems, disabling firewalls, swappiness, Selinux and installing Java.Good Experience in Planning, Installing and Configuring Hadoop Cluster in Cloudera and Hortonworks Distributions.
  • Installing and configuring Hadoop eco system like pig, hive.
  • Experience in importing and exporting the data using Sqoop from HDFS to Relational Database systems/mainframe and vice-versa.
  • Experience in importing and exporting the logs using Flume from outside sources using different flume agents.
  • Hands on experience in Zookeeper and ZKFC in managing and configuring in NameNode failure scenariosExperience in deploying Hadoop 2.0(YARN).

PROFESSIONAL EXPERIENCE:

Confidential

Bigdata Architect

Responsibilities:

  • Design, deploy and architect big data clusters and ingestion frameworks
  • Responsible for 4 clusters ranging from Dev, Test, Stage and prod
  • Responsible for Cluster maintenance, Monitoring, commissioning and decommissioning Data nodes, Troubleshooting, Manage and review data backups, Manage & review log files. Changing the configurations based on the requirements of the users for the better performance of the jobs
  • Automated deployments using cluster shell, Ansible playbooks cut down the execution time by
  • 50%
  • Designed, Installed and configured MAPR cluster.
  • Cluster benchmarking with memory, network, IO and disk tests and comparing them with DFSio and sorting tests.
  • Implementing POC for big data tools like CASK.
  • Configured Spark on Yarn on a MAPR Cluster decreasing the spark programs run time by 50% Configuring security with Truststores and keystores for the DATA and REST.
  • Implementing security using Kerberos and Active Directory
  • Created Authorization policies for all the Ecosystem components in Ranger, Sentry like tools Troubleshooting developer issues on a daily basis addressing performance bottlenecks and mentoring on using the big data tools
  • Implemented cluster on a virtual environment using SAN disks
  • Infrastructure support for the whole MDA clusters integrating with various KPIs like custom
  • Ingest framework, Tableau, SSIS and CDAP.
  • Performed upgrades, backups on Ambari and HDP Stacks on secured cluster Capacity scheduling by setting up queues for different teams making sure enough cluster resources are assigned based on the team's requirements like the variety, velocity and volume
  • Configuring clusters by Setting up Users, Groups and System Settings, configure topology, Configuring Volumes and Jobs logs and scheduling
  • Data access and protection setting up data access, configure client NFS Access, configure and setup control access to the cluster, configure snapshots and mirrors
  • Monitoring the cluster by configuring and responding to Alarms, balancing cluster resourcesmanaging logs and snapshots, adding and removing services
  • Replacing Failed Disks, removing disks, perform node maintenance and adding nodes

Skills: Microsoft Office Suite, SSIS, SAN, POC, Rest

Confidential

Bigdata Admin

Responsibilities:

  • Responsible for Cluster maintenance, Monitoring, commissioning and decommissioning Data nodes, Troubleshooting, Manage and review data backups, Manage & review log files.
  • Changing the configurations based on the requirements of the users for the better performance of the jobs
  • Automated deployments by writing Ansible playbooks
  • Tuning Hadoop cluster for multi-tenancy and optimal resource utilization, performance and high availability
  • Troubleshooting Hadoop/Yarn/MapReduce
  • Experience tuning and troubleshooting MapReduce and general YARN jobs for optimal performance, both on the Command Line as well as via provided UI's.
  • Providing High Availability NameNodes using Quorum Journal Managers and Zookeeper Failover Controllers Extensive knowledge of various configurations parameters available in mapred-size.xml and yarn-site.xml and locations of various log files Usage of JobHistoryServer (and newer App Timeline Server) UI & API to monitor cluster usage and overall job health, forensics etc
  • Implemented Dominant Resource Calculator and capacity scheduler queues to enhance the performance of the cluster
  • Experience with Hive, HiveServer2, and the Hive Metastore, tuning and optimizing. Knowledge of
  • Hive SQL and SQL syntax and usage (selects, inserts, joins). Bonus points for experience with Tez, and newer file formats such as ORC (Optimized Row Columnar) and why they may be superior to Text, Sequence and RC formats, but also what are some tradeoffs
  • Using Ambari administering large Hadoop clusters (> 100 to > 1000's of physical nodes) Ambari REST API for automating common tasks along with the monitoring of the overall cluster health. ozie knowledge and experience developing, deploying Oozie workflows, including coordinator flows and Oozie actions such as the HiveAction
  • Deploying and maintaining a zookeeper ensemble within and without Hadoop Experience using Sqoop to transfer data into and out of HDFS & RDBMS Experience tuning & maintaining HBase
  • Adding and removing Data Nodes and Services using Ambari

Skills: Hive, SQL, Hadoop HDFS, UI/UX, restAPI

Confidential

Hadoop Admin

Responsibilities:

  • Currently working as admin in HortonWorks Hadoop distribution for 4 clusters ranges from POC to PROD.
  • Responsible for Cluster maintenance, Monitoring, commissioning and decommissioning Data nodes, Troubleshooting, Manage and review data backups, Manage & review log files.
  • Day to day responsibilities includes solving developer issues, providing access to new users and providing instant solutions to reduce the impact and documenting the same and preventing future issues.
  • Experienced on adding/installation of new components and removal of them through Ambari.
  • Monitoring systems and services through Ambari dashboard to make the clusters available for the business.
  • Architecture design and implementation of deployment, configuration management, backup, and disaster recovery systems and procedures.
  • Hand on experience on cluster up gradation and patch upgrade without any data loss and with proper backup plans.
  • Changing the configurations based on the requirements of the users for the better performance of the jobs.
  • Experienced in Ambari-alerts configuration for various components and managing the alerts.
  • Provided security and authentication with ranger where ranger admin provides administration and user sync adds the new users to the cluster.
  • Good troubleshooting skills on Hue, which provides GUI for developers/business users for day to day activities.
  • Developed MapReduce programs to cleanse the data in HDFS obtained from heterogeneous data sources to make it suitable for ingestion into Hive schema for analysis.
  • Implemented complex MapReduce programs to perform joins on the Map side using distributed cache
  • Setup flume for different sources to bring the log messages from outside to Hadoop Hdfs.
  • Implemented NameNode HA in all environment to provide high availability of clusters.
  • Capacity scheduler implementation in all environments to provide balancing on resources allocation

Skills: Gui, Mapreduce, Hadoop HDFS, Ambari, POC

Confidential

QA Lead/Scrum Master

Responsibilities:

  • Guided Scrum team define their purpose, so that the team finds their own definition of success for an organization new to Agile methodology Facilitate daily stand-ups and scrum ceremonies for two scrum teams.
  • Worked closely with Product Owners, coordinating product backlog grooming and story estimation.
  • Report Confidential daily Scrum of Scrum meetings. Track burn down, issues and progress in TFS. Work with component teams to resolve issues.
  • Improved team velocity by incorporating capacity planning into sprint planning sessions.
  • Implemented Issues Tracking in TFS to effectively track impediments against user stories.
  • Successfully accelerated Scrum team to a successful project completion faster than Waterfall with a minimum number of defects for SL and Conventional products
  • Organized and facilitated reviews, retrospectives, release planning, demos and other Scrum-related meetings.
  • Facilitate scrum ceremonies (grooming, sprint planning, retrospectives, daily stand-ups, etc.)
  • Removed team impediments on a daily basis to allow the team to deliver the sprint goals and deliverables.
  • Organizing all testing efforts and working with different teams to ensure the success of the product
  • Work with Project Managers/Owners to work thru requirement needs for the clients and provide high level requirements Track and communicate team velocity and sprint/release progress.
  • Assisted team with making appropriate commitments through story selection, sizing and task definition and participated proactively in developing and maintaining team standards, tools and best practices reducing development time.
  • Communicate with other management, engineers, product managers and support specialists on product issues.
  • Adapted SCRUM processes, as needed, acted as a change agent (technical, application) Identified product backlog items, led daily Scrums, monitored and adapted sprint execution

Skills: Scrum, SQL, TFS, Agile, Waterfall

We'd love your feedback!