We provide IT Staff Augmentation Services!

Sr. Hadoop Developer/admin Resume

0/5 (Submit Your Rating)

Warwick, RI

SUMMARY

  • Over 7+ years of professional experience in Software Development & Requirement Analysis in Agile work environment with 4+ years of Big Data Ecosystems experience in ingestion, storage, querying, processing and analysis of Big Data.
  • Experience in dealing with Apache Hadoop components like HDFS, MapReduce, Hive, HBase, Pig, Sqoop,Nifi,Oozie, Python, Spark, Cassandra,MongoDBetc.
  • Good understanding/knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, Secondary Name nodeandMapReduce concepts.
  • Experienced managing No - SQL DB on largeHadoopdistribution Systems such as: Cloudera, Hortonworks, Map M series etc.
  • Experienced developingHadoopintegration for data ingestion, data mapping and data process capabilities.
  • Extensive work in ETL process consisting of data transformation, data sourcing, mapping, conversion.
  • Strong understanding of Data Modeling and experience with Data Cleansing, Data Profiling and Data analysis.
  • Experience working on Docker hub, creating Docker images and handling multiple images primarily for middleware installations and domain configuration.
  • Experience in extracting source data from Sequential files, XML files, Excel files, transforming and loading it into teh target data warehouse.
  • Strong experience with Java/J2EE technologies such as Core Java, JSP, HTML, JavaScript, JSON.
  • Involved in databasedesign, creating Tables, Views, Stored Procedures, Functions, Triggers and Indexes.
  • Good understanding of service oriented architecture (SOA) and web services like XML, XSD, XSDL, SOAP.
  • Good Knowledge about scalable, secure cloud architecture based on Amazon Web Services, leveraging AWS cloud services: EC2, Cloud Formation, VPC, S3, etc.
  • Good Knowledge onHadoopCluster architecture and monitoring teh cluster.
  • Expertise in setting up standards and processes forHadoopbased application design and implementation.
  • Experience in importing and exporting data using Sqoop from Relational Database Systems to HDFS and vice-versa.
  • Experience in managingHadoopclusters using Cloudera Manager.

TECHNICAL SKILLS

Big Data Technologies: HDFS, Hive, AWS, Map Reduce, Pig, Sqoop, Oozie, Spark.

Scripting Languages: Shell, Python, Perl.

Tools: Quality center v11.0\ALM, TOAD, JIRA, HP QTP, Selenium, Test NG, JUnit.

Programming Languages: Java, C, C++, SQL, PL/SQL, No SQL.

QA methodologies: Waterfall, Agile, V-model.

Front End Technologies: HTML, CSS, XML, JavaScript.

Java Frameworks: MVC, jQuery, Apache Struts2.0, spring and Hibernate.

Defect Management: Jira, Quality Center.

Domain Knowledge: GSM, WAP, GPRS, CDMA and UMTS (3G).

Web Services: SOAP (JAX-WS), SOA, Restful (JAX-RS).

Application Servers: Apache Tomcat, Web Logic Server, Web Sphere, JBoss.

Databases: Oracle 11g, MySQL, IBM DB2, NoSQL Databases,HBase,Nifi,MongoDB,Cassandra, Data Enterprise 4.6.1.

Operating Systems: Linux, UNIX, MAC, Windows NT/98/2000/XP/Vista, Windows 7.

PROFESSIONAL EXPERIENCE

Sr. Hadoop Developer/Admin

Confidential, Warwick, RI

Responsibilities:

  • Currently working as admin on Cloudera (CDH 552) distribution for 4 clusters ranges from POC to PROD.
  • Responsible for Cluster maintenance, Monitoring, commissioning and decommissioning Data nodes, Troubleshooting, Manage and review data backups, Manage & review log files.
  • Day to day responsibilities includes solvingdeveloperissues, deployments moving code from one environment to other environment, providing access to new users and providing instant solutions to reduce teh impact and documenting teh same and preventing future issues.
  • Collaborating with application teams to install operating systemandHadoopupdates, patches, version upgrades.
  • Worked with cloud services like Amazon Web Services (AWS) and involved in ETL, Data Integration and Migration.
  • Monitored workload, job performance and capacity planning using Cloudera Manager.
  • Worked onDockerto drive local development service instances (e.g. Cassandra, Elastic search) and containerization of build pipelines.
  • Involved in Analyzing system failures, identifying root causes, and recommended course of actions.
  • Used Spark-Streaming APIs to perform necessary transformations and actions on teh fly for building teh common learner data model which gets teh data from Kafka in near real time and Persists into Cassandra.
  • Imported logs from web servers with Flume to ingest teh data into HDFS.
  • InstalledDockerRegistry for local upload and download ofDockerimages and even fromDockerhub.
  • Retrieved data from HDFS into relational databases with Sqoop Parsed cleansed and mined useful and meaningful data in HDFS using MapReduce for further analysis.
  • Implemented custom interceptors for flume to filter data and defined channel selectors to multiplex teh data into different sinks.
  • Partitioned and queried teh data in Hive for further analysis by teh BI team.
  • Involved in extracting teh data from various sources intoHadoopHDFS for processing.
  • Worked on analyzingHadoopcluster and different big data analytic tools including Pig, HBase database andSqoop.
  • Creating collections and configurations, register a Lily HBase Indexer configuration with teh Lily HBase Indexer Service.
  • Commissioned and Decommissioned nodes on CDH5Hadoopcluster on Red hat LINUX.
  • Experience in configuring teh Storm in loading teh data from MYSQL to HBASE using JMS.
  • Exported teh analyzed data to teh relationaldatabases using Sqoop for visualization and to generate reports for teh BI team.
  • Installed Oozie workflow engine to run multiple Hive and pig jobs.
  • Troubleshooting, debugging & fixing Talend specific issues, while maintaining teh health and performance of teh ETL environment.
  • Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.

Environment: HDFS, Map Reduce, Hive, Hue, Pig, AWS, Cassandra, Flume, Oozie, Sqoop, CDH552, ApacheHadoop26, Spark, Storm, Cloudera Manager, Red Hat, MySQL and Oracle.

Hadoop Developer/Admin

Confidential, Tampa, FL

Responsibilities:

  • Working as administrator in Hortonworks (HDP 2242) distribution for 10 clusters ranges from POC to PROD.
  • Hands on experience in writing MR jobs for cleansing teh data and to copy it to AWS cluster form our cluster.
  • Worked onDockercontainer snapshots, attaching to a running container, removing images, managing directory structures and managing containers.
  • Provided security and autantication with ranger where ranger admin provides administration and user sync adds teh new users to teh cluster.
  • Good troubleshooting skills on Hue, which provides GUI fordevelopers/business users for day to day activities.
  • Developed MapReduce programs to cleanse teh data in HDFS obtained from heterogeneous data sources to make it suitable for ingestion into Hive schema for analysis.
  • Setup flume for different sources to bring teh log messages from outside toHadoopHDFS.
  • Implemented NameNode HA in all environments to provide high availability of clusters.
  • Working experience on maintaining MySQLdatabases creation and setting up teh users and maintain teh backup of cluster metadata databases with Cron jobs
  • Setting up MySQL master and slave replications and halping business applications to maintain their data in MySQL Servers.
  • Managed and reviewed Log files as a part of administration for troubleshooting purposes Communicate and escalate issues appropriately.
  • As an admin followed standard Back up policies to make sure teh high availability of cluster.
  • Involved in Analyzing system failures, identifying root causes, and recommended course of actions.
  • Monitored multiple clusters environments using Ambari Alerts, Metrics and NagiOS.

Environment: Hadoop, Map Reduce, AWS, HDFS, Pig, Hive, HBase, MapReduce, Flume, Hortonworks, Eclipse, MYSQL, UNIX Shell Scripting

Hadoop Developer

Confidential, Pleasanton, CA

Responsibilities:

  • Installed and configuredHadoopMapReduce, HDFS and developed multiple MapReduce jobs in Java for data cleansing and pre-processing.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Proactively monitored systems and services, architecture design and implementation ofHadoopdeployment, configuration management, backup, and disaster recovery systems and procedures.
  • Used Flume to collect, aggregate, and store teh web log data from different sources like web servers, mobile and network devices and pushed to HDFS.
  • Supported Map Reduce Programs those are running on teh cluster.
  • Involved in loading data from UNIX file system to HDFS, configuring Hive and writing HiveUDF’s.
  • Built Big Data solutions using HBase handling multiple records.
  • Created HBase tables to store variable data formats from millions of data rows.
  • Developed data pipeline using Flume and Javamap reduce to ingest employee browsing data into HBase/HDFS for analysis.
  • Utilized Java and MySQL from day to day to debug and fix issues with client processes.
  • Extracted data from oracle database and spreadsheets and staged into a single place and applied business logic to load them in teh central oracle database.
  • Implemented partitioning, dynamic partitions and buckets in pig and HIVE
  • Used Hive and Pig to analyze data in HDFS to identify issues and behavioral patterns.
  • Created internal and externalHive tables and defined static and dynamic partitions for optimized performance.
  • Handled different types of joins in Scala like Map joins, bucker map joins, sorted bucket map joins.
  • Implemented different machine learning techniques in Scala and using machine learning library.
  • Working knowledge in creating Stored Procedures, Triggers, User-Defined Functions, Views, Indexes, User Profiles, Analytical Functions using T-SQL, SQL Server, PL/SQL.
  • Worked with QA lead/managers to designing automation testing big data jobs.

Environment: Hadoop, MapReduce, HDFS, Hive, Scala, Sqoop,CouchDB, Flume, Tomcat 6., SQL language, Oracle, XML, Eclipse.

Hadoop Developer

Confidential, Bowie, MD

Responsibilities:

  • Involved with ingesting data received from various relational database providers, on HDFS for analysis and other big data operations.
  • Wrote MapReduce jobs to perform operations like copying data on HDFS and defining job flows on EC2 server, load and transform large sets of structured, semi-structured and unstructured data.
  • Creating Hive tables to import large data sets from various relational databases using Sqoop and export teh analyzed data back for visualization and report generation by teh BI team.
  • Developed Spark scripts by using Java, and Python,shellcommands as per teh requirement.
  • Used Spark API over ClouderaHadoopYARN to perform analytics on data in Hive.
  • Developed Scala scripts using both Data frames/SQL/Data sets and RDD/MapReduce in Spark 1.6 for Data Aggregation, queries and writing data back into OLTP system through Sqoop.
  • Optimizing of existing algorithms inHadoopusing Spark Context, Spark-SQL, Data Frames and Pair RDD's.
  • Response to value-added services based on clients' profiles and purchasing habits.
  • Design and implement MapReduce jobs to support distributed processing using JAVA, Hive and Apache Pig.
  • Maintenance of data importing scripts using Hive and MapReduce jobs.
  • Developed and maintain several batch jobs to run automatically depending on business requirements.
  • Unit testing and Deploying for internal usage monitoring performance of solution.

Environment: ApacheHadoop, Hive, PIG, HDFS, Spark, Scala, Java Map-Reduce, Core Java, GIT, Jenkins, UNIX, MYSQL, Eclipse, Sqoop and Cloudera Distribution and MySql.

JAVA Developer

Confidential

Responsibilities:

  • Involved in teh Business Requirement Analysis, Design, Coding, Testing and Support.
  • Implemented Agile Methodology that includes weekly meeting with business analysts and monthly sprint review with clients.
  • Used Spring MVC Framework to develop teh application by implementing Controller, Services classes.
  • Used Annotation based Spring Framework for auto wiring and injecting teh required dependencies to implement business logic.
  • Involved in writing Hibernate Annotations and Hibernate Configuration files to persist data into database.
  • Used Hibernate Query language(HQL) to perform queries against teh database.
  • Worked as one of teh Core Developers of teh team.
  • Used JSP and JavaScript to develop teh front end.
  • Used CVS for version control across common source code used by developers.
  • Used different design patterns like Data Access Object (DAO), Data Transfer Object (DTO) and Business Delegate to develop teh application.
  • Developed user interface using HTML, jquery, Ajax and JavaScript.
  • Prepared Test Cases and Unit Testing performed using Junit.
  • Applied partial business logic writing Stored Procedures and Functions using PL SQL in Oracle DB.
  • Exposed and wrote services as RESTful web service.
  • Used various SQLscripts for querying Tables and modifying teh tables using oracledeveloper.
  • Used ANT scripts to build and deploy application.
  • Extensively Worked with Eclipse and Jboss to develop and deploy teh complete application.

Environment: and Tools: j2EE 1.4, Struts 2, Hibernate, Spring, JavaScript, SOAP, WSDL, JSP, JSTL, Oracle, CSS, HTML, DHTML, JUnit, CAST, WASD.

We'd love your feedback!