We provide IT Staff Augmentation Services!

Hadoop Developer Resume

0/5 (Submit Your Rating)

Irving, TX

SUMMARY

  • 9+ years of experience in Application analysis, Design, Development,
  • Maintenance/Supporting web and Client - server based applications in Java/J2EE technologies which includes 3 years of experience with Big Data management andHadooprelated components HDFS, MapReduce, Pig, Hive,YARN,Sqoop,Flume, Crunch, Spark, Strom, Scala,kafkabased on Big Data platforms.
  • Delivered 3 projects (full SDLC) using big data technologies like Hadoop, Oozie and NoSQL
  • Extensive working noledge in setting up and running Clusters, monitoring, Data analytics, Sentiment analysis,Predictive analysis, Data presentation with big data world.
  • Expertise in design and implementation of Banking, E-Commerce and Retail domains.
  • Experience managing Cloudera distribution of Hadoop(CDH4&5 using Cloudera Manager),MapR, and Horton works.
  • Excellent understanding of NoSQL databases like HBase,Cassandra and MongoDB.
  • Hands on experience working on structured, unstructured data with various file formats such as Avrodata files, xml files, Jsonfiles, sequence files using Map Reduce programs.
  • Work experience with cloud configurations likeAmazon web services (AWS).
  • Implemented custom business logic and performedjoin optimization,secondary sorting, custom sorting using Map Reduce programs.
  • Experienced testing and running of Map Reducepipelineson Apache Crunch.
  • Composed Map Reduce pipelines with the halp of UDF’s and also implemented joining and data aggregation using Apache Crunch.
  • Developed MapReduce jobs, Used different optimization techniques to improve performance in Map Reduce Programs.
  • Expertise in Data load management, importing & exporting data using SQOOP,Apache Kafka, Spark Streaming and FLUME
  • Implemented business logic using Pig scripts. Wrote custom Pig UDF’s to analyze data
  • Performed differentETLoperations using Pigfor joiningoperations and transformations on data to join, clean, aggregate and analyze data.
  • Experience with Oozie Workflow Engine to automate and parallelize Hadoop, MapReduceand Pig jobs
  • Experienced with handling different file formats like xml, json, RC and ORC formats in Hive using different SerDe's.
  • Extensive experience with wiring SQL queries using HiveQL to perform analytics on structured data.
  • Experience in performing data validation using HIVE dynamic partitioning and bucketing.
  • Experienced in importing and exporting data between RDBMS and TeraData intoHDFS using Sqoop.
  • Experienced in handling streaming data like web server log data using flume.
  • Worked on design and implemented a Cassandra based database and related webservices for storing unstructured data.
  • Experience on YARN environment with Strom, Spark, Kafka and Avro.
  • Hands on experience in Scala, Kafka and Strom.
  • Experience in implementing algorithms for analyzing using spark.
  • Experience in implementingSpark using Scalaand SparkSQL for fasteranalyzing and processing of data.
  • Experience in getting data from various sources into HDFS and building reports using Tableau with spark.
  • Experience in creating tables on top of data on AWS S3obtained from different data sources and providing them to analytics team building reports using Tableau.
  • Extensive Hands on experience with Accessing and perform CURD operations against HBase data using Java APIand implementing time series data management.
  • Experienced in monitoringHadoop clusters using Cloudera ManagerTool.
  • Involved in various datamining tasks such as pattern mining, classification and clustering techniques.
  • Expert noledge over J2EE Design Patterns like MVC Architecture, Singleton, Factory Pattern, Front Controller, Session Facade, Business Delegate and Data Access Object for building J2EE Applications.
  • Extensive experience in Java/J2EE programming - Spring, Hibernate, SOAP/Rest web services,JMS, JNDI, EJB,JAX-WS.
  • Expertise with Application servers and web servers likeOracle Web Logic,IBM WebSphere, Apache Tomcat, JBOSS and VMware.
  • Proven expertise in implementing IOC/Dependency Injection features in various aspects of Spring Framework.
  • Expertise Working with XML Parsers likes SAX, DOM, and XStream
  • Good Knowledge of using IDE Tools like Eclipse, NetBeans,SpringSource Tool for Java/J2EE application development.
  • Experienced in developing the unit test cases usingBeetest,MRUnit, Junitand Easy Mock.
  • KnowledgeonSplunkfor logging mechanism.
  • Knowledge on Build tool Jenkins.
  • Experience in using Mavenand ANTfor build automation.
  • Experience in using version control and configuration management tools like SVN, CVSand Tortoise SVN.
  • Experience working in environments using Agile (SCRUM) andWaterfall methodologies.
  • Good noledge analyzing data usingPython development and scripting for HadoopStreaming.
  • Involved in gathering logs information and managing hardware issues using Netezza.
  • Experience in designing applications using UML Diagrams like Class Diagram, Component Diagram, Sequence Diagrams, and Deployment Diagram using MS Visio, Rational Rose.
  • Expertise in database modeling, administration and development usingSQLand PL/SQL in Oracle (8i, 9i and 10g), MySQL, Teradata, DB2and SQL Server environments.
  • Experienced in using Operating Systems like Windows 98 / 2000 / XP/7/ UNIX,LINUX.

TECHNICAL SKILLS

Hadoop/Big Data: HDFS, Map Reduce, Hive, Pig, YARN,Sqoop, Flume, Oozie, Cloudera manager,Crunch, Strom, Scala, Kafka, Spark, AWS, RHadoop

Database/NoSql: SQL.Pl/SQL,HBase,Cassandra

J2EE Frameworks: Hibernate,Spring, JMS, EJB, JSF

XML/Web Services: SOAP/ Rest

Methodologies: Agile, Waterfall

Build Tools: Maven, ANT, Log4j.¬

Scripting languages: JavaScript, HTML, HTML5, XML, Python, JSP and Servlets

Tools: and other: Rational Rose, Microsoft Visio, CSS

Operating Systems: Linux/Unix,WINDOWS

PROFESSIONAL EXPERIENCE

Confidential

Hadoop Developer

Responsibilities:

  • Involved in Analysis, Design, Development and Testing process based on the new business requirements.
  • Experience managing Cloudera distribution of Hadoop CDH5.5using Cloudera Manager.
  • Wrote java processes for generating different sample log types which generates different random patterns.
  • Wrote script for generating flumeconfg files for different logs types and for different environments.
  • Worked on setting up flume agents in dev and uat environments.
  • Working on data ingestions, integrating with flume.
  • Experience in refactoring the existing spark batch process for different logs written in scala.
  • Writing workflows and scheduling using oozie.
  • Deploying oozie jobs in DEV and UAT for testing the hourly jobs by parsing the generated sample logs in the same environments.
  • End to end testing the ingestion pipeline in UAT from generating logs to saving the transformed data in HDFS.
  • Responsible in handling streaming data like web server log data using flume.
  • Experience in working with Hive for processing the raw data.
  • Involved in daily SCRUM meetings to discuss the development/progress ofSprints and was active in making scrum meetings more productive.

Confidential

Data Engineer

Responsibilities:

  • Experience working with analysts and product owners for getting exact requirement from clients when they provide some new data sources.
  • Extensively used Amazon s3 as cloud storage for all the ingested and transformed data.
  • Migrated all data and tables from Spark 0.9 to Spark 1.4. Created a parallel process for data all ingestions which uses spark 1.4 and tan comparing and validating data to 0.9 there by killing processes in 0.9.
  • Developed java code to get data from different type’s sources which uses REST API, screen scraping from diverse websites depending on type of source and ingested it into a central repository.
  • Responsible for developing Hive queries involving bucketing, partitioning and multiple UDF implementations.
  • Involved in developing UDF’s and UDAF’s to implement customized transformations.
  • Extensively used Sqoop to get data from RDBMS sources like Teradata and Netezza.
  • Validated if data TEMPhas been normalized and loaded into S3 buckets using Hive queries.
  • Wrote custom Hive UDF’s based on business requirement.
  • Developed an automated system which can import data from any type of source to any type of destination.
  • Developed SparkSQL automation components and responsible for modifying java component to directly connect to thrift server.
  • Experience in writing job flows in python which integrates developed java code and shell scripts which run onAzkabanserver.
  • Scheduled and monitored jobs using Azkaban.
  • Experience in spinning up anAmazon EMR cluster with custom configuration and resource allocation.
  • Experience in providing on call support for any cluster issues etc. Also experienced in solving data issues when a job fails in production cluster because of cluster setup issue, working on it on the fly when analytics team creates an incident.
  • Experience is setting up custom slack notifications where notification can be sent to a group/channel or as a personal message for failure jobs.
  • Did a POC on comparing CSV files with parquet files on big data sets by performing aggregations and observing time responses.
  • Experience in solving data issues or defects raised by QA team.
  • Involved in daily SCRUM meetings to discuss the development/progress ofSprints and was active in making scrum meetings more productive.

Confidential

Hadoop Developer

Responsibilities:

  • Participated in Gathering requirements, analyze requirements and design technical documents for business requirements.
  • Involved different phases in big data projects like data acquiring, data processing and data serving using dash boards.
  • Implemented different aggregations and filtering operations using JavaAPI in Cassandra
  • Responsible for Datamodeling in Cassandra and deciding the row key and different column families in Cassandra.
  • Import/export data from Oracle data base to/from HDFS using Sqoop, Hue and JDBC.
  • Gathered data from different sources like Internet, sensors, user behavior using Flume and Kafka and moved to HDFSand implementedOptimized join baseusingMapReduce programs.
  • Implemented Custom Input formats dat handlewide range of input files received from java applications to process in MapReduce.
  • Implemented joins and data aggregation using Apache Crunch.
  • Writing MapReduce pipelineprograms for testing using Apache Crunch.
  • Experienced in unit testingMapReduceprograms using MRUnit.
  • Divided each data set in to corresponding categories by following MapReduce Binning design pattern.
  • Implemented FilterMappers to eliminate un-necessary records and perform data and schema validation.
  • Experience in using Pig as an ETL tool for event joins, filters, transformations and pre- aggregations.
  • Created partitions, bucketing across state in Hive to handle structured data.
  • Implemented Dash boards dat handleHiveQL queries internally like Aggregation functions, basic hive operations, and different kind of join operations.
  • Implemented business logic based on state in Hive using Generic UDF's.
  • Involved in creating data-models for customer data using Cassandra Query Language.
  • Managing and scheduling batch Jobs on a Hadoop Cluster using Oozie.
  • Created production jobs using Oozie work flows dat integrated different actions like MapReduce, Sqoop, and Hiveby utilizing fork and join operations in Oozie.
  • Experience in managing and reviewing Hadoop Log files.
  • Experienced with monitoring Cluster using Cloudera managerandAmbari.
  • Involved in managing Hadoop distributions using Hortonworks.
  • Worked on building BI reports inTableauwith Sparkusing SparkSQL.
  • Experience in deploying data from various sources into HDFS and building reports using Tableau.
  • Implemented Spark using Scalaand SparkSQL for faster testing and processing of data.
  • Implemented data ingestion and handling clusters in real time processing using kafka.
  • Experience with Core Distributed computing and Data Mining Library using ApacheSpark.
  • Experienced in data modelling in hive implementing hive-indexing.
  • Experience in utilizing spark machine learning techniques implemented in scala.
  • Experienced in configuring maven builds dat integrated dependencies check styles, test coverage's.
  • Responsible for analyzing multi-platform applications using python.
  • Developed MapReduce jobs in Python for data cleaning and data processing.
  • Designing Test Plans, Test Cases and performed System Testing.
  • Involved in daily SCRUM meetings to discuss the development/progress of Sprints and was active in making scrum meetings more productive.
  • Experience in integrating RHadoop for categorization and statistical analysis to generate reports.

Confidential

Hadoop Developer

Responsibilities:

  • Involved in Analysis, Design, Development and Testing processof phase 1.
  • Interacting with the client and designing Technical design document from Business Requirements for the development team.
  • Access HBase from Java using Java API to perform CRUD operations on HBase data to perform real time analytics on hospitals.
  • Created hive external tables with data residing in Hbase by implementingHBase-Hive integration.
  • Experienced in loading and transforming large sets of structured, semi-structuredand unstructured data Hadoop concepts.
  • Experienced with different APIs to access HBASE like thrift, Java API.
  • Exported the patterns analyzed back to Teradata usingSqoop.
  • Implemented Region servers to track secure based regions in HBase.
  • Experienced join data from different data sources using Pig Join operations.
  • Experienced with implementing analytics on providing insights using pig sampling.
  • Implemented Range practitioners in using Total Order Practitioner to implement global sorting in Map Reduce.
  • Experienced with handling Sequence file for better optimization, small file problems, file based structures like log files in Map Reduce.
  • Experienced with handling Avro data files using Avro tools and dynamic schema generation tools in Map Reduce programs.
  • Experienced with Optimizing Hive queries by enabling certain features in Hive.
  • Experienced with implementing incremental updates in hive by fallowing four-step strategies.
  • Experienced with implementing Optimized Hive joins like Map-Side joins and Bucketed Map joins.
  • Integrating bulk data into Cassandra file system using MapReduce programs.
  • Performed quick searching, sorting and grouping for analyzing data using Datastax Cassandra.
  • Experienced with moving data from Teradata/ RDBMS systems into HDFS using Sqoop.
  • Integrated Ooziewc client with java applications to build work flows into application.
  • Implemented Unit test cases using MRUnit, Easy Mock and Junit.
  • Involved in application design like Sequence Diagrams, Class Diagrams using Microsoft VISIO tool.

Confidential, Irving, TX

J2EE Developer

Responsibilities:

  • Involved in Designing and analyzing previous code.
  • Involved in design of ICD (Interface control design) sequence diagrams, class diagrams and data flow diagrams.
  • Involved in Bugfixes(Refactoring) of written code.
  • Implemented RESTfulwebservices with JERSEY implementation for providing the webservices to shipping client calls.
  • Worked with Design patterns like Service Factory, Singleton and Factory Pattern.
  • Worked with Servlets transformation layer.
  • Worked with RESTful service for providing the services in JSON.
  • Worked with core java 1.6 in process layer.
  • In transaction layer used calling backend transaction EJB object using the service factory pattern.
  • Used Confidential Logger for logging purpose.
  • Worked with JAXB for marshaling and parsing of data from back end transaction.
  • Involved in writing J-Unit test cases.
  • Total project was built with Ant. Involved in writing the build script and build properties.
  • Used SVN for source code versioning and code repository.

Confidential

J2EE Developer

Responsibilities:

  • Participated in Sprint meetings using AGILE development methodology.
  • Involved in all the phases of the life cycle of the project from requirements gathering to quality assurance testing.
  • Developed Class diagrams, Sequence diagramsusing Rational Rose.
  • Responsible in developing Rich Web Interface modules with Struts tags,JSP, JSTL, CSS,JavaScript, Ajax, GWT.
  • Developed presentation layer using Struts framework, and performed validations using Struts Validator plugin.
  • Created SQL script for the Oracle database
  • Implemented the Business logic using Java Spring Transaction Spring AOP.
  • Implemented persistence layer using Spring JDBC to store and update data in database.
  • Produced web service using WSDL/SOAP standard.
  • Implemented J2EE design patterns like Singleton Pattern with Factory Pattern.
  • Extensively involved in the creation of the Session Beans and MDB, using EJB 3.0.
  • Used Hibernate framework for Persistence layer.
  • Extensively involved in writingStored Proceduresfor data retrieval and data storage and updates inOracledatabase using Hibernate.
  • Deployed and built the application using Maven.
  • Performed testingusingJunit.
  • Used JIRA to track bugs.
  • Extensively used Log4j for logging throughout the application.
  • Produced a Web service using REST with Jersey implementation for providing customer information.
  • Used SVN for source code versioning and code repository.

We'd love your feedback!