Hadoop Developer Resume
Irving, TX
SUMMARY
- 9+ years of experience in Application analysis, Design, Development,
- Maintenance/Supporting web and Client - server based applications in Java/J2EE technologies which includes 3 years of experience with Big Data management andHadooprelated components HDFS, MapReduce, Pig, Hive,YARN,Sqoop,Flume, Crunch, Spark, Strom, Scala,kafkabased on Big Data platforms.
- Delivered 3 projects (full SDLC) using big data technologies like Hadoop, Oozie and NoSQL
- Extensive working noledge in setting up and running Clusters, monitoring, Data analytics, Sentiment analysis,Predictive analysis, Data presentation with big data world.
- Expertise in design and implementation of Banking, E-Commerce and Retail domains.
- Experience managing Cloudera distribution of Hadoop(CDH4&5 using Cloudera Manager),MapR, and Horton works.
- Excellent understanding of NoSQL databases like HBase,Cassandra and MongoDB.
- Hands on experience working on structured, unstructured data with various file formats such as Avrodata files, xml files, Jsonfiles, sequence files using Map Reduce programs.
- Work experience with cloud configurations likeAmazon web services (AWS).
- Implemented custom business logic and performedjoin optimization,secondary sorting, custom sorting using Map Reduce programs.
- Experienced testing and running of Map Reducepipelineson Apache Crunch.
- Composed Map Reduce pipelines with teh help of UDF’s and also implemented joining and data aggregation using Apache Crunch.
- Developed MapReduce jobs, Used different optimization techniques to improve performance in Map Reduce Programs.
- Expertise in Data load management, importing & exporting data using SQOOP,Apache Kafka, Spark Streaming and FLUME
- Implemented business logic using Pig scripts. Wrote custom Pig UDF’s to analyze data
- Performed differentETLoperations using Pigfor joiningoperations and transformations on data to join, clean, aggregate and analyze data.
- Experience with Oozie Workflow Engine to automate and parallelize Hadoop, MapReduceand Pig jobs
- Experienced with handling different file formats like xml, json, RC and ORC formats in Hive using different SerDe's.
- Extensive experience with wiring SQL queries using HiveQL to perform analytics on structured data.
- Experience in performing data validation using HIVE dynamic partitioning and bucketing.
- Experienced in importing and exporting data between RDBMS and TeraData intoHDFS using Sqoop.
- Experienced in handling streaming data like web server log data using flume.
- Worked on design and implemented a Cassandra based database and related webservices for storing unstructured data.
- Experience on YARN environment with Strom, Spark, Kafka and Avro.
- Hands on experience in Scala, Kafka and Strom.
- Experience in implementing algorithms for analyzing using spark.
- Experience in implementingSpark using Scalaand SparkSQL for fasteranalyzing and processing of data.
- Experience in getting data from various sources into HDFS and building reports using Tableau with spark.
- Experience in creating tables on top of data on AWS S3obtained from different data sources and providing them to analytics team building reports using Tableau.
- Extensive Hands on experience with Accessing and perform CURD operations against HBase data using Java APIand implementing time series data management.
- Experienced in monitoringHadoop clusters using Cloudera ManagerTool.
- Involved in various datamining tasks such as pattern mining, classification and clustering techniques.
- Expert noledge over J2EE Design Patterns like MVC Architecture, Singleton, Factory Pattern, Front Controller, Session Facade, Business Delegate and Data Access Object for building J2EE Applications.
- Extensive experience in Java/J2EE programming - Spring, Hibernate, SOAP/Rest web services,JMS, JNDI, EJB,JAX-WS.
- Expertise with Application servers and web servers likeOracle Web Logic,IBM WebSphere, Apache Tomcat, JBOSS and VMware.
- Proven expertise in implementing IOC/Dependency Injection features in various aspects of Spring Framework.
- Expertise Working with XML Parsers likes SAX, DOM, and XStream
- Good Knowledge of using IDE Tools like Eclipse, NetBeans,SpringSource Tool for Java/J2EE application development.
- Experienced in developing teh unit test cases usingBeetest,MRUnit, Junitand Easy Mock.
- KnowledgeonSplunkfor logging mechanism.
- Knowledge on Build tool Jenkins.
- Experience in using Mavenand ANTfor build automation.
- Experience in using version control and configuration management tools like SVN, CVSand Tortoise SVN.
- Experience working in environments using Agile (SCRUM) andWaterfall methodologies.
- Good noledge analyzing data usingPython development and scripting for HadoopStreaming.
- Involved in gathering logs information and managing hardware issues using Netezza.
- Experience in designing applications using UML Diagrams like Class Diagram, Component Diagram, Sequence Diagrams, and Deployment Diagram using MS Visio, Rational Rose.
- Expertise in database modeling, administration and development usingSQLand PL/SQL in Oracle (8i, 9i and 10g), MySQL, Teradata, DB2and SQL Server environments.
- Experienced in using Operating Systems like Windows 98 / 2000 / XP/7/ UNIX,LINUX.
TECHNICAL SKILLS
Hadoop/Big Data: HDFS, Map Reduce, Hive, Pig, YARN,Sqoop, Flume, Oozie, Cloudera manager,Crunch, Strom, Scala, Kafka, Spark, AWS, RHadoop
Database/NoSql: SQL.Pl/SQL,HBase,Cassandra
J2EE Frameworks: Hibernate,Spring, JMS, EJB, JSF
XML/Web Services: SOAP/ Rest
Methodologies: Agile, Waterfall
Build Tools: Maven, ANT, Log4j.¬
Scripting languages: JavaScript, HTML, HTML5, XML, Python, JSP and Servlets
Tools: and other: Rational Rose, Microsoft Visio, CSS
Operating Systems: Linux/Unix,WINDOWS
PROFESSIONAL EXPERIENCE
Confidential
Hadoop Developer
Responsibilities:
- Involved in Analysis, Design, Development and Testing process based on teh new business requirements.
- Experience managing Cloudera distribution of Hadoop CDH5.5using Cloudera Manager.
- Wrote java processes for generating different sample log types which generates different random patterns.
- Wrote script for generating flumeconfg files for different logs types and for different environments.
- Worked on setting up flume agents in dev and uat environments.
- Working on data ingestions, integrating with flume.
- Experience in refactoring teh existing spark batch process for different logs written in scala.
- Writing workflows and scheduling using oozie.
- Deploying oozie jobs in DEV and UAT for testing teh hourly jobs by parsing teh generated sample logs in teh same environments.
- End to end testing teh ingestion pipeline in UAT from generating logs to saving teh transformed data in HDFS.
- Responsible in handling streaming data like web server log data using flume.
- Experience in working with Hive for processing teh raw data.
- Involved in daily SCRUM meetings to discuss teh development/progress ofSprints and was active in making scrum meetings more productive.
Confidential
Data Engineer
Responsibilities:
- Experience working with analysts and product owners for getting exact requirement from clients when they provide some new data sources.
- Extensively used Amazon s3 as cloud storage for all teh ingested and transformed data.
- Migrated all data and tables from Spark 0.9 to Spark 1.4. Created a parallel process for data all ingestions which uses spark 1.4 and then comparing and validating data to 0.9 their by killing processes in 0.9.
- Developed java code to get data from different type’s sources which uses REST API, screen scraping from diverse websites depending on type of source and ingested it into a central repository.
- Responsible for developing Hive queries involving bucketing, partitioning and multiple UDF implementations.
- Involved in developing UDF’s and UDAF’s to implement customized transformations.
- Extensively used Sqoop to get data from RDBMS sources like Teradata and Netezza.
- Validated if data has been normalized and loaded into S3 buckets using Hive queries.
- Wrote custom Hive UDF’s based on business requirement.
- Developed an automated system which can import data from any type of source to any type of destination.
- Developed SparkSQL automation components and responsible for modifying java component to directly connect to thrift server.
- Experience in writing job flows in python which integrates developed java code and shell scripts which run onAzkabanserver.
- Scheduled and monitored jobs using Azkaban.
- Experience in spinning up anAmazon EMR cluster with custom configuration and resource allocation.
- Experience in providing on call support for any cluster issues etc. Also experienced in solving data issues when a job fails in production cluster coz of cluster setup issue, working on it on teh fly when analytics team creates an incident.
- Experience is setting up custom slack notifications where notification can be sent to a group/channel or as a personal message for failure jobs.
- Did a POC on comparing CSV files with parquet files on big data sets by performing aggregations and observing time responses.
- Experience in solving data issues or defects raised by QA team.
- Involved in daily SCRUM meetings to discuss teh development/progress ofSprints and was active in making scrum meetings more productive.
Confidential
Hadoop Developer
Responsibilities:
- Participated in Gathering requirements, analyze requirements and design technical documents for business requirements.
- Involved different phases in big data projects like data acquiring, data processing and data serving using dash boards.
- Implemented different aggregations and filtering operations using JavaAPI in Cassandra
- Responsible for Datamodeling in Cassandra and deciding teh row key and different column families in Cassandra.
- Import/export data from Oracle data base to/from HDFS using Sqoop, Hue and JDBC.
- Gatheird data from different sources like Internet, sensors, user behavior using Flume and Kafka and moved to HDFSand implementedOptimized join baseusingMapReduce programs.
- Implemented Custom Input formats that handlewide range of input files received from java applications to process in MapReduce.
- Implemented joins and data aggregation using Apache Crunch.
- Writing MapReduce pipelineprograms for testing using Apache Crunch.
- Experienced in unit testingMapReduceprograms using MRUnit.
- Divided each data set in to corresponding categories by following MapReduce Binning design pattern.
- Implemented FilterMappers to eliminate un-necessary records and perform data and schema validation.
- Experience in using Pig as an ETL tool for event joins, filters, transformations and pre- aggregations.
- Created partitions, bucketing across state in Hive to handle structured data.
- Implemented Dash boards that handleHiveQL queries internally like Aggregation functions, basic hive operations, and different kind of join operations.
- Implemented business logic based on state in Hive using Generic UDF's.
- Involved in creating data-models for customer data using Cassandra Query Language.
- Managing and scheduling batch Jobs on a Hadoop Cluster using Oozie.
- Created production jobs using Oozie work flows that integrated different actions like MapReduce, Sqoop, and Hiveby utilizing fork and join operations in Oozie.
- Experience in managing and reviewing Hadoop Log files.
- Experienced with monitoring Cluster using Cloudera managerandAmbari.
- Involved in managing Hadoop distributions using Hortonworks.
- Worked on building BI reports inTableauwith Sparkusing SparkSQL.
- Experience in deploying data from various sources into HDFS and building reports using Tableau.
- Implemented Spark using Scalaand SparkSQL for faster testing and processing of data.
- Implemented data ingestion and handling clusters in real time processing using kafka.
- Experience with Core Distributed computing and Data Mining Library using ApacheSpark.
- Experienced in data modelling in hive implementing hive-indexing.
- Experience in utilizing spark machine learning techniques implemented in scala.
- Experienced in configuring maven builds that integrated dependencies check styles, test coverage's.
- Responsible for analyzing multi-platform applications using python.
- Developed MapReduce jobs in Python for data cleaning and data processing.
- Designing Test Plans, Test Cases and performed System Testing.
- Involved in daily SCRUM meetings to discuss teh development/progress of Sprints and was active in making scrum meetings more productive.
- Experience in integrating RHadoop for categorization and statistical analysis to generate reports.
Confidential
Hadoop Developer
Responsibilities:
- Involved in Analysis, Design, Development and Testing processof phase 1.
- Interacting with teh client and designing Technical design document from Business Requirements for teh development team.
- Access HBase from Java using Java API to perform CRUD operations on HBase data to perform real time analytics on hospitals.
- Created hive external tables with data residing in Hbase by implementingHBase-Hive integration.
- Experienced in loading and transforming large sets of structured, semi-structuredand unstructured data Hadoop concepts.
- Experienced with different APIs to access HBASE like thrift, Java API.
- Exported teh patterns analyzed back to Teradata usingSqoop.
- Implemented Region servers to track secure based regions in HBase.
- Experienced join data from different data sources using Pig Join operations.
- Experienced with implementing analytics on providing insights using pig sampling.
- Implemented Range practitioners in using Total Order Practitioner to implement global sorting in Map Reduce.
- Experienced with handling Sequence file for better optimization, small file problems, file based structures like log files in Map Reduce.
- Experienced with handling Avro data files using Avro tools and dynamic schema generation tools in Map Reduce programs.
- Experienced with Optimizing Hive queries by enabling certain features in Hive.
- Experienced with implementing incremental updates in hive by fallowing four-step strategies.
- Experienced with implementing Optimized Hive joins like Map-Side joins and Bucketed Map joins.
- Integrating bulk data into Cassandra file system using MapReduce programs.
- Performed quick searching, sorting and grouping for analyzing data using Datastax Cassandra.
- Experienced with moving data from Teradata/ RDBMS systems into HDFS using Sqoop.
- Integrated Ooziewc client with java applications to build work flows into application.
- Implemented Unit test cases using MRUnit, Easy Mock and Junit.
- Involved in application design like Sequence Diagrams, Class Diagrams using Microsoft VISIO tool.
Confidential, Irving, TX
J2EE Developer
Responsibilities:
- Involved in Designing and analyzing previous code.
- Involved in design of ICD (Interface control design) sequence diagrams, class diagrams and data flow diagrams.
- Involved in Bugfixes(Refactoring) of written code.
- Implemented RESTfulwebservices with JERSEY implementation for providing teh webservices to shipping client calls.
- Worked with Design patterns like Service Factory, Singleton and Factory Pattern.
- Worked with Servlets transformation layer.
- Worked with RESTful service for providing teh services in JSON.
- Worked with core java 1.6 in process layer.
- In transaction layer used calling backend transaction EJB object using teh service factory pattern.
- Used Confidential Logger for logging purpose.
- Worked with JAXB for marshaling and parsing of data from back end transaction.
- Involved in writing J-Unit test cases.
- Total project was built with Ant. Involved in writing teh build script and build properties.
- Used SVN for source code versioning and code repository.
Confidential
J2EE Developer
Responsibilities:
- Participated in Sprint meetings using AGILE development methodology.
- Involved in all teh phases of teh life cycle of teh project from requirements gathering to quality assurance testing.
- Developed Class diagrams, Sequence diagramsusing Rational Rose.
- Responsible in developing Rich Web Interface modules with Struts tags,JSP, JSTL, CSS,JavaScript, Ajax, GWT.
- Developed presentation layer using Struts framework, and performed validations using Struts Validator plugin.
- Created SQL script for teh Oracle database
- Implemented teh Business logic using Java Spring Transaction Spring AOP.
- Implemented persistence layer using Spring JDBC to store and update data in database.
- Produced web service using WSDL/SOAP standard.
- Implemented J2EE design patterns like Singleton Pattern with Factory Pattern.
- Extensively involved in teh creation of teh Session Beans and MDB, using EJB 3.0.
- Used Hibernate framework for Persistence layer.
- Extensively involved in writingStored Proceduresfor data retrieval and data storage and updates inOracledatabase using Hibernate.
- Deployed and built teh application using Maven.
- Performed testingusingJunit.
- Used JIRA to track bugs.
- Extensively used Log4j for logging throughout teh application.
- Produced a Web service using REST with Jersey implementation for providing customer information.
- Used SVN for source code versioning and code repository.
