We provide IT Staff Augmentation Services!

Hadoop Developer Resume

5.00/5 (Submit Your Rating)

PhoeniX

PROFESSIONAL SUMMARY:

  • Overall 8+ years of overall experience with strong emphasis on Design, Development, Implementation, Testing and Deployment of Software Applications.
  • Over 5+ years of comprehensive IT experience in BigData and Big DataAnalytics, Hadoop, HDFS, Map Reduce, YARN, Hadoop Ecosystem and ShellScripting.
  • Highly capable for processing large sets of Structured, Semi - structured and Unstructured datasets and supporting BigData applications.
  • Hands on experience with Hadoop Ecosystem components like MapReduce (Processing), HDFS (Storage), YARN, Sqoop, Pig, Hive, HBase, Oozie, ZooKeeper and Spark for data storage and analysis.
  • Expertise in transferring data between a Hadoop ecosystem and structured data storage in a RDBMS such as MY SQL, Oracle, Teradata and DB2 using Sqoop.
  • Experience in NoSQL databases like MongoDB, HBase and Cassandra.
  • Have excellent knowledge on Python Collections and Multi-Threading.
  • Skilled experience in Python with proven expertise in using new tools and technical developments
  • Experience in ApacheSpark cluster and streams processing using Spark Streaming
  • Worked on several python packages like numpy, scipy, pytables etc.
  • Expertise in moving large amounts of log, streaming event data and Transactional data using Flume.
  • Experience in developing Map Reduce jobs in Java for data cleaning and pre-processing.
  • Expertise in writing PigLatin, Hive Scripts and extended their functionality using UserDefined Functions (UDF's).
  • Expertise in handling structured arrangement of data within certain limits (Data Layout's) using Partitions and Bucketing in Hive.
  • Expertise in preparing interactive Data Visualization's using Tableau Software from different sources.
  • Hands on experience in developing workflows that execute MapReduce, Sqoop, Pig,Hive and Shellscripts using Oozie.
  • Experience working with Cloudera HueInterface and Impala.
  • Hands on experience developing Solr Indexes using MapReduceIndexer Tool.
  • Expertise in Object-oriented analysis and design (OOAD) like UML and use of various design patterns.
  • Experience in Java, JSP, Servlets, EJB, Web Logic, Web Sphere, Hibernate, SpringJBoss, JDBC, RMI, Java Script, Ajax, Jquery, XML and HTML.
  • Fluent with the core Java concepts like I/O, Multi-threading, Exceptions, RegEx, Data Structures and Serialization.
  • Performed unit testing using Junit Testing Framework and Log4J to monitor the error logs.
  • Good Knowledge of Python and Python WebFramework Django.
  • Experienced with Python frameworks like Webapp2 and, Flask.
  • Experience in process improvement, normalization/de-normalization, data extraction, cleansing and manipulation.
  • Extensively used Informatica Power Centre for Extraction, Transformation and Loading process.
  • Experience in Dimensional Data Modelling using Star and Snow Flake Schema.
  • Worked on reusable code known as Tie outs to maintain the data consistency.
  • Converting requirement specification, Source system understanding into Conceptual, Logical and physical Data Model, Data flow (DFD).
  • Hands on JAXWS, JSP, Servlets, Struts, WebLogic, WebSphere, Hibernate, spring, JBoss,JDBC, RMI, Java Script, Ajax, JQuery, Linux, UNIX, WSDL, XML, HTML, AWS and Scala and Vertica.
  • Expertise in working with transactional databases like Oracle, SQLserver, MySQL, and Db2.
  • Expertise in developing SQLqueries, Stored Procedures and excellent development experience with Agile Methodology.
  • Ability to adapt to evolving technology, strong sense of responsibility and accomplishment.
  • Excellent leadership, interpersonal, problem solving and time management skills.
  • Excellent communication skills both Written (documentation) and Verbal (presentation).

TECHNICAL SKILLS:

Languages/Tools: Java, C, C++, C#,Scala, VB, XML, HTML/XHTML, HDML, DHTML.

Big Data: HDFS, MapReduce, HIVE, PIG, HBase, SQOOP, Oozie, Zookeeper, Spark, Mahout, Kafka, Storm, Cassandra, Solr, Impala,Greenplum, MongoDB

Web/Distributed Technologies: J2EE, Servlets, JSP, Struts, Hibernate, JSF, JSTL,EJB,RMI,JNI, XML,JAXP,XSL,XSLT, UML, MVC,STRUTS, Spring, Corba, Java Threads.

Browser Languages/Scripting: HTML, XHTML, CSS, XML, XSL, XSD, XSLT, Java script, HTML DOM, DHTML, AJAX.

App/Web Servers: IBM Websphere BEA Web logic, Jdeveloper, Apache Tomcat, JBoss.

GUI Environment: Swing, AWT, Applets.

Messaging & Web Services Technology: SOAP, WSDL,UDDI, XML, SOA, JAX-RPC, IBM WebSphere MQ v5.3, JMS.

Testing &Case Tools: JUnit, Log4j, Rational Clear case, CVS, ANT, Maven, JBuilder.

Configuration Management: Chef, Puppet, Ansible, Docker.

Build Tools: CVS, Subversion, GIT, Ant, Maven, Gradle, Hudson, TeamCity, Jenkins, Chef, Puppet, Ansible, Docker.

CI Tools: Jenkins, Bamboo.

Scripting Languages: Python, Shell (Bash), Perl, PowerShell, Ruby, Groovy, PowerShell.

Monitoring Tools: Nagios, Cloud Watch, JIIRA, Bugzilla and Remedy.

Databases: NO SQL Oracle, MS SQL Server 2000, DB2, MS Access & My SQL, Teradata. Cassandra, Greenplum and MongoDB

Operating systems: Windows, Solaris, Unix, Linux (Red Hat 'SUSELinux), Sun Solaris, Ubuntu, CentOS.

PROFESSIONAL EXPERIENCE:

Confidential, Phoenix

Hadoop Developer

Responsibilities:

  • Involved in end to end data processing like ingestion, processing, and quality checks and splitting.
  • Real time streaming the data using Spark Streaming with Kafka
  • Developed Spark scripts by using Scala as per the requirement.
  • Load the data into SparkRDD and performed in-memory data computation to generate the output response.
  • Performed different types of transformations and actions on the RDD to meet the business requirements.
  • Developed a data pipeline using Kafka, Spark and Hive to ingest, transform and analysing data.
  • Also worked on analysing Hadoop cluster and different bigdata analytic tools including Pig, HBase and Sqoop.
  • Involved in loading data from UNIX file system to HDFS.
  • Created HBase tables to store variable data formats of PII data coming from different portfolios.
  • Implemented best offer logic using Pig scripts and Pig UDFs.
  • Responsible to manage data coming from various sources.
  • Installed and configured Hive and also written Hive UDFs.
  • Experience on loading and transforming of large sets of structured, semi structured and unstructured data.
  • Cluster coordination services through Zookeeper.
  • Exported the analysed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Analysed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Responsible for setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
  • Installed and configured Hadoop Map Reduce, HDFS.
  • Developed multiple Map Reduce jobs in java for data cleaning and pre-processing.
  • Installed and configured Pig.
  • Involved in managing and reviewing Hadoop log files.
  • Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
  • Developing Scripts and Batch Job to schedule various Hadoop Program.
  • Responsible for writing Hive queries for data analysis to meet the business requirements.
  • Responsible for creating Hive tables and working on them using HiveQL.
  • Responsible for importing and exporting data into HDFS and Hive using Sqoop.
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
  • Designed and implemented Map Reduce based large-scale parallel relation-learning system.
  • Involved in scheduling Oozie workflow engine to run multiple Hive jobs

Environment: Hadoop, MapReduce2.7.2, Hive2.0, Pig0.16, Sqoop2, Java, Oozie, HBase0.98.19, Kafka0.10.1.1, Spark2.0, Scala2.12.0, Eclipse, Linux, Oracle, Teradata.

Confidential, San Francisco, CA

Hadoop Developer

Responsibilities:

  • Developed a process for Sqooping data from multiple sources like SQLServer, Oracle and Teradata.
  • Responsible for creation of mapping document from source fields to destination fields mapping.
  • Developed a shell script to create staging, landing tables with the same schema like the source and generate the properties which are used by Ooziejobs.
  • Developed Oozie workflow's for executing Sqoop and Hive actions.
  • Worked with NoSQL databases like Hbase in creating Hbase tables to load large sets of semi structured data coming from various sources.
  • Involved in building databaseModel, APIs and Views utilizing python, in order to build an interactive web based solution
  • Performance optimizations on Spark/Scala. Diagnose and resolve performance issues.
  • Responsible for developing Python wrapper scripts which will extract specific date range using Sqoop by passing custom properties required for the workflow.
  • Developed scripts to run Oozie workflows, capture the logs of all jobs that run on cluster and create a metadata table which specifies the execution times of each job.
  • Developed Hivescripts for performing transformation logic and also loading the data from staging zone to final landing zone.
  • Developed monitoring and notification tools using Python.
  • Worked on Parquet File format to get a better storage and performance for publish tables.
  • Involved in loading transactional data into HDFS using Flume for Fraud Analytics.
  • Developed Python utility to validate HDFS tables with source tables.
  • Designed and developed UDF'S to extend the functionality in both PIG and HIVE.
  • Import and Export of data using Sqoop between MySQL to HDFS on regular basis.
  • Managed datasets using Panda data frames and MySQL, queried MYSQL database queries from python using Python-MySQL connector and MySQL dB package to retrieve information.
  • Developed and tested many features for dashboard using Python, Java, Bootstrap, CSS, JavaScript and JQuery.
  • Responsible to check-in the developed code into Harvest for release management which is a part of CI/CD.
  • Involved in using CA7 tool to setup dependencies at each level (Table Data, File and Time).
  • Automated all the jobs for pulling data from FTP server to load data into Hive tables using Oozieworkflows.
  • Involved in developing Spark code using Scala and Spark-SQL for faster testing and processing of data and exploring of optimizing it using SparkContext, Spark-SQL, PairRDD's, Spark YARN.
  • Migrating the needed data from Oracle, MySQL in to HDFS using Sqoop and importing various formats of flat files in to HDFS.

Environment: Hadoop, HDFS2.6.3, Hive1.0.1, HBase0.98.12.1, Zookeeper3.5.1, Oozie, Impala1.4.1, Java(jdk1.6), ClouderaCDH 3, Oracle, Teradata SQL Server, UNIX Shell Scripting, Flume1.6.0, Scala2.11.6, Spark1.5.0, Sqoop1.4.6, Python3.5.1.

Confidential -Austin, TX

Hadoop Developer

Responsibilities:

  • Responsible for understanding the requirements and implementing the security using AD Groups for the Dataset.
  • Involved Low level design for MR, Hive, Impala, Shellscripts to process data.
  • Worked on ETL scripts to pull the data from DB2/Oracle Data Base into HDFS.
  • Experience in utilizing Spark machine learning techniques implemented in Scala.
  • Involved in POC development and unit testing using Spark and Scala.
  • Created Partitioned Hive tables and worked on them using Hive.
  • Installing and configuring Hive, Sqoop, Flume, Oozie on the Hadoop clusters.
  • Involved in scheduling Oozie workflow engine to run multiple Hive and Pig jobs.
  • Develop and implement Python/Django applications.
  • Developed a process for the Batch ingestion of CSV Files, Sqoop from different sources and also generating views on the data source using ShellScripting and Python.
  • Integrated a shellscript to create Collections/morphline, SolrIndexes on top of table directories using MapReduce Indexer Tool within Batch Ingestion Framework.
  • Implemented partitioning, dynamic partitions and buckets in HIVE.
  • Developed HiveScripts to create the views and apply transformation logic in the Target Database.
  • Involved in the design of Data Mart and Data Lake to provide faster insight into the Data.
  • Involved in using Stream Sets Data Collector tool and created Data Flows for one of the streaming application.
  • Experienced in using Kafka as a data pipeline between JMS (Producer) and Spark Streaming Application (Consumer)
  • Involved in the development of Spark Streaming application for one of the data source using Scala, Spark by applying the transformations.
  • Skilled in using collections in Python for manipulating and looping through different user defined objects.
  • Wrote a Python module to connect and view the status of an ApacheCassandra instance.
  • Developed a script in Scala to read all the Parquet Tables in a Database and parse them as Json files, another script to parse them as structured tables in Hive.
  • Designed and Maintained Oozie workflows to manage the flow of jobs in the cluster.
  • Configured Zookeeper for Cluster co-ordination services.
  • Generated PythonDjango forms to record data of online users and used PyTest for writing test cases
  • Developed a unit test script to read a Parquet file for testing PySpark on the cluster.
  • Involved in exploration of new technologies like AWS, Apache Flink, and Apache NIFIetc which can increase the business value.

Environment: Hadoop, HDFS, Hive, HBase, Zookeeper, Impala, Cloudera, Oracle, SQL Server, UNIX Shell Scripting, Flume, Scala, Spark, Sqoop, Python, kafka, PySpark.

Confidential, Houston, TX

Hadoop Developer

Responsibilities:

  • Involved in review of functional and non-functional requirements.
  • Facilitated knowledge transfer sessions.
  • Installed and configured Hadoop Mapreduce, HDFS, Developed multiple MapReduce jobs in java for data cleaning and pre-processing.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Experienced in defining job flows.
  • Experienced in managing and reviewing Hadoop log files.
  • Extracted files from RDBMS through Sqoop and placed in HDFS and processed.
  • Experienced in running Hadoop streaming jobs to process terabytes of xml format data.
  • Load and transform large sets of structured, semi structured and unstructured data.
  • Responsible to manage data coming from various sources.
  • Got good experience with NOSQL database such as HBase
  • Supported Map Reduce Programs those are running on the cluster.
  • Installed and configured Hive and also written HiveUDFs.
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
  • Gained very good business knowledge on health insurance, claim processing, fraud suspect identification, appeals process etc.
  • Developed a custom File System plug in for Hadoop so it can access files on Data Platform.
  • This plugin allows Hadoop MapReduce programs, HBase, Pig and Hive to work unmodified and access files directly.
  • Designed and implemented Mapreduce-based large-scale parallel relation-learning system
  • Written the programs in Spark using Scala and used RDD for transformations and performed actions on them.

Environment: Java 6, Eclipse, Oracle 10g, Linux Red Hat. Linux, MapReduce, HDFS, Hive, Java (JDK 1.6), MapReduce, Spark, Oracle 11g / 10g, PL/SQL, SQL*PLUS, Toad 9.6, Windows NT, UNIX Shell Scripting.

Confidential

Java Developer

Responsibilities:

  • Responsible for understanding the scope of the project and requirements gathering
  • Created the database, user, environment, activity and class diagram for the project (UML).
  • Implemented the database using oracle database engine.
  • Created an entity object (business rules and policy, validation logic, default value logic, security).
  • Web application development using J2EE, JSP, Servlets, JDBC, JavaBeans, Struts, Ajax, Custom Tags, EJB, Hibernate, Ant, Junitand ApacheLog4j, Web Services, Message queue(MQ).
  • Designing GUI prototype using ADF 11G GUI component before finalizing it for development.
  • Experience in using version controls such as CVS, PVCS.
  • Involved in consuming, producing Restful web services using JAX-RS.
  • Collaborated with ETL/Informatica team to determine the necessary data modules and UI designs to support Cognos reports.
  • Junit was used for unit testing for the integration testing tool.
  • Created modules using task flow with bounded and unbounded.
  • Generating WSDL (web services) and create work flow using BPEL.
  • Created the skin for the layout.
  • Made integrated testing for the application.
  • Created dynamic report and using JFreechart.

Environment: Java, Servlets, JSF, Adf rich client UI framework ADF-BC (BC4J) 11g, Web Services using Oracle SOA, Oracle Web Logic.

Confidential

Software Programmer

Responsibilities:

  • Installed, configured and administration of WebSphere Application Server 6.1 Network Deployment on Windows Server.
  • Involved in the analysis & design of the application using Rational Rose.
  • Developed the various action classes to handle the requests and responses.
  • Designed and created JavaObjects, JSP pages, JSF, JavaBeans and Servlets to achieve various business functionalities.
  • Created validation methods using JavaScript and Backing Beans.
  • Involved in writing client side validations using JavaScript, CSS.
  • Involved in the design of the Referential Data Service module to interface with various databases using JDBC.
  • Used Hibernate framework to persist the employee work hours to the database.
  • Developed classes and interface with underlying web services layer.
  • Prepared documentation and participated in preparing user's manual for the application.
  • Prepared Use Cases, Business Process Models and Data flow diagrams, User Interface models.
  • Gathered & analyzed requirements for EAuto, designed process flow diagrams.
  • Defined business processes related to the project and provided technical direction to development workgroup.
  • Analyzed the legacy and the Financial Data Warehouse.
  • Participated in Data base design sessions, Database normalization meetings.
  • Managed Change Request Management and Defect Management.
  • Managed UAT testing and developed test strategies, test plans, reviewed QA test plans for appropriate test coverage.
  • Involved in Developing JSP's, action classes, form beans, response beans, EJB's.
  • Extensively used XML to code configuration files.
  • Developed PL/SQL stored procedures, triggers.
  • Performed functional, integration, system and validation testing.

Environment: Java, J2EE,JSP, JCL, DB2, Struts, SQL, PL/DSQL, Eclipse, Oracle, Windows XP, HTML, CSS, JavaScript, and XML.

We'd love your feedback!