Hadoop Developer Resume
PhoeniX
PROFESSIONAL SUMMARY:
- Overall 8+ years of overall experience with strong emphasis on Design, Development, Implementation, Testing and Deployment of Software Applications.
- Over 5+ years of comprehensive IT experience in BigData and Big DataAnalytics, Hadoop, HDFS, Map Reduce, YARN, Hadoop Ecosystem and ShellScripting.
- Highly capable for processing large sets of Structured, Semi - structured and Unstructured datasets and supporting BigData applications.
- Hands on experience with Hadoop Ecosystem components like MapReduce (Processing), HDFS (Storage), YARN, Sqoop, Pig, Hive, HBase, Oozie, ZooKeeper and Spark for data storage and analysis.
- Expertise in transferring data between a Hadoop ecosystem and structured data storage in a RDBMS such as MY SQL, Oracle, Teradata and DB2 using Sqoop.
- Experience in NoSQL databases like MongoDB, HBase and Cassandra.
- Have excellent knowledge on Python Collections and Multi-Threading.
- Skilled experience in Python with proven expertise in using new tools and technical developments
- Experience in ApacheSpark cluster and streams processing using Spark Streaming
- Worked on several python packages like numpy, scipy, pytables etc.
- Expertise in moving large amounts of log, streaming event data and Transactional data using Flume.
- Experience in developing Map Reduce jobs in Java for data cleaning and pre-processing.
- Expertise in writing PigLatin, Hive Scripts and extended their functionality using UserDefined Functions (UDF's).
- Expertise in handling structured arrangement of data within certain limits (Data Layout's) using Partitions and Bucketing in Hive.
- Expertise in preparing interactive Data Visualization's using Tableau Software from different sources.
- Hands on experience in developing workflows that execute MapReduce, Sqoop, Pig,Hive and Shellscripts using Oozie.
- Experience working with Cloudera HueInterface and Impala.
- Hands on experience developing Solr Indexes using MapReduceIndexer Tool.
- Expertise in Object-oriented analysis and design (OOAD) like UML and use of various design patterns.
- Experience in Java, JSP, Servlets, EJB, Web Logic, Web Sphere, Hibernate, SpringJBoss, JDBC, RMI, Java Script, Ajax, Jquery, XML and HTML.
- Fluent with the core Java concepts like I/O, Multi-threading, Exceptions, RegEx, Data Structures and Serialization.
- Performed unit testing using Junit Testing Framework and Log4J to monitor the error logs.
- Good Knowledge of Python and Python WebFramework Django.
- Experienced with Python frameworks like Webapp2 and, Flask.
- Experience in process improvement, normalization/de-normalization, data extraction, cleansing and manipulation.
- Extensively used Informatica Power Centre for Extraction, Transformation and Loading process.
- Experience in Dimensional Data Modelling using Star and Snow Flake Schema.
- Worked on reusable code known as Tie outs to maintain the data consistency.
- Converting requirement specification, Source system understanding into Conceptual, Logical and physical Data Model, Data flow (DFD).
- Hands on JAXWS, JSP, Servlets, Struts, WebLogic, WebSphere, Hibernate, spring, JBoss,JDBC, RMI, Java Script, Ajax, JQuery, Linux, UNIX, WSDL, XML, HTML, AWS and Scala and Vertica.
- Expertise in working with transactional databases like Oracle, SQLserver, MySQL, and Db2.
- Expertise in developing SQLqueries, Stored Procedures and excellent development experience with Agile Methodology.
- Ability to adapt to evolving technology, strong sense of responsibility and accomplishment.
- Excellent leadership, interpersonal, problem solving and time management skills.
- Excellent communication skills both Written (documentation) and Verbal (presentation).
TECHNICAL SKILLS:
Languages/Tools: Java, C, C++, C#,Scala, VB, XML, HTML/XHTML, HDML, DHTML.
Big Data: HDFS, MapReduce, HIVE, PIG, HBase, SQOOP, Oozie, Zookeeper, Spark, Mahout, Kafka, Storm, Cassandra, Solr, Impala,Greenplum, MongoDB
Web/Distributed Technologies: J2EE, Servlets, JSP, Struts, Hibernate, JSF, JSTL,EJB,RMI,JNI, XML,JAXP,XSL,XSLT, UML, MVC,STRUTS, Spring, Corba, Java Threads.
Browser Languages/Scripting: HTML, XHTML, CSS, XML, XSL, XSD, XSLT, Java script, HTML DOM, DHTML, AJAX.
App/Web Servers: IBM Websphere BEA Web logic, Jdeveloper, Apache Tomcat, JBoss.
GUI Environment: Swing, AWT, Applets.
Messaging & Web Services Technology: SOAP, WSDL,UDDI, XML, SOA, JAX-RPC, IBM WebSphere MQ v5.3, JMS.
Testing &Case Tools: JUnit, Log4j, Rational Clear case, CVS, ANT, Maven, JBuilder.
Configuration Management: Chef, Puppet, Ansible, Docker.
Build Tools: CVS, Subversion, GIT, Ant, Maven, Gradle, Hudson, TeamCity, Jenkins, Chef, Puppet, Ansible, Docker.
CI Tools: Jenkins, Bamboo.
Scripting Languages: Python, Shell (Bash), Perl, PowerShell, Ruby, Groovy, PowerShell.
Monitoring Tools: Nagios, Cloud Watch, JIIRA, Bugzilla and Remedy.
Databases: NO SQL Oracle, MS SQL Server 2000, DB2, MS Access & My SQL, Teradata. Cassandra, Greenplum and MongoDB
Operating systems: Windows, Solaris, Unix, Linux (Red Hat 'SUSELinux), Sun Solaris, Ubuntu, CentOS.
PROFESSIONAL EXPERIENCE:
Confidential, Phoenix
Hadoop Developer
Responsibilities:
- Involved in end to end data processing like ingestion, processing, and quality checks and splitting.
- Real time streaming the data using Spark Streaming with Kafka
- Developed Spark scripts by using Scala as per the requirement.
- Load the data into SparkRDD and performed in-memory data computation to generate the output response.
- Performed different types of transformations and actions on the RDD to meet the business requirements.
- Developed a data pipeline using Kafka, Spark and Hive to ingest, transform and analysing data.
- Also worked on analysing Hadoop cluster and different bigdata analytic tools including Pig, HBase and Sqoop.
- Involved in loading data from UNIX file system to HDFS.
- Created HBase tables to store variable data formats of PII data coming from different portfolios.
- Implemented best offer logic using Pig scripts and Pig UDFs.
- Responsible to manage data coming from various sources.
- Installed and configured Hive and also written Hive UDFs.
- Experience on loading and transforming of large sets of structured, semi structured and unstructured data.
- Cluster coordination services through Zookeeper.
- Exported the analysed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Analysed large amounts of data sets to determine optimal way to aggregate and report on it.
- Responsible for setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
- Installed and configured Hadoop Map Reduce, HDFS.
- Developed multiple Map Reduce jobs in java for data cleaning and pre-processing.
- Installed and configured Pig.
- Involved in managing and reviewing Hadoop log files.
- Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
- Developing Scripts and Batch Job to schedule various Hadoop Program.
- Responsible for writing Hive queries for data analysis to meet the business requirements.
- Responsible for creating Hive tables and working on them using HiveQL.
- Responsible for importing and exporting data into HDFS and Hive using Sqoop.
- Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
- Designed and implemented Map Reduce based large-scale parallel relation-learning system.
- Involved in scheduling Oozie workflow engine to run multiple Hive jobs
Environment: Hadoop, MapReduce2.7.2, Hive2.0, Pig0.16, Sqoop2, Java, Oozie, HBase0.98.19, Kafka0.10.1.1, Spark2.0, Scala2.12.0, Eclipse, Linux, Oracle, Teradata.
Confidential, San Francisco, CA
Hadoop Developer
Responsibilities:
- Developed a process for Sqooping data from multiple sources like SQLServer, Oracle and Teradata.
- Responsible for creation of mapping document from source fields to destination fields mapping.
- Developed a shell script to create staging, landing tables with the same schema like the source and generate the properties which are used by Ooziejobs.
- Developed Oozie workflow's for executing Sqoop and Hive actions.
- Worked with NoSQL databases like Hbase in creating Hbase tables to load large sets of semi structured data coming from various sources.
- Involved in building databaseModel, APIs and Views utilizing python, in order to build an interactive web based solution
- Performance optimizations on Spark/Scala. Diagnose and resolve performance issues.
- Responsible for developing Python wrapper scripts which will extract specific date range using Sqoop by passing custom properties required for the workflow.
- Developed scripts to run Oozie workflows, capture the logs of all jobs that run on cluster and create a metadata table which specifies the execution times of each job.
- Developed Hivescripts for performing transformation logic and also loading the data from staging zone to final landing zone.
- Developed monitoring and notification tools using Python.
- Worked on Parquet File format to get a better storage and performance for publish tables.
- Involved in loading transactional data into HDFS using Flume for Fraud Analytics.
- Developed Python utility to validate HDFS tables with source tables.
- Designed and developed UDF'S to extend the functionality in both PIG and HIVE.
- Import and Export of data using Sqoop between MySQL to HDFS on regular basis.
- Managed datasets using Panda data frames and MySQL, queried MYSQL database queries from python using Python-MySQL connector and MySQL dB package to retrieve information.
- Developed and tested many features for dashboard using Python, Java, Bootstrap, CSS, JavaScript and JQuery.
- Responsible to check-in the developed code into Harvest for release management which is a part of CI/CD.
- Involved in using CA7 tool to setup dependencies at each level (Table Data, File and Time).
- Automated all the jobs for pulling data from FTP server to load data into Hive tables using Oozieworkflows.
- Involved in developing Spark code using Scala and Spark-SQL for faster testing and processing of data and exploring of optimizing it using SparkContext, Spark-SQL, PairRDD's, Spark YARN.
- Migrating the needed data from Oracle, MySQL in to HDFS using Sqoop and importing various formats of flat files in to HDFS.
Environment: Hadoop, HDFS2.6.3, Hive1.0.1, HBase0.98.12.1, Zookeeper3.5.1, Oozie, Impala1.4.1, Java(jdk1.6), ClouderaCDH 3, Oracle, Teradata SQL Server, UNIX Shell Scripting, Flume1.6.0, Scala2.11.6, Spark1.5.0, Sqoop1.4.6, Python3.5.1.
Confidential -Austin, TX
Hadoop Developer
Responsibilities:
- Responsible for understanding the requirements and implementing the security using AD Groups for the Dataset.
- Involved Low level design for MR, Hive, Impala, Shellscripts to process data.
- Worked on ETL scripts to pull the data from DB2/Oracle Data Base into HDFS.
- Experience in utilizing Spark machine learning techniques implemented in Scala.
- Involved in POC development and unit testing using Spark and Scala.
- Created Partitioned Hive tables and worked on them using Hive.
- Installing and configuring Hive, Sqoop, Flume, Oozie on the Hadoop clusters.
- Involved in scheduling Oozie workflow engine to run multiple Hive and Pig jobs.
- Develop and implement Python/Django applications.
- Developed a process for the Batch ingestion of CSV Files, Sqoop from different sources and also generating views on the data source using ShellScripting and Python.
- Integrated a shellscript to create Collections/morphline, SolrIndexes on top of table directories using MapReduce Indexer Tool within Batch Ingestion Framework.
- Implemented partitioning, dynamic partitions and buckets in HIVE.
- Developed HiveScripts to create the views and apply transformation logic in the Target Database.
- Involved in the design of Data Mart and Data Lake to provide faster insight into the Data.
- Involved in using Stream Sets Data Collector tool and created Data Flows for one of the streaming application.
- Experienced in using Kafka as a data pipeline between JMS (Producer) and Spark Streaming Application (Consumer)
- Involved in the development of Spark Streaming application for one of the data source using Scala, Spark by applying the transformations.
- Skilled in using collections in Python for manipulating and looping through different user defined objects.
- Wrote a Python module to connect and view the status of an ApacheCassandra instance.
- Developed a script in Scala to read all the Parquet Tables in a Database and parse them as Json files, another script to parse them as structured tables in Hive.
- Designed and Maintained Oozie workflows to manage the flow of jobs in the cluster.
- Configured Zookeeper for Cluster co-ordination services.
- Generated PythonDjango forms to record data of online users and used PyTest for writing test cases
- Developed a unit test script to read a Parquet file for testing PySpark on the cluster.
- Involved in exploration of new technologies like AWS, Apache Flink, and Apache NIFIetc which can increase the business value.
Environment: Hadoop, HDFS, Hive, HBase, Zookeeper, Impala, Cloudera, Oracle, SQL Server, UNIX Shell Scripting, Flume, Scala, Spark, Sqoop, Python, kafka, PySpark.
Confidential, Houston, TX
Hadoop Developer
Responsibilities:
- Involved in review of functional and non-functional requirements.
- Facilitated knowledge transfer sessions.
- Installed and configured Hadoop Mapreduce, HDFS, Developed multiple MapReduce jobs in java for data cleaning and pre-processing.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Experienced in defining job flows.
- Experienced in managing and reviewing Hadoop log files.
- Extracted files from RDBMS through Sqoop and placed in HDFS and processed.
- Experienced in running Hadoop streaming jobs to process terabytes of xml format data.
- Load and transform large sets of structured, semi structured and unstructured data.
- Responsible to manage data coming from various sources.
- Got good experience with NOSQL database such as HBase
- Supported Map Reduce Programs those are running on the cluster.
- Installed and configured Hive and also written HiveUDFs.
- Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
- Gained very good business knowledge on health insurance, claim processing, fraud suspect identification, appeals process etc.
- Developed a custom File System plug in for Hadoop so it can access files on Data Platform.
- This plugin allows Hadoop MapReduce programs, HBase, Pig and Hive to work unmodified and access files directly.
- Designed and implemented Mapreduce-based large-scale parallel relation-learning system
- Written the programs in Spark using Scala and used RDD for transformations and performed actions on them.
Environment: Java 6, Eclipse, Oracle 10g, Linux Red Hat. Linux, MapReduce, HDFS, Hive, Java (JDK 1.6), MapReduce, Spark, Oracle 11g / 10g, PL/SQL, SQL*PLUS, Toad 9.6, Windows NT, UNIX Shell Scripting.
Confidential
Java Developer
Responsibilities:
- Responsible for understanding the scope of the project and requirements gathering
- Created the database, user, environment, activity and class diagram for the project (UML).
- Implemented the database using oracle database engine.
- Created an entity object (business rules and policy, validation logic, default value logic, security).
- Web application development using J2EE, JSP, Servlets, JDBC, JavaBeans, Struts, Ajax, Custom Tags, EJB, Hibernate, Ant, Junitand ApacheLog4j, Web Services, Message queue(MQ).
- Designing GUI prototype using ADF 11G GUI component before finalizing it for development.
- Experience in using version controls such as CVS, PVCS.
- Involved in consuming, producing Restful web services using JAX-RS.
- Collaborated with ETL/Informatica team to determine the necessary data modules and UI designs to support Cognos reports.
- Junit was used for unit testing for the integration testing tool.
- Created modules using task flow with bounded and unbounded.
- Generating WSDL (web services) and create work flow using BPEL.
- Created the skin for the layout.
- Made integrated testing for the application.
- Created dynamic report and using JFreechart.
Environment: Java, Servlets, JSF, Adf rich client UI framework ADF-BC (BC4J) 11g, Web Services using Oracle SOA, Oracle Web Logic.
Confidential
Software Programmer
Responsibilities:
- Installed, configured and administration of WebSphere Application Server 6.1 Network Deployment on Windows Server.
- Involved in the analysis & design of the application using Rational Rose.
- Developed the various action classes to handle the requests and responses.
- Designed and created JavaObjects, JSP pages, JSF, JavaBeans and Servlets to achieve various business functionalities.
- Created validation methods using JavaScript and Backing Beans.
- Involved in writing client side validations using JavaScript, CSS.
- Involved in the design of the Referential Data Service module to interface with various databases using JDBC.
- Used Hibernate framework to persist the employee work hours to the database.
- Developed classes and interface with underlying web services layer.
- Prepared documentation and participated in preparing user's manual for the application.
- Prepared Use Cases, Business Process Models and Data flow diagrams, User Interface models.
- Gathered & analyzed requirements for EAuto, designed process flow diagrams.
- Defined business processes related to the project and provided technical direction to development workgroup.
- Analyzed the legacy and the Financial Data Warehouse.
- Participated in Data base design sessions, Database normalization meetings.
- Managed Change Request Management and Defect Management.
- Managed UAT testing and developed test strategies, test plans, reviewed QA test plans for appropriate test coverage.
- Involved in Developing JSP's, action classes, form beans, response beans, EJB's.
- Extensively used XML to code configuration files.
- Developed PL/SQL stored procedures, triggers.
- Performed functional, integration, system and validation testing.
Environment: Java, J2EE,JSP, JCL, DB2, Struts, SQL, PL/DSQL, Eclipse, Oracle, Windows XP, HTML, CSS, JavaScript, and XML.
