Sr. Hadoop Developer Resume
Carlstadt, NJ
PROFESSIONAL SUMMARY:
- Over 8 years of experience with emphasis on Big Data technologies, development and design of Java based enterprise applications
- Expertise in the creation of On - premise and Cloud Data Lake
- Experience working with Cloudera, Hortonworks, AWS and Pivotal Distributions of Hadoop
- Expertise in HDFS, MapReduce, Spark, Hive, Impala, Pig, Sqoop, HBase, Oozie, Flume, Kafka, Storm, Solr and various other ecosystem components.
- Expertise in Spark framework for batch and real time data processing
- Experience in working with BI team and transform big data requirements into Hadoop centric technologies.
- Experience in performance tuning the Hadoop cluster by gathering and analyzing the existing infrastructure.
- Working experience on designing and implementing complete end-to-end Hadoop Infrastructure including PIG, HIVE, Sqoop, Oozie, Flume and zookeeper.
- Experience in converting MapReduce applications to Spark.
- Experience in handline messaging services using Apache Kafka.
- Experience in working with flume to load the log data from multiple sources directly into HDFS
- Experience in Data migration from existing data stores and mainframe NDM(Network Data mover) to Hadoop
- Good Knowledge with NonSQL Databases - Cassandra, Mongo DB and HBase.
- Experience in handling multiple relational databases: MySQL, SQL Server, PostgreSQL and Oracle.
- Experience in supporting data analysis projects using Elastic Map Reduce on the Amazon Web Services (AWS) cloud. Exporting and importing data into S3.
- Experience in designing both time driven and data driven automated workflows using Oozie.
- Experience in supporting analysts by administering and configuring HIVE.
- Experience in running Pig and Hive scripts.
- Experience in fine-tuning Mapreduce jobs for better scalability and performance.
- Developed various Map Reduce applications to perform ETL workloads on terabytes of data.
- Performed Importing and exporting data into HDFS and Hive using Sqoop.
- Experience in writing shell scripts to dump the sharded data from Landing Zones to HDFS.
- Worked on predictive modeling techniques like Neural Networks, Decision Trees and Regression Analysis.
- Experience in Data mining and Business Intelligence tools such as Tableau, SAS Enterprise Miner, JMP and Enterprise Guide, IBM SPSS modeler and MicroStratergy.
TECHNICAL SKILLS:
Languages: Java, Python, Shell, J2EE, C, C++, SQL, PL/SQL, HTML, DHTML, Java Script, XML, UML,Cold Fusion.
J2EE Technologies: JSP, Servlets, Tag Libraries, JSTL, EJB, JNDI, JDBC, JMS.
Frameworks: Apache Struts, Spring AOP, Hibernate, Junit,Greenplum Database, Apache Axis, JSF using ICEFaces
Application/Web Servers: IBM Web Sphere, Apache Jakarta Tomcat, BEA Weblogic, JEE5 Web services.
Web Services & XML: XML, XHTML, XSL, XSLT, CSS, SOAP, WSDL, SAX and DOM parsers,SOA.
IDE/ Tools: IBM Web Sphere Studio Application Developer WSAD 5.1.2, RAD 6.1/7, Eclipse
RDBMS: HBase, Oracle, DB2, SQL server 2000, Mysql
Hadoop/Bigdata: HDFS, MapReduce, Hive, Pig, Sqoop, Flume, Oozie, and ZooKeeper
Version Control: Rational Clear Case, Visual Source Safe, CVS
Methodoligies: Agile and Test Driven Development, SCRUM
PROFESSIONAL EXPERIENCE:
Confidential, Carlstadt, NJ
Sr. Hadoop Developer
Responsibilities:
- Working on the creation of business rules for Confidential Stores in Pig.
- Imported data from legacy systems to Hadoop using Sqoop and Apache Camel.
- Used Pig for data transformation
- Used Sqoop for ETL between Hadoop and structured database(RDBMS).
- Worked on Talend for string manipulations, lookup handling and ETL jobs
- Used Apache Spark for real time and batch processing
- Used Scala for developing many Spark applications
- Used Oozie for job scheduling
- Used Apache Kafka for handling log messages that are handled by multiple systems
- Used Scala in accessing the Hadoop data using Map R.
- Expert in parallel processing of the data in the Hadoop Clusters to perform a set of compuations using MPP databases.
- Worked on implementation of AWS Redshift.
- Used Talend as the support to the data integration platforms and OLAP applications.
- Design and implementation of High Availability feature for Search Engine
- Volume testing to calculate cluster's throughput
- Helped the team to increase the Cluster size from 22 to 30 Nodes and for the implementation of Hadoop security (Kerberos, Sentry and Hbase ACLs)
- Expert in data compression using Snappy.
- Worked on HCatalog, which allows PIG and Map Reduce to take advantage of the SerDE data format transformation definitions are already written on HIVE
- Worked on DevOps tools like Chef, Artifactory and Jenkins to configure and maintain the production environment
- Use Pig to transform data into various formats
- Stored processed tables in Cassandra from HDFS for applications to access the data in real time
- Worked on writing UDFs in Python for Pig
- Created ORCFile tables from the existing non-ORCFile Hive tables
Environment: Horton-works Data Platform 2.2, AWS, Pig, Hive, Spark, Kafka, Scala, Cassandra, Sqoop, Apache Camel, Oozie, HCatalog, Chef, Jenkins, Talend, Python, Artifactory, Avro, IBM Data Studio
Confidential, Chevy chase, Maryland
Sr. Hadoop Developer
Responsibilities:
- Analyzing the data and using PIG, HIVE for the loading of the data into HDFS.
- Vast use of Shell scripting for the loading of data into HDFS.
- Worked with the QA and Production team in the data loading process.
- Worked on TWS for the scheduling of the jobs.
- Data processing involved working on ANT BUILDS.
- Data loading involved creating Hive tables and partitions based on the requirement.
- Worked on various types of SERDE
- SSA (Standard Source Adapters), standard set of java and python libraries which AT&T is building to ensure consistency of load and extract job code for manageability, scalability and maintenance efficiency.
- Worked on HCatalog which allows PIG and Map Reduce to take advantage of the SerDE data format transformation definitions that we write for HIVE
- Worked on different UDFs in Python and Java, these are used in the MES solution to provide a way for DA’s and developers to encrypt, or decrypt, data.
- Worked on the implementation of Apache Knox for providing a single point authentication of Hadoop services.
- Worked on different file formats (ORCFILE, RCFILE, SEQUENCEFILE, TEXTFILE) and different Compression Codecs (GZIP, SNAPPY, LZO).
Environment: Hadoop, HDFS, Pig, Sqoop, Hive, Horton works distribution, Shell Scripting, Ubuntu, Linux Red Hat, JSON, Python.
Confidential, Bluebell, PA
Hadoop Developer
Responsibilities:
- Responsible for installing and configuring Hadoop MapReduce, HDFS also developed various MapReduce jobs for data cleaning.
- Installed and configured Hive to create tables for the unstructured data in HDFS
- Hold good expertise on major components in Hadoop Ecosystem including Hive, PIG, HBase, HBase-Hive Integration, Sqoop and Flume.
- Involved in loading data from UNIX file system to HDFS
- Expert in parallel processing(MPP) of databases to perform computation of a single task on the Hadoop cluster.
- Responsible for managing and scheduling jobs on Hadoop Cluster
- Responsible for importing and exporting data(RDBMS) into HDFS and Hive using Sqoop
- Experienced in running Hadoop streaming jobs to process terabytes of xml format data
- Experienced in managing Hadoop log files
- Worked on managing data coming from different sources
- Wrote HQL queries to create tables and loaded data from HDFS to make it structured, load and transform large sets of structured, semi structured and unstructured data
- Extensively worked on Hive for generating transforming files from different analytical formats to .txt i.e. text files enabling to view the data for further analysis
- Created Hive tables, loaded them with data and wrote hive queries that run internally in MapReduce.
- Wrote and modified store procedures enabling to load and modify data according to the project requirements
- Developed a continuous deployment pipeline using Jerkins, Chef and Shell scripts.
- Responsible for developing PIG Latin scripts enabling the extraction of data from the web server output files to load into HDFS
- Responsible for data injection and extracting to hadoop cluster from mainframe database.
- Worked on exploring the options from mainframe to Pig and Hive.
- Extensively used Flume to collect the log files from the web servers and then integrated these files into HDFS
- Responsible for implementing schedulers on Job Tracker enabling them to effectively use the resources available in the cluster for any given MapReduce jobs.
- Constantly worked on tuning the performance of the queries in Hive and Pig, making the queries work even more powerfully in processing and retrieving the data
- Supported Map Reduce Programs running on the cluster
- Created external tables in Hive and loaded the data into these tables
- Hands on experience in database performance tuning and data modelling
- Monitored the cluster coordination using Zookeeper
Environment: Hadoop, HDFS, MapReduce, Hadoop distribution of Cloudera, Hive, Cloudera, MapR, Java (jdk1.6), DataStax, Flat files, UNIX Shell Scripting, Oracle 11g 10g, PL SQL, SQL*PLUS, Toad 9.6, Windows NTp, Pig, Sqoop, HBase, Shell Scripting, Ubuntu, Linux Red Hat.
Confidential, Bloomington, IL
J2EE Developer
Responsibilities:
- Extensively used Hibernate in data access layer to access and update information in the database.
- Extensively used ICEFaces framework for its User Interface components and help navigation within the website.
- Customize CSS with ICEFaces Style- sheets for different styles.
- Used Perl and shell scripting to automate the batch process and run SQL scripts.
- Developed Web services -RESTful for getting credit card information from third party.
- Used SAX parser for parsing XML files.
- Used JMS API for asynchronous communication by putting the messages in the Message queue, such as PDF, Excel report generation.
- Involved to work with another developer to migrate an existing MS Access application to Cold Fusion
- Implemented various design patterns in the project such as Business Delegate, Session Façade, Data Transfer Object, Data Access Object, Service Locator and Singleton.
- Developed Stored Procedures for Oracle 10g database.
- Performed unit testing using JUNIT framework and used Test Cases for testing Action Classes.
- By Using SOA application we reused the software components.
- Used ANT scripts to build the application and deployed on WebSphere Application Server.
- Used Rational Clear Case and Clear Quest for version control and change management.
Environment: Java 1.5, J2EE, Hibernate, JMS, JSF, ICEFaces 3.2.0, XML, RestFul, JDBC, JavaScript, UML, Perl, HTML, JNDI, CVS, JUnit, Adobe ColdFusion, WebSphere Server 6.1, RAD 7, SOA, Rational Rose, Rational Clearcase, Rational Clear Quest, Oracle 10g.
Confidential
J2EE Developer
Responsibilities:
- Involved in Requirement Analysis, Development and Documentation.
- Used MVC architecture (Jakarta Struts framework) for Web tier.
- Participation in developing form-beans and action mappings required for struts implementation and validation framework using struts.
- Development of front-end screens with JSP Using Eclipse.
- Involved in Development of Medical Records module. Responsible for development of the functionality using Struts and EJB components.
- Coding for DAO Objects using JDBC (using DAO pattern).
- XML and XSDs are used to define data formats.
- Implemented J2EE design patterns value object singleton, DAO for the presentation tier, business tier and Integration Tier layers of the project.
- Involved in Bug fixing and functionality enhancements.
- Designed and developed excellent Logging Mechanism for each order process using Log4J.
- Involved in writing Oracle SQL Queries.
- Involved in Check-in and Checkout process using CVS.
- Developed additional functionality in the software as per business requirements.
- Involved in requirement analysis and complete development of client side code.
- Followed Sun standard coding and documentation standards.
- Participation in project planning with business analysts and team members to analyse the business requirements and translated business requirements into working software.
- Developed software application modules using disciplined software development process.
Environment: Java, J2EE, JSP, EJB, ANT, STRUTS1.2, Log4J, Weblogic 7.0, JDBC, MyEclipse, Windows XP, CVS, Oracle.
Confidential
Java Developer
Responsibilities:
- Technical responsibilities included high level architecture and rapid development.
- Design architecture following J2EE MVC framework.
- Developed interfaces using HTML, JSP pages and Struts -Presentation View.
- Involved in designing & developing web-services using SOAP and WSDL.
- Developed and implemented Servlets running under JBoss.
- Used J2EE design Patterns for the Middle Tier development.
- Used J2EE design patterns and Data Access Object (DAO) for the business tier and integration Tier layer of the project.
- Created UML class diagrams that depict the code’s design and its compliance with the functional requirements.
- Developed various EJBs for handling business logic and data manipulations from database.
- Designed and developed the UI using Struts view component, JSP, HTML, CSS and JavaScript.
- Implemented CMP entity beans for persistence of business logic implementation
- Development of database interaction code to JDBC API making extensive use of SQL Query Statements and advanced prepared statement.
- Involved in writing Spring Configuration XML files that contains declarations and other dependent objects declaration.
- Inspection/Review of quality deliverables such as Design Documents.
- Involved in creation running of Test Cases for JUnit Testing.
- Experience in implementing Web Services using SOAP, REST and XML/HTTP technologies.
- Used Log4J to print the logging, debugging, warning, info on the server console.
- Wrote SQL Scripts,Stored procedures and SQL Loader to load reference data.
Environment: J2EE (Java Servlets, JSP, Struts), MVC Framework, Apache Tomcat, JBoss, Oracle8i.
